Why the engine holds the dice
The model marking its own homework is worthless
Long scripted playthroughs found what a DM asked to referee itself actually does. It fudged downward out of politeness. It repeated whole sentences byte for byte. It promoted a typo into a permanent character. It lost your gold between two context windows. It announced rules no book contains.
Every one of those was fixed by moving the job out of the prompt and into code with a test on it: seeded rolls, a resource ledger, an echo guard, a rule verifier that quotes the book. The measurements that proved each fix are written down in the project's own records, and several of them are on the home page.