Quick conclusion
Four of the models from part one. The same day. A different job.
In part one, every draft I bought invented experience that was not in the material I supplied. Eleven submissions, eleven failures, and paying ten times more did not change it.
So I stopped changing the model and changed the job. Instead of “write the article,” I asked for something smaller: take these research notes and reorganize them.
Invented figures went to zero. All four models. Then I compared each output against the source and found what had gone missing instead: the sentences that said “not confirmed yet.”
Two of the four deleted them. The one that deleted the most was the fastest and the shortest.
Why the first test was unfair
I did not work this out by studying my rubric. I worked it out by looking at the material.
Part one asked the models to rebuild an article about running this site. That material only exists here. It is my own operating data, and no instruction is going to give a model access to it. If you require an experience that cannot be retrieved, you get an invented one.
The notes I used this time come from another site I run, where the articles are built from public web research instead of from my own logs. That difference is the whole point. On that kind of material, a model’s own reading can reach the source. It is not being asked to know something only I know.
That was the actual reason I re-ran it. The tidier version, that I had built a test which required invention and then graded models for inventing, is also true. It just came afterward.
The job I gave them
One prompt, one pass, no conversation. Take two research files, roughly 9,200 tokens of them, and return six sections: confirmed facts with sources, a table of figures, everything still unconfirmed, contradictions between sources with both sides kept, notes for whoever writes the article, and a self-check.
The four models are the same ones from part one, and I am keeping the labels. A is the cheapest model you already met, and B is the one that scored best on agreement.
C is not in this round. D and E come from the higher tier I tested at the end of part one.
All four together cost under five cents.
Then I did the thing I should have done the first time. I wrote a check that compares input to output instead of reading the output on its own: figures that appear from nowhere, names that appear from nowhere, links that appear from nowhere, and whether every uncertain item in the source is still uncertain in the result.
Nobody made anything up
Zero invented figures, from all four models. These are the same models that had failed this exact item in every submission they made earlier that day.
I want to be careful about what that proves. The models did not improve. Nothing about them changed between the two runs.
The invention in part one was not a defect I fixed. It was a requirement I had written into the job without noticing.
All four also followed the six-section format exactly. Format instructions get followed. That was true in part one too. What does not get followed is “keep the meaning.”
Then I checked what was missing
The source notes carry 24 lines flagged as uncertain: unverified figures, pages that disagree, values someone needed to look at again before publishing.
| Model | Source length kept | Uncertain lines returned | Caveats lost | Result |
|---|---|---|---|---|
| A | 97% | 29 | 0 | Pass |
| E | 79% | 30 | 0 (2 flagged, verified intact) | Conditional |
| B | 47% | 12 | 8 | Fail |
| D | 26% | 11 | 11 | Fail |
The correlation is not subtle. The harder a model compressed, the more of the “not confirmed yet” went first.
Which makes sense. In ordinary writing, hedges are the most cuttable thing on the page. In a research note the hedge is the content.
Delete it and the sentence still reads fine. It just reads as a fact now.
That is worse than invention, and it is why I had missed it. An invented figure leaves a trace you can find. A caveat that has been removed looks exactly like a fact that was never in doubt.
Here is the clearest case. The source had a nine-row table where the number of columns changed from row to row, with a warning attached that said so. A reproduced all nine rows and kept the warning sitting right underneath.
D kept three rows, renamed the table so it only claimed to cover those three, and dropped the warning. Nothing D wrote is false. What is left just looks clean, and worse, it looks checked.
D also dropped the line saying older third-party pages no longer match the current source. The original notes had marked that as the single most important thing to get right.
The model that passed barely summarized anything
A is the cheapest model in the lineup, about a third of a cent for the job, and it was the only pass. That reads like a good headline and it is the wrong one.
A passed because it barely summarized. It kept 97 percent of the source length and mostly moved things into the six sections. What I bought there is not judgment. It is formatting, the manual labor of putting each piece where it belongs.
A was not spotless either. It echoed back the name of one of the files I had sent it, which is not a phrase that exists inside the notes. My log records that run as “check,” not “clear.”
E, the most expensive of the four, preserved meaning just as well. Two caveats my check reported as missing turned out to be intact, reworded.
But E did two things I could not wave through. Where the source gave a bare path, E supplied the hostname and produced a complete link. The link is correct. It is also not in the source, and on this job, supplying what is missing is the exact behavior I am screening for.
E then ran out of output budget and stopped partway, so its self-check never arrived at all.
Of the three self-checks that did arrive, all three claimed a perfect eight out of eight. Measured against the source, each was wrong on up to two of the five items I can check by machine. Part one’s conclusion survives the change of job: a self-check is worth requiring as a gap to measure, never as a claim to read.
What this check cannot do
The source notes said the service supported two platforms. In reality only one account tier had access to the first one. A human fact-check caught that before publication, and it was logged as the most serious correction in that article.
A and B both wrote “both platforms,” faithfully.
So this check measures whether the contractor damaged the material. It does not measure whether the material is true. It does not replace fact-checking, and the day I start feeling like it does is the day it becomes dangerous.
I moved my own goalpost, and here it is
Before running anything, I wrote down the pass condition: zero caveats lost, not negotiable afterward.
On the first run, my check reported that A had lost two. I looked at them. Both were still there in full, with the warning symbol stripped off the front. So I split the measure in two, content gone versus marker gone, and judged on the first.
I still believe that is what “lost” always meant. But I changed a criterion mid-experiment, after seeing a result, which is precisely the move you are supposed to distrust. Without that change, the only model that passed would have failed.
The other direction is just as real. A check that cries wolf gets switched off, and then it starts rejecting good work. Next time I write both tiers down before the run instead of during it.
What I approved
A is now cleared for this one job, reorganizing research notes, on the condition that every output goes through the input-to-output check and I fix whatever it flags.
That is not a statement of trust. The honest answer to “why did you approve it” is not that the test made me comfortable. It is that the cost and the risk balanced at this price.
Which is exactly why monitoring is a condition and not a good intention. This is the same move as replacing rules with gates: the check writes a line to a log every time it runs, so if the log stops growing, the check stopped running. I would rather learn that from a missing line than from a published article.
FAQ
Does this mean cheap models are good at summarizing?
The opposite, on this data. The model that compressed hardest lost the most, and the only pass compressed almost nothing. Asking for a summary is asking for judgment about what matters, and that is the part that failed. Asking for your own material rearranged is a different request, and that one worked.
Why not just tell the model to keep every caveat?
I did. It is one of the six required sections, and all four models wrote that section. Two of them wrote it while dropping items on the way in. Instructions produce the shape of compliance. Only comparing input to output produces the fact of it.
Would a more expensive model fix this?
Not reliably. The two models that preserved everything were the cheapest and the most expensive of the four. Price predicted nothing again. Output length predicted everything.
Do you still read the output yourself?
All of it, for now. The check catches things added and things removed. It cannot catch a plausible rearrangement that quietly changes what a figure refers to, and that part is still mine.
Next steps
Two experiments in, and I keep learning more about my own process than about the models. Part one measured the wrong thing. Part two measured the right thing and found that the interesting failure is a silent one.
The next test is not another model. It is splitting the work: one model does the research pass, a different one checks it with no memory of having written it, and I find out whether the total comes out better than any single model does alone. I have not run that yet, so I have nothing to report about it.
After that, this stops being an experiment and becomes a quarterly measurement. The log I mentioned has four rows in it today, which is not a finding. Ask me in three months.