The Gate That Approved the Broken Draft
A new meaning-check flagged two clean AI drafts and missed two broken ones. Here’s what that reveals about trusting an automated writing gate.
A new meaning-check flagged two clean AI drafts and missed two broken ones. Here’s what that reveals about trusting an automated writing gate.
My own tool hid every free model on OpenRouter, including one that turned out to be GLM-5.3-Flash. Free changed nothing about how I judge a draft.
I tested whether a second AI model can fact-check the first one’s draft against the source. It caught 4 of 10 fabrications, and once it noted the mismatch and still marked the sentence sourced.
I built a safety check for AI delegation, rolled it out to a second site, and gave it two real jobs. Neither one used it. Here is what that tells me.
I gave four cheap models a smaller job: reorganize my research notes. Invented figures went to zero. Two of the four deleted the caveats instead.
I scored eleven AI-written drafts of the same article, from $0.003 to $0.10 each. All eleven failed the same item, and paying ten times more did not raise the score.