The Gate That Approved the Broken Draft
A new meaning-check flagged two clean AI drafts and missed two broken ones. Here’s what that reveals about trusting an automated writing gate.
A new meaning-check flagged two clean AI drafts and missed two broken ones. Here’s what that reveals about trusting an automated writing gate.
My own tool hid every free model on OpenRouter, including one that turned out to be GLM-5.3-Flash. Free changed nothing about how I judge a draft.
I tested whether a second AI model can fact-check the first one’s draft against the source. It caught 4 of 10 fabrications, and once it noted the mismatch and still marked the sentence sourced.
I built a safety check for AI delegation, rolled it out to a second site, and gave it two real jobs. Neither one used it. Here is what that tells me.
I gave four cheap models a smaller job: reorganize my research notes. Invented figures went to zero. Two of the four deleted the caveats instead.
I scored eleven AI-written drafts of the same article, from $0.003 to $0.10 each. All eleven failed the same item, and paying ten times more did not raise the score.
After four incidents, I stopped trusting written rules alone and started wiring Claude Code hooks into scripts that block the risky action outright.
I had 3 days of full access to a new frontier model before my subscription capacity reset. Here’s why I used it to audit my own AI automation system instead of building something new, and what it found.
The biggest upgrade to my AI fact-checking wasn’t a smarter model or a longer prompt — it was switching the context. Here’s the simple move that catches what an AI’s self-check misses.
Scheduling a WordPress post isn’t the finish line. Here’s my Draft, Scheduled, Synced flow, the post-publish check most people skip, and the two silent failures that taught me to close the loop.