I recently froze a table of behavioral metrics computed from 281 of my own AI-agent sessions and ran it through two separate checks before I let myself trust any conclusion drawn from it. Both checks passed. Neither one verified the thing I actually needed to know — and both failures turned out to have the exact same shape.
The underlying table is observational, closer to a photograph of my own routing policy than a fact about the models sitting in it, and I've already written about that failure mode on its own. This post assumes that caveat and goes one layer past it: what happens after you've accepted it, built real checks anyway, and watched them both come back green.
The corpus was 281 sessions, broken into 4,818 threads and 4,415 rows of behavioral metrics, frozen so nothing in it could shift under me mid-analysis. I had my own relaunched main model read the frozen table and write down its claims first, sealed before any outside model saw the data. Then I sent the identical table to six outside vendor families and asked each one to surface claims of its own. Only three of those six made it back cleanly — one hit a free-tier quota of zero, one threw an internal error, one had its response stream truncated mid-read. Between my own sealed read and the three outside reads that did return, four independent families ended up looking at the same numbers: my relaunched model, xAI's Grok, DeepSeek, and a Google open model.
Then I ran the part of the check that actually matters: a deterministic auditor that pulled every number any of the four models had cited, recomputed it directly against the frozen database, and flagged anything off by more than 5%, or any number that didn't correspond to anything in the data at all. 70 citations went in. 70 came back PASS. Zero hallucinated numbers, across four models with no shared training lineage.
That's a genuinely good result, not a strawman I'm setting up to knock down. It rules out a real failure mode — models inventing statistics that merely sound plausible. What it doesn't rule out is a model reading a real, correctly cited number and drawing the wrong conclusion from it. That happened once, cleanly enough to use as the example.
One row in the table belonged to Claude Sonnet 5's main-role threads — a small population, nine of them — where four behavioral metrics, including tool_error_rate, all read exactly 0.000. DeepSeek's model cited that 0.000 correctly and concluded it meant precise, error-free tool use, then built a routing suggestion on top of that reading. My own model, looking at the same row, flagged it as a likely measurement artifact instead. Sonnet 5 barely acts in a main-role thread in this corpus, and most of these rates are derived from file edits it almost never makes there: with no edits underneath the ratio, the rate isn't low, it's undefined and defaults to zero. A median tool-error rate of zero across those nine sparse threads is the same kind of non-signal — tie-degenerate, not a track record. The zero meant "not measured," not "no errors."
The audit never had a chance to catch this, because there was nothing wrong with the citation to catch. DeepSeek quoted 0.000 accurately. The mistake happened one inferential step later, in the word "therefore" — and a citation-accuracy audit has no way to see that step. It checks arithmetic, not reasoning.
The fan-out audit was casual by design: let several models free-associate over a frozen table and grade the arithmetic afterward. One metric got stricter treatment. Before looking at the data, I pre-registered a formal equivalence test — TOST, two one-sided tests — on same-file re-edit rate, comparing my relaunched main model's main-role threads against Claude Opus 4.8's, the older flagship. I committed to the margin in advance, 0.75 of the pooled standard deviation, and to a simple rule: the test only passes if the whole confidence interval around the observed difference sits inside that margin.
| Quantity | Value |
|---|---|
| Relaunched model, main-role sessions (n) | 54 |
| Claude Opus 4.8, main-role sessions (n) | 145 |
| Observed mean difference | +0.089 |
| Pre-registered equivalence margin | ±0.222 (0.75 × pooled SD) |
| 90% confidence interval | [+0.015, +0.163] |
| Pre-registered test verdict | Pass |
| Adversarial review verdict | Unsupported |
By the rule I'd committed to, this passed clean: the full interval sat inside the margin. That's the good kind of check — a number graded against a threshold I picked before I knew whether it would be convenient. So I treated "passed" as license to say the two models were close enough on this axis to stop worrying about it.
I sent the result to two adversarial reviewers from two different outside model families — one instructed to attack the statistics, one instructed to attack the operational reasoning. Both came back with the same verdict: unsupported. Not wrong about the arithmetic — wrong about what I'd let the arithmetic mean. Three reasons, and all three held up under a second look:
Line the two failures up and they match. The citation audit's definition of "pass" was: did you copy the digits. The equivalence test's definition of "pass" was: does the gap sit inside a zone I chose in advance. Neither one asked the question I actually wanted answered — did you understand this, and are these two models actually interchangeable. A passing score on either check told me something true and narrow, and I filled in a broader claim myself that the check never made.
The honest version of where this leaves me: the observed behavioral signal for telling my relaunched main model and Claude Opus 4.8 apart is weak — one metric, one axis, a real but small and bounded gap. That's the whole claim that survives. "They're interchangeable" is overreach. "Route on price alone" is overreach. Both were the conclusions I was reaching for before the adversarial pass, and both got killed by reviewers doing the one thing neither check I'd built was designed to do.
One more detail worth keeping: the reviewer that attacked the operational reasoning in Act Two was DeepSeek's model — the same family that misread Sonnet 5's zeroed row in Act One. Same family, opposite outcome, two different jobs. The lesson isn't "trust this vendor less." It's that a blind spot belongs to the check, not permanently to whichever model happens to be running it. A citation audit will always be blind to interpretation, no matter which model passes it. An adversarial refute pass will catch what a citation audit can't, no matter which model runs it either.
Two changes, both narrow on purpose:
Q. What does a numeric-citation audit actually verify?
It verifies that a cited number matches its source within tolerance, nothing more. My audit recomputed all 70 numbers that four different model families had cited from a frozen dataset, checked each one against the source within 5%, and found zero mismatches and zero hallucinated numbers. That's a real result, but it's a copy-fidelity check. It can't tell you whether a model drew the right conclusion from a number it quoted correctly.
Q. How can a model cite a number accurately and still misread it?
One row in my dataset, Claude Sonnet 5's main-role threads, a population of just nine, had four behavioral metrics read exactly zero, including its tool-error rate. That's a measurement artifact, not a performance: this model barely acts in a main-role thread here, so the edit-derived rates have no edits underneath them and default to zero, while a median tool-error rate over nine sparse threads is tie-degenerate rather than a real track record. One outside model, DeepSeek, cited the zero correctly and concluded it meant precise, error-free tool use, then built a routing suggestion on that reading. The citation was accurate. The inference one step later wasn't, and a citation-accuracy audit has no way to catch that, because it only checks arithmetic.
Q. What is a TOST equivalence test, and why did mine still get overturned after it passed?
TOST, two one-sided tests, checks whether an observed difference is small enough to sit entirely inside a margin you commit to before seeing the data. Mine did: comparing my relaunched main model (54 main-role sessions) to Claude Opus 4.8 (145 main-role sessions) on same-file re-edit rate, the mean difference was +0.089 with a 90% confidence interval of [+0.015, +0.163], fully inside my pre-registered margin of ±0.222. Two adversarial reviewers from different model families still called it unsupported: the margin was wide enough to pass almost anything, the same interval excluded zero — meaning a real, directional gap exists — and the metric never touched correctness, task success, or latency, the things that would actually justify a routing decision.
Q. Was either check worthless, then?
No. The citation audit still rules out a real failure mode, models fabricating statistics, across four independently trained model families with zero hallucinations. The equivalence test still ruled out a large difference between the two models. Both are genuine, useful floors. The mistake was letting a pass stand in for a broader claim neither check was built to make.
Q. What changed in how you run these checks now?
Two things. Any analysis I delegate to an LLM now gets a second, adversarial interpretation pass from a different model family, not just a citation-accuracy pass, specifically instructed to argue against the conclusions rather than just check the digits. And any equivalence claim has to report whether its confidence interval excludes zero, win or lose, because a wide margin with a directional interval is a different finding from a tight margin with an interval that straddles zero, even though both would formally pass.