We checked 162 quoted AI benchmark gaps against their own noise. 20 hold up.
Counts use each benchmark's own published intervals where they exist; where none exists, the gap is marked as not checkable.
Six frontier launch posts and nine public leaderboards, every quoted gap tested against the benchmark's own sampling noise, with a method frozen before the data was collected. Most gaps don't clear it. For 28 of them, the published data isn't even enough to check.
The scoreboard problem in one example
On SWE-bench Multilingual, the top two models score 72.7% and 66.3%. That looks decisive. It is a gap on 300 tasks, and this benchmark needs about 24 tasks of daylight before a gap at this size clears its own sampling noise. Six positions on that leaderboard sit inside the noise of first place.
What we did
We took the six most recent frontier launch announcements (Anthropic, OpenAI, Google and Mistral, September 22 to October 6) and nine public leaderboard views (SWE-bench's six, Terminal-Bench, Aider polyglot, SWE-bench Pro V2). We extracted every eligible head-to-head gap: 44 from the launch posts, 118 adjacent pairs from the leaderboards, 162 in all. Then we asked one question per gap: is it bigger than the sampling noise of the benchmark it comes from? The method (Wilson 95% intervals, Newcombe difference, a paired check where per-task results exist) was frozen before any data was read, and every number, source and hash is in the open repo.
| Post | Date | Headline comparisons | Separated | Not checkable |
|---|---|---|---|---|
| P1 Anthropic Opus 5.5 | 22 Sep | 13 | 8 | 2 |
| P2 OpenAI GPT-6 | 22 Sep | 0 | 0 | 0 |
| P3 Anthropic Sonnet 5.5 | 28 Sep | 7 | 3 | 2 |
| P4 OpenAI GPT-6.1 | 29 Sep | 0 | 0 | 0 |
| P5 Google Gemini 4 Argon | 30 Sep | 15 | 0 | 15 |
| P6 Mistral Large 4 | 6 Oct | 9 | 0 | 9 |
| All six posts | 44 | 11 | 28 |
Each launch post's headline comparisons, from the frozen FINDINGS read with the 8 October score-type audit.
Finding one: most quoted gaps don't clear the noise
Of the 44 launch-post gaps, 11 clear it, 5 don't and 28 can't be checked. Of the 118 leaderboard pairs, 9 clear it and 109 don't. Only 14 pairs anywhere had published per-task results allowing the stronger paired test; one clears it. Small benchmarks are the worst: Terminal-Bench has 66 tasks, so on the simple reading two models need to be about 18 points apart, and 25 of its 27 quoted gaps are smaller than that. Read with its own published intervals where they match, 4 of the 27 separate, 15 don't and 8 can't be checked.
| Benchmark | Items | Smallest gap that clears the noise | In points |
|---|---|---|---|
| SWE-bench Verified | 500 | 31 | 6.2 pp |
| SWE-bench Multilingual | 300 | 24 | 8.0 pp |
| Terminal-Bench 4.0 | 66 | points only | 18.2 pp |
| Aider polyglot | 225 | points only | 9.3 pp |
The smallest gap that clears each benchmark's own sampling noise, from the frozen FINDINGS. Terminal-Bench 4.0 and Aider polyglot scores are not one run over their items (Terminal-Bench averages repeated trials per task; Aider allows a second repair attempt), so their gap is given in points only, by the repo's gap rule.
Finding two: for most comparisons, you can't even check
This is the result we didn't expect. Many published scores are not one run over N questions; they're averages over repeated runs, with the repeat counts often unstated and harness, tools and effort settings differing between the two sides. For 483 of the 617 comparisons we examined (including an appendix of effort-curve contrasts), the published data does not support an honest repeat-aware interval at all. We withdraw those as claims rather than guessing, and 175 of them had looked separated under the simple reading. The replication crisis question for AI benchmarks is not whether the gap is significant. It is whether anyone published enough for the question to be answerable.
Finding three: when sources do publish error bars, verdicts move in both directions
Terminal-Bench publishes real 95% intervals. Using them, first place does separate from positions 3 through 15; second place is the only one the test can't tell apart from first. Four gaps that looked unresolved under the simple reading become separated under the source's own bars.* Error bars are not a pedantic garnish; they change the reading both ways.
* One of the four, on Chartography, uses an interval read off Chartography's published graphs, not stated numbers.
Also worth knowing: GPQA Diamond, AIME 2025 and MMLU-Pro, until recently fixtures of every launch post, appear in none of the six. And one result that looked clearly separated compares a lab-run score against a competitor's self-reported one; it is now marked not checkable.
Check any gap yourself
The gap calculator takes two scores and an N and gives the verdict in questions, not percentage points. The data, method, archived-source hashes and every one of the 617 rows are in the repo. None of this says two models are equal; no separation detected means this benchmark, at this size, with this data, cannot tell.
Method per the Driftproof paper. A gap you can't distinguish from noise is not a lead; it's a coin flip with a press release.