Driftproof

We checked 162 quoted AI benchmark gaps against their own noise. 20 hold up.

Counts use each benchmark's own published intervals where they exist; where none exists, the gap is marked as not checkable.

Six frontier launch posts and nine public leaderboards, every quoted gap tested against the benchmark's own sampling noise, with a method frozen before the data was collected. Most gaps don't clear it. For 28 of them, the published data isn't even enough to check.

The scoreboard problem in one example

On SWE-bench Multilingual, the top two models score 72.7% and 66.3%. That looks decisive. It is a gap on 300 tasks, and this benchmark needs about 24 tasks of daylight before a gap at this size clears its own sampling noise. Six positions on that leaderboard sit inside the noise of first place.

What we did

We took the six most recent frontier launch announcements (Anthropic, OpenAI, Google and Mistral, September 22 to October 6) and nine public leaderboard views (SWE-bench's six, Terminal-Bench, Aider polyglot, SWE-bench Pro V2). We extracted every eligible head-to-head gap: 44 from the launch posts, 118 adjacent pairs from the leaderboards, 162 in all. Then we asked one question per gap: is it bigger than the sampling noise of the benchmark it comes from? The method (Wilson 95% intervals, Newcombe difference, a paired check where per-task results exist) was frozen before any data was read, and every number, source and hash is in the open repo.

Headline comparisons in each launch post: separated, not separated and not checkableseparatednot separatednot checkableP1 Anthropic Opus 5.5, 22 Sep832P2 OpenAI GPT-6, 22 Sep0 comparisonsP3 Anthropic Sonnet 5.5, 28 Sep322P4 OpenAI GPT-6.1, 29 Sep0 comparisonsP5 Google Gemini 4 Argon, 30 Sep15P6 Mistral Large 4, 6 Oct9
Headline comparisons in each launch post. Filled bars are separated, hollow bars are not, dashed bars are not checkable, and each bar is labelled with its count. P2 and P4 have zero comparisons. From the frozen FINDINGS read with the 8 October score-type audit; the detailed charts are in the data repository.
PostDateHeadline comparisonsSeparatedNot checkable
P1 Anthropic Opus 5.522 Sep1382
P2 OpenAI GPT-622 Sep000
P3 Anthropic Sonnet 5.528 Sep732
P4 OpenAI GPT-6.129 Sep000
P5 Google Gemini 4 Argon30 Sep15015
P6 Mistral Large 46 Oct909
All six posts441128

Each launch post's headline comparisons, from the frozen FINDINGS read with the 8 October score-type audit.

Finding one: most quoted gaps don't clear the noise

Of the 44 launch-post gaps, 11 clear it, 5 don't and 28 can't be checked. Of the 118 leaderboard pairs, 9 clear it and 109 don't. Only 14 pairs anywhere had published per-task results allowing the stronger paired test; one clears it. Small benchmarks are the worst: Terminal-Bench has 66 tasks, so on the simple reading two models need to be about 18 points apart, and 25 of its 27 quoted gaps are smaller than that. Read with its own published intervals where they match, 4 of the 27 separate, 15 don't and 8 can't be checked.

BenchmarkItemsSmallest gap that clears the noiseIn points
SWE-bench Verified500316.2 pp
SWE-bench Multilingual300248.0 pp
Terminal-Bench 4.066points only18.2 pp
Aider polyglot225points only9.3 pp

The smallest gap that clears each benchmark's own sampling noise, from the frozen FINDINGS. Terminal-Bench 4.0 and Aider polyglot scores are not one run over their items (Terminal-Bench averages repeated trials per task; Aider allows a second repair attempt), so their gap is given in points only, by the repo's gap rule.

Finding two: for most comparisons, you can't even check

This is the result we didn't expect. Many published scores are not one run over N questions; they're averages over repeated runs, with the repeat counts often unstated and harness, tools and effort settings differing between the two sides. For 483 of the 617 comparisons we examined (including an appendix of effort-curve contrasts), the published data does not support an honest repeat-aware interval at all. We withdraw those as claims rather than guessing, and 175 of them had looked separated under the simple reading. The replication crisis question for AI benchmarks is not whether the gap is significant. It is whether anyone published enough for the question to be answerable.

Finding three: when sources do publish error bars, verdicts move in both directions

Terminal-Bench publishes real 95% intervals. Using them, first place does separate from positions 3 through 15; second place is the only one the test can't tell apart from first. Four gaps that looked unresolved under the simple reading become separated under the source's own bars.* Error bars are not a pedantic garnish; they change the reading both ways.

* One of the four, on Chartography, uses an interval read off Chartography's published graphs, not stated numbers.

Also worth knowing: GPQA Diamond, AIME 2025 and MMLU-Pro, until recently fixtures of every launch post, appear in none of the six. And one result that looked clearly separated compares a lab-run score against a competitor's self-reported one; it is now marked not checkable.

Check any gap yourself

The gap calculator takes two scores and an N and gives the verdict in questions, not percentage points. The data, method, archived-source hashes and every one of the 617 rows are in the repo. None of this says two models are equal; no separation detected means this benchmark, at this size, with this data, cannot tell.

Method per the Driftproof paper. A gap you can't distinguish from noise is not a lead; it's a coin flip with a press release.