Driftproof

Benchmark gap calculator

Pick a benchmark and type two scores to see whether the gap clears the benchmark's own noise, or could come down to luck in which questions were asked.

How this works

Two scores (item noise)

Too close to call a winner

A is ahead by 10 questions.

Difference needed for stronger evidence: about 6.2 points.

72% beat 70% by 10 questions, but with only 500 questions a gap this small can happen from sampling noise.

Show the statistics

no separation detected at this sample size

72.0% vs 70.0% on 500 questions is a lead of 10 questions; at this size, noise alone can span 30 questions.

Score A 72.0%, interval 67.9 to 75.8; Score B 70.0%, interval 65.8 to 73.9. Result: no separation detected at this sample size.60%65%70%75%80%Score AScore B-10-50510A minus B
Score A, Wilson 95%
67.9 to 75.8
Score B, Wilson 95%
65.8 to 73.9
A minus B, Newcombe
-3.6 to 7.6 points

Three presets, at the edge of noise

AIME 2025

30 questions (source). A lead needs at least 8 questions, 26.7 points, before it clears this benchmark's sampling noise.

SWE-bench Verified

500 questions (source). A lead needs at least 31 questions, 6.2 points, before it clears this benchmark's sampling noise.

MMLU-Pro

12032 questions (source). A lead needs at least 153 questions, 1.3 points, before it clears this benchmark's sampling noise.

What this does not tell you

  • Item noise is not run noise: the two tabs answer different questions and use different rules.
  • No separation detected is not equal ability: it means this sample size could not tell them apart.
  • Both intervals assume independent questions.

How this works

Method frozen 7 Oct 2026: Wilson 95% intervals and Newcombe's hybrid difference, written out in gap.js; the verdict words are the methodology page's.

Mean and standard deviation in percent, and the run count, per arm. This is the rule the published reports use: bands that do not overlap read separated. Methodology.

On the runs tab, the spread between runs is one standard deviation, in points.

Piece: AI benchmark gaps vs sampling noise, by Maverick, applies this method to the gaps quoted in launch posts and on public leaderboards.

Presets and their sources

Method per the Driftproof paper (driftproofhq.com/paper): the band rule and the verdict words.