Too close to call a winner
A is ahead by 10 questions.
Difference needed for stronger evidence: about 6.2 points.
72% beat 70% by 10 questions, but with only 500 questions a gap this small can happen from sampling noise.
Pick a benchmark and type two scores to see whether the gap clears the benchmark's own noise, or could come down to luck in which questions were asked.
Too close to call a winner
A is ahead by 10 questions.
Difference needed for stronger evidence: about 6.2 points.
72% beat 70% by 10 questions, but with only 500 questions a gap this small can happen from sampling noise.
For each side: its average score, how far its runs usually sit from that average, and how many runs it had.
no separation detected at this sample size
72.0% vs 70.0% on 500 questions is a lead of 10 questions; at this size, noise alone can span 30 questions.
67.9 to 75.865.8 to 73.9-3.6 to 7.6 points30 questions (source). A lead needs at least 8 questions, 26.7 points, before it clears this benchmark's sampling noise.
500 questions (source). A lead needs at least 31 questions, 6.2 points, before it clears this benchmark's sampling noise.
12032 questions (source). A lead needs at least 153 questions, 1.3 points, before it clears this benchmark's sampling noise.
Method frozen 7 Oct 2026: Wilson 95% intervals and Newcombe's hybrid difference, written out in gap.js; the verdict words are the methodology page's.
Mean and standard deviation in percent, and the run count, per arm. This is the rule the published reports use: bands that do not overlap read separated. Methodology.
On the runs tab, the spread between runs is one standard deviation, in points.
Piece: AI benchmark gaps vs sampling noise, by Maverick, applies this method to the gaps quoted in launch posts and on public leaderboards.
Method per the Driftproof paper (driftproofhq.com/paper): the band rule and the verdict words.