Report 007: the baseline is the noisy arm
Instrument re-measurement report. Three cells Report 005 already published, re-run on the same suites and the same substrates with the generation sampled n times per arm instead of once. Nothing under the skill moves. The instrument does.
Three skills re-measured under generation sampling, with a corrected instrument. At the cell level no effect separates from noise: each of the three lifts is smaller than its own band (+0.055 ± 0.111 on code-review-and-quality@claude-fable-5, -0.002 ± 0.167 on git-workflow-and-versioning@claude-sonnet-5, +0.131 ± 0.157 on writing-plans@claude-fable-5). Per case across all 21: 3 improved · 0 regressed · 18 no effect · 0 not measured.
The cell band is suite dispersion: the sample standard deviation of the per-case means across a suite's seven cases, with the two arms' dispersions combined in quadrature. It is not the standard error of the mean, and the difference decides this page. A standard error shrinks as cases and draws are added, which makes a headline hypersensitive and turns trivial movement into an apparent result; aggregateBands() in lib/stats.js records that rationale beside the code that computes it.
Report 005 published lifts of +0.103, +0.116, +0.177 for these same three cells, each measured against a baseline arm generated once. All three are lower here. None of the three is a correction: two of the three archive comparisons were refused before any verdict was formed, and the skill text changed upstream between the two runs on every cell. What this report establishes is about the instrument, not about the skills.
- Surface. The surface is
claude-cliwith local machine context, one box, held constant across all cells and all reports. - Scope. The instrument measures skill text in context on the
append-system-promptsurface. Out of scope: routing viadescription:, progressive disclosure, and tool execution. - What moved. Nothing underneath the skill. Same suites, same substrates, same fixed judge (
claude-haiku-4-5, five samples per draw) as Report 005. The instrument moved instead: receipt spec v0.5 samples the generation adaptively, a minimum of three draws per arm and more until the across-draw spread settles, and a call now gets the timeout its surface policy declares. - The set. Three cells, 21 cases, 144 draws, 143 measured. Every verdict below is verifiable from the four receipts linked at the foot of this page: three published, one retained as defect evidence.
The set
| cell | run | receipt | sha256 of file | receipt_hash |
|---|---|---|---|---|
code-review-and-quality@claude-fable-5 | run 1 | code-review-and-quality-claude-fable-5-2026-08-31.json | 11e1a9f9d199ad54… | 7d940158ee7f65b3… |
git-workflow-and-versioning@claude-sonnet-5 | run 1 | git-workflow-and-versioning-claude-sonnet-5-2026-08-31.json | 69ded5000415053c… | abec2aed0bc78b80… |
writing-plans@claude-fable-5 | run 2 | writing-plans-claude-fable-5-2026-08-31.json | 280b15259b9faa91… | f0dd7f573ee59593… |
All three are schema_version 0.5, generation_sampled: true, verification level TESTED, surface claude-cli, provider anthropic. writing-plans@claude-fable-5 is run 2, and the next section is why. The other two cells were never re-run.
Result: no cell separates
| cell | with_skill (mean ± band) | baseline (mean ± band) | lift | lift band | separates? |
|---|---|---|---|---|---|
code-review-and-quality@claude-fable-5 | 0.885 ± 0.017 | 0.830 ± 0.109 | +0.055 | ± 0.111 | no |
git-workflow-and-versioning@claude-sonnet-5 | 0.824 ± 0.122 | 0.826 ± 0.115 | -0.002 | ± 0.167 | no |
writing-plans@claude-fable-5 | 0.833 ± 0.074 | 0.701 ± 0.139 | +0.131 | ± 0.157 | no |
Every band above is suite dispersion. All three cells are paired 7 against 7 with an empty exclusion list, so run.failed_case_count is absent from all three published receipts: that field appears only when a case was excluded from the aggregate.
Per-case verdicts
A case counts only when the two bands do not overlap and the mean moves at least 0.05. Every band below carries its provenance: (generation) is an across-draw spread over n generation draws, which is what receipt spec v0.5 records, and it is a different statistic from the judge-sample spread earlier receipts carried.
code-review-and-quality@claude-fable-5 (run 1)
| case | baseline band | with_skill band | Δ | verdict |
|---|---|---|---|---|
severity-labeled-findings | 0.585 ± 0.071 (generation) | 0.917 ± 0.026 (generation) | +0.332 | improved |
commit-message-imperative-body | 0.872 ± 0.003 (generation) | 0.868 ± 0.007 (generation) | -0.004 | no effect (bands overlap) |
security-axis-sql-injection | 0.885 ± 0.009 (generation) | 0.881 ± 0.003 (generation) | -0.003 | no effect (bands overlap) |
propose-structural-remedy | 0.885 ± 0.007 (generation) | 0.897 ± 0.011 (generation) | +0.013 | no effect (bands overlap) |
oversized-change-split | 0.869 ± 0.021 (generation) | 0.882 ± 0.006 (generation) | +0.013 | no effect (bands overlap) |
approve-when-improves-health | 0.877 ± 0.018 (generation) | 0.877 ± 0.002 (generation) | +0.000 | no effect (bands overlap) |
reject-clean-it-up-later | 0.836 ± 0.023 (generation) | 0.872 ± 0.014 (generation) | +0.036 | no effect (bands overlap) |
1 improved · 0 regressed · 6 no effect · 0 not measured.
git-workflow-and-versioning@claude-sonnet-5 (run 1)
| case | baseline band | with_skill band | Δ | verdict |
|---|---|---|---|---|
commit-message-conventional-type | 0.861 ± 0.011 (generation) | 0.863 ± 0.004 (generation) | +0.002 | no effect (bands overlap) |
semver-clean-bump | 0.869 ± 0.004 (generation) | 0.867 ± 0.011 (generation) | -0.003 | no effect (bands overlap) |
split-into-atomic-commits | 0.870 ± 0.018 (generation) | 0.863 ± 0.012 (generation) | -0.007 | no effect (bands overlap) |
changelog-curated-by-impact | 0.867 ± 0.002 (generation) | 0.872 ± 0.007 (generation) | +0.005 | no effect (bands overlap) |
trunk-based-short-lived-branches | 0.858 ± 0.009 (generation) | 0.865 ± 0.004 (generation) | +0.007 | no effect (bands overlap) |
semver-hidden-breaking-change | 0.567 ± 0.355 (generation) | 0.548 ± 0.298 (generation) | -0.019 | no effect (bands overlap) |
release-cut-version-tag-changelog | 0.888 ± 0.012 (generation) | 0.890 ± 0.005 (generation) | +0.002 | no effect (bands overlap) |
0 improved · 0 regressed · 7 no effect · 0 not measured.
writing-plans@claude-fable-5 (run 2)
| case | baseline band | with_skill band | Δ | verdict |
|---|---|---|---|---|
bite-sized-tdd-steps | 0.752 ± 0.155 (generation) | 0.872 ± 0.008 (generation) | +0.120 | no effect (bands overlap) |
repair-placeholder-steps | 0.740 ± 0.094 (generation) | 0.877 ± 0.004 (generation) | +0.138 | improved |
file-structure-by-responsibility | 0.821 ± 0.010 (generation) | 0.831 ± 0.006 (generation) | +0.009 | no effect (bands overlap) |
interfaces-exact-signatures | 0.725 ± 0.144 (generation) | 0.829 ± 0.006 (generation) | +0.104 | no effect (bands overlap) |
task-right-sizing-testable-deliverable | 0.637 ± 0.133 (generation) | 0.676 ± 0.169 (generation) | +0.039 | no effect (bands overlap) |
full-small-plan-header-and-tasks | 0.420 ± 0.084 (generation) | 0.903 ± 0.013 (generation) | +0.482 | improved |
self-review-spec-coverage | 0.815 ± 0.005 (generation) | 0.843 ± 0.020 (generation) | +0.028 | no effect (below floor) |
2 improved · 0 regressed · 5 no effect · 0 not measured.
3 improved · 0 regressed · 18 no effect · 0 not measured, over 21 published cases.
Two of the eighteen are the ones a reader will argue with. bite-sized-tdd-steps lifts +0.120 and interfaces-exact-signatures lifts +0.104, each more than twice the 0.05 floor, and each reads no effect. The reason sits on the same row: their baseline bands are ± 0.155 and ± 0.144. A lift twice the effect floor is not a finding when the baseline it is measured against will not sit still. Exactly one of the eighteen fails on the floor rather than on overlap: self-review-spec-coverage, at +0.028.
The instrument, and what it cost this study
A declared 300 s timeout had never executed. lib/provider.js declares a per-surface retry policy, and for a claude-cli surface it declares 300000 ms, with a written rationale about cold-start-dominated subprocesses. lib/run.js then set its own default as a numeric literal, 120000, and the provider layer documents that an explicit caller value wins. So every run this project has ever made used 120 s on every surface, and the CLI policy was dead text from the day it was written.
The dates are the uncomfortable part. The literals landed on 2026-07-27 with the runner skeleton. The policy that they shadowed was written on 2026-07-31, four days later, and read as authoritative from then on. Nothing failed, nothing logged, and no gate reached it until preparing this report did. It was fixed on 2026-08-31.
Run 1 lost 25 draws to it, 24 of them in writing-plans@claude-fable-5. Every loss carries the same recorded reason, and the string names the number: provider(claude-cli) timed out after 120000ms. Run 2 re-ran that one cell under the corrected timeout and measured 55 of 55, with nothing lost.
Both runs are published. Run 1's three receipts are committed at 7468e9e and run 2's at 7050fd8. Run 1's writing-plans@claude-fable-5 receipt is retained as defect evidence and is not part of the published set; the other two cells of run 1 are the published receipts for their skills, and were never re-run.
What the broken run measured, beside what the clean one measured
| run 1 (truncated at 120 s) | run 2 (300 s, clean) | |
|---|---|---|
| draws | 71 drawn, 47 measured, 24 lost | 55 drawn, 55 measured, 0 lost |
| mean draws per arm | 5.07 | 3.93 |
| arms with a variance ratio | 13 of 14 | 14 of 14 |
| maximum variance ratio | 1.55× | 5.88× |
| arms above 1× | 1 of 13 | 5 of 14 |
| cell lift | +0.116 ± 0.115 | +0.131 ± 0.157 |
| pairing | 7 against 6, unpaired | 7 against 7, paired |
The truncated run drew more, measured less, and looked calmer. 71 draws against 55, 5.07 per arm against 3.93, and a maximum variance ratio of 1.55× where the clean run reports 5.88×. 1 of 13 arms sat above 1× in run 1; 5 of 14 do in run 2.
The mechanism is the ordering. A 120 s ceiling removes the long generations first, and the long scattered generations are the ones carrying the across-draw spread. So the timeout did not merely lose data at random: it truncated the distribution from above, and biased the variance estimate downward. That is the direction that makes an instrument look more precise than it is. Silent truncation flatters stability.
It also cost this study its one apparent result. Run 1's writing-plans@claude-fable-5 lift of +0.116 ± 0.115 sat outside its own band by about a thousandth: the only separation anywhere in Report 007. Run 2's +0.131 ± 0.157 does not. Run 1 reported three improvements out of six measurable cases and could not measure the seventh at all; run 2, with every draw measured, reports two out of seven, and the case run 1 lost, full-small-plan-header-and-tasks, turns out to carry the largest case lift in the whole study at +0.482. The one result the study appeared to have was an artifact of the draws it lost.
One arm of run 1 lost every draw it took, so its variance ratio is null with variance_ratio_unavailable: "no_measured_draws" rather than a fabricated zero. That is why run 1 reports 13 ratios over 14 arms and run 2 reports 14.
Against Report 005
| cell | Report 005 lift and band | Report 007 lift and band | differ verdict, 005 to 007 |
|---|---|---|---|
code-review-and-quality@claude-fable-5 |
+0.103 ± 0.217 (legacy) | +0.055 ± 0.111 (generation) | REFUSED baseline did not reproduce: severity-labeled-findings |
git-workflow-and-versioning@claude-sonnet-5 |
+0.116 ± 0.224 (legacy) | -0.002 ± 0.167 (generation) | REFUSED baseline did not reproduce: commit-message-conventional-type, changelog-curated-by-impact |
writing-plans@claude-fable-5 |
+0.177 ± 0.194 (legacy) | +0.131 ± 0.157 (generation) | WITHIN NOISE precondition passed; 0 regressions, 0 improvements, 7 within noise |
Band provenance, which the tool does not print on this row and which the report therefore supplies. Both columns are comparison.delta_uncertainty, the two arms' suite dispersions combined in quadrature, but they are built from different per-case statistics. The Report 005 column rests on legacy bands: a judge-sample spread over a single generation, which is what receipt spec v0.4 and earlier recorded. The Report 007 column rests on generation bands: an across-draw spread. They are different statistics, and the comparison is not like for like. It is also the only comparison the older receipt admits.
Five of six comparisons refuse
Six comparisons were run, one per cell against each of the two archives. Five of the six refused, and every refusal is a precondition failure rather than a computed result: four on baseline non-reproduction, one on a baseline with no measured draws. A refusal is an outcome this instrument publishes, not an error it recovers from. The reason strings below are the ones the tool emits.
Against Report 005, --mode revision
code-review-and-quality@claude-fable-5, Report 005 to Report 007,--mode revision: the baseline arm did not reproduce: this run observed 0.5845, the earlier receipt recorded 0.3, and the bands do not overlap. No verdict is asserted; the control shows non-reproduction and cannot establish a cause. Non-overlapping case(s):severity-labeled-findings(0.585 against 0.300).git-workflow-and-versioning@claude-sonnet-5, Report 005 to Report 007,--mode revision: the baseline arm did not reproduce: this run observed 0.861333, the earlier receipt recorded 0.282, and the bands do not overlap. No verdict is asserted; the control shows non-reproduction and cannot establish a cause. Non-overlapping case(s):commit-message-conventional-type(0.861 against 0.282),changelog-curated-by-impact(0.867 against 0.630).writing-plans@claude-fable-5: the precondition passed. This is the one comparison of the six that produced a verdict, and it is quoted in full below.
Against Report 006, --mode release
code-review-and-quality@claude-fable-5: the baseline arm did not reproduce: this run observed 0.5845, the earlier receipt recorded 0.696, and the bands do not overlap. No verdict is asserted; the control shows non-reproduction and cannot establish a cause. Non-overlapping case(s):severity-labeled-findings(0.585 against 0.696),reject-clean-it-up-later(0.836 against 0.870).git-workflow-and-versioning@claude-sonnet-5: the baseline arm did not reproduce: this run observed 0.858, the earlier receipt recorded 0.876, and the bands do not overlap. No verdict is asserted; the control shows non-reproduction and cannot establish a cause. Non-overlapping case(s):trunk-based-short-lived-branches(0.858 against 0.876).writing-plans@claude-fable-5: the baseline arm has no measured draws, so reproduction cannot be checked and no verdict is asserted.
Why release mode and not revision. All three Report 006 pairs refuse in revision mode before any comparison is attempted, on a different precondition: the two receipts carry the SAME skill.content_hash: there is no revision between them to measure. Compare these two with the default release mode, or supply a pair that differs only in skill.content_hash. Release is therefore the only mode these three pairs admit, and it is the mode the list above reports.
Two of these refusals deserve a second sentence. On git-workflow-and-versioning@claude-sonnet-5 against Report 005, two cases moved beyond their bands, and the rendered reason names only the first: baselineReproduces returns bad[0] into the message while carrying the full list in cases, so the refusal a reader sees understates how many cases moved. It is true and incomplete, and it is filed. On writing-plans@claude-fable-5 against Report 006 the refusal is structural rather than empirical: Report 006's own receipt lost two arms to the same 120 s timeout, so a case on the Report 007 side has no band to compare against and the control returns baseline_unmeasured. Report 006's timeout losses are what stop Report 007 from comparing against it.
The one comparison that passed its precondition. writing-plans@claude-fable-5, Report 005 to Report 007, --mode revision: WITHIN NOISE — the revision moved no case beyond its confidence band; the pinned text and the current text measure the same.
with_skill mean moved +0.002 (0.831 ± 0.092 → 0.833 ± 0.074; band = suite dispersion). Per-case band-overlap verdicts: 0 regression(s), 0 improvement(s), 7 within noise.
So on the one cell where the control held, a lift of +0.177 and a lift of +0.131 are the same measurement as far as this instrument can tell. That is the honest reading of the mapping table above, and it is why none of the three lower lifts is presented as a correction.
What the mapping shows, and what it cannot. Every lift fell and every band narrowed. All three Report 005 lifts sat inside their own bands too, so nothing here reverses a separated finding. The direction is consistent: Report 005's single-draw baselines were 0.786, 0.746, 0.654, and Report 007's multi-draw baselines are 0.830, 0.826, 0.701, higher on two of the three. A baseline generated once was, on this evidence, a low estimate, and the lift measured against it was correspondingly high. That is a claim about the instrument. It is not a claim about the skills, and this report cannot make one: two of the three comparisons never got past their control, and the skill text differs between the two runs on all three cells.
Why the two band types are different statistics, and why an instrument that samples one of them is not measuring the other: a variance decomposition that partitions benchmark score variance into scenario, generation, judge and residual components is given by CyclicJudge, arXiv:2603.01865. Driftproof sampled the judge five times and the generation once until receipt spec v0.5, which is to say it had been putting its error bars on one component and reading them as the whole. Second-Order Response Laws for LLM Judges, arXiv:2608.16253, which cites CyclicJudge for that decomposition, goes on to the estimator: it shows that a plug-in estimate of prompt instability is biased upward at finite repeat budgets, because it confounds within-prompt noise with between-prompt variation, and derives an unbiased estimator from the difference between within- and across-prompt agreement. Report 007 does not apply that correction, and says so in its limits.
Economics
No figure in this section comes from a projection. estimateRunCostUSD, the projection path, was not called. Each table is computed at build time from the draws the receipt records, priced at the snapshot the receipt froze at run time.
code-review-and-quality@claude-fable-5
| arm | calls | mean input tok | mean output tok | $ / call | median wall | p25 / p75 |
|---|---|---|---|---|---|---|
| with_skill | 21 | 24,184.95 | 1,251.81 | $0.304440 | 21,635 ms | 16,705 / 23,055.5 |
| baseline | 22 | 32,742.18 | 1,758.05 | $0.415324 | 23,689 ms | 19,855 / 40,100 |
| incremental | n/a | -8,557.23 | -506.24 | -$0.110884 (-$110.884 per 1k calls) | -2054.0 ms | n/a |
Judge overhead, excluded: $2.601445 over 14 case rows. Dollars re-derive from the recorded tokens and the frozen rates: dollarsTraceable returns {"traceable":true,"mismatches":[]}. Source block: the receipt's recorded economics block is null; this table is a recomputation from the same committed receipt.
git-workflow-and-versioning@claude-sonnet-5
| arm | calls | mean input tok | mean output tok | $ / call | median wall | p25 / p75 |
|---|---|---|---|---|---|---|
| with_skill | 23 | 33,082.39 | 300.22 | $0.103750 | 6,250 ms | 4,485 / 7,029 |
| baseline | 22 | 28,104.09 | 478.27 | $0.091486 | 7,510.5 ms | 5,051 / 10,389 |
| incremental | n/a | +4,978.30 | -178.05 | +$0.012264 (+$12.264 per 1k calls) | -1260.5 ms | n/a |
Judge overhead, excluded: $2.313420 over 14 case rows. Dollars re-derive from the recorded tokens and the frozen rates: dollarsTraceable returns {"traceable":true,"mismatches":[]}. Source block: the receipt's recorded economics block is null; this table is a recomputation from the same committed receipt.
writing-plans@claude-fable-5
| arm | calls | mean input tok | mean output tok | $ / call | median wall | p25 / p75 |
|---|---|---|---|---|---|---|
| with_skill | 24 | 37,688.21 | 5,206.83 | $0.637224 | 43,164.5 ms | 21,832.5 / 95,723.5 |
| baseline | 31 | 55,531.29 | 3,998.16 | $0.755221 | 52,638 ms | 37,538 / 60,930 |
| incremental | n/a | -17,843.08 | +1,208.67 | -$0.117997 (-$117.997 per 1k calls) | -9473.5 ms | n/a |
Judge overhead, excluded: $3.253265 over 14 case rows. Dollars re-derive from the recorded tokens and the frozen rates: dollarsTraceable returns {"traceable":true,"mismatches":[]}. Source block: present and verified, and the recomputation matches it field for field across both arms and the incremental.
2 of 3 cells are cheaper with the skill than without it, at -$110.884 per thousand calls on code-review-and-quality@claude-fable-5 and -$117.997 per thousand calls on writing-plans@claude-fable-5. The mechanism is visible in the tokens rather than asserted: mean input falls by 8,557 and 17,843 tokens per call, which is the baseline's sprawl becoming unnecessary. The one cell that costs more, git-workflow-and-versioning@claude-sonnet-5, costs +$12.264 per thousand calls, and it is also the cell with no measured benefit at all: a lift of -0.002 and zero improved cases. Latency falls in 3 of 3.
A caution the tables carry and a reader should not skip: call counts are draw counts, and draw counts differ per arm. writing-plans@claude-fable-5's baseline drew 31 calls against with_skill's 24, because the baseline was unstable enough to keep drawing. Per-call means are therefore computed over differently sized draw sets in the two arms. That is correct, since a draw is a call and each measured draw contributes one usage row, but it means the cost delta and the score delta are averages over unequal n, and a reader who assumes matched n will misread both.
Two mechanical notes on the absolute columns. A CLI call carries a fixed harness preamble of roughly 25k input tokens that is identical in both arms and cancels in the delta, which is why the incremental row is the one to read. Cached input tokens are costed at the list input rate, which over-estimates; cached_tokens is recorded per call so a reader can recompute on other assumptions.
Disclosures
The published set is not fully clean, and one absorbed draw sits inside it. code-review-and-quality@claude-fable-5's severity-labeled-findings baseline arm drew 5 and measured 4: one generation call was lost to the 120 s timeout, and the band for that arm was computed from the 4 that survived. It produced no failed_timeout, no exclusion and no failed_case_count, because 4 measured draws is a legal draw set. That is silent absorption, and it is the same failure mode this report's instrument section describes, sitting in a cell this report publishes. It is one draw of 144 and it does not move the verdict, which reads improved at +0.332 on a wide baseline band either way. But "run 2 fixed it" is true of writing-plans@claude-fable-5 only. The other two cells were never re-run, and this one still carries the absorbed draw.
Run 1's receipts were sealed with null economics, and the reader was fixed after the sealing. Both run-1 receipts record call_count: 0 and null in every arm field. generationUsages read only the case-level usage, and a v0.5 receipt puts usage on the draws, so the arm figures came out empty while the draws carried complete usage the whole time. The judge block was never affected, because judge_usage stayed at case level, and that asymmetry is what located the defect. The economics tables above are recomputed from the same committed receipts through the corrected armEconomics, with no re-run and no edit to any receipt. writing-plans@claude-fable-5's run-2 receipt was sealed after the fix, and its recorded block reproduces field for field under recomputation, which is the control on the other two.
The per-case table above is computed by a path with no command behind it. driftproof diff compares two receipts. There is no shipped command that renders within-report, per-case, skill-against-baseline verdicts: a receipt carries one aggregate comparison.delta, and lib/verdict.js reads that aggregate through the effect floor with no band test at all. This page therefore calls the shipped rule functions directly, bandOf from lib/reuse.js and the band-and-floor rule from lib/diff.js, over each receipt's own case rows. No number is re-derived by hand, but the headline per-case table has no command a reader can run to get it, and no assertion over that command. It is filed as a gap in the tool, disclosed here rather than after someone notices.
Surface. The surface is claude-cli with local machine context, one box, held constant across all cells and all reports.
Scope. The instrument measures skill text in context on the append-system-prompt surface. Out of scope: routing via description:, progressive disclosure, and tool execution.
Two observations, offered as observations
Neither of the two sections below is a finding. Nothing in them was put to a band-separation test, and this report claims no verdict from them. They are stated because they are the most legible things in the clean run, and because leaving them out would be a choice about what a reader gets to see.
Time to settle
Draws are adaptive: a minimum of three per arm, and more until the across-draw spread settles. stopping_reason reads min_reached when three sufficed and stabilised when the arm needed more.
| cell | arms | settled at the 3-draw minimum | needed extra draws | total drawn | mean draws / arm |
|---|---|---|---|---|---|
code-review-and-quality@claude-fable-5 (run 1) | 14 | 13 | 1 | 44 | 3.14 |
git-workflow-and-versioning@claude-sonnet-5 (run 1) | 14 | 12 | 2 | 45 | 3.21 |
writing-plans@claude-fable-5 (run 2) | 14 | 8 | 6 | 55 | 3.93 |
The arms that needed extra draws are, with two exceptions, baseline arms. In writing-plans@claude-fable-5 five of the six unstable arms are baselines. The with-skill arm settles at the minimum almost everywhere; the baseline is what will not sit still.
Variance asymmetry
The ratio below is an arm's across-draw spread over its mean judge-sample spread, the variance_ratio the receipt records. Above 1× means generation-level noise exceeds judge-level noise: the axis this instrument sampled once, against the axis it sampled five times.
| cell | arms with a ratio | min | median | max | arms above 1× |
|---|---|---|---|---|---|
code-review-and-quality@claude-fable-5 (run 1) | 14 | 0.19× | 0.50× | 1.14× | 2 of 14 |
git-workflow-and-versioning@claude-sonnet-5 (run 1) | 14 | 0.18× | 0.63× | 10.26× | 5 of 14 |
writing-plans@claude-fable-5 (run 2) | 14 | 0.17× | 0.73× | 5.88× | 5 of 14 |
The distribution is not symmetric and the median misleads on its own: most arms are quiet at generation level and a few are extremely loud. git-workflow-and-versioning@claude-sonnet-5's semver-hidden-breaking-change runs 10.26× on baseline and 6.06× on with_skill, one case supplying most of that cell's dispersion. Medians here are median() from lib/value.js, which averages the two middle values on an even count.
Baseline spread against with-skill spread, writing-plans@claude-fable-5 run 2
| case | baseline sd | with_skill sd | ratio |
|---|---|---|---|
bite-sized-tdd-steps | 0.1549 | 0.0080 | 19.36× |
repair-placeholder-steps | 0.0943 | 0.0042 | 22.66× |
file-structure-by-responsibility | 0.0103 | 0.0061 | 1.68× |
interfaces-exact-signatures | 0.1444 | 0.0064 | 22.46× |
task-right-sizing-testable-deliverable | 0.1335 | 0.1687 | 0.79× |
full-small-plan-header-and-tasks | 0.0843 | 0.0133 | 6.33× |
self-review-spec-coverage | 0.0050 | 0.0200 | 0.25× |
5 of the seven cases have a baseline wider than the with-skill arm, three of them by more than nineteen times. code-review-and-quality@claude-fable-5 shows the same shape more weakly (5 of seven wider, maximum 8.00×) and git-workflow-and-versioning@claude-sonnet-5 weakly again (5 of seven, maximum 2.73×).
What the shape suggests, without this report having tested it, is that the skill's most visible effect on this corpus is on consistency rather than on the mean. That would also explain why so few cases separate: a wide baseline band is precisely what prevents a real mean difference from clearing an overlap test. This is a hypothesis the design of Report 007 cannot settle. A mean-difference test is not an instrument for a variance-reduction effect, and nothing here was pre-registered as a variance claim, so the shape is recorded and left open.
Why we looked, not what we confirmed. Two lines of prior work motivated these two tables and neither is evidence for them. Prompt rankings are unstable under ordinary evaluation variability, and ranking by a lower confidence bound rather than by a point estimate is a documented response: On the Stability of Prompt Ranking in Large Language Model Evaluation, arXiv:2606.24381. A wide baseline band is exactly the condition under which that selection rule matters, and Driftproof does not currently apply one. The variance-component framing above is CyclicJudge, arXiv:2603.01865, cited again here for the same decomposition.
Related work
| work | identifier | cited for |
|---|---|---|
| CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation 2026-03-02 |
arXiv:2603.01865 | The variance decomposition that partitions benchmark score variance into scenario, generation, judge and residual components. It is why a judge-sample band and an across-draw band are different statistics, and why an instrument that samples one of them is not measuring the other. |
| Second-Order Response Laws for LLM Judges: Debiased Estimation of Prompt Instability 2026-08-17 |
arXiv:2608.16253 | The estimator, rather than the decomposition. It formalises the split between sampling noise within a prompt and systematic difference across prompts as a second-order response law, and shows that the usual plug-in measure of prompt instability is biased upward at finite repeat budgets because it confounds the two; unbiased estimators follow from the difference between within- and across-prompt agreement. It cites CyclicJudge for the variance decomposition. Cited here as the nearest published treatment of the small-repeat-budget estimator problem this report's three-to-six draws per arm sit inside, and as a bias pointing the opposite way from the truncation bias measured above. This report applies no such correction. |
| On the Stability of Prompt Ranking in Large Language Model Evaluation 2026-06-23 |
arXiv:2606.24381 | A stability-aware selection rule based on a lower confidence bound, which accounts for performance and variance together instead of ranking on a point estimate. The selection rule a wide baseline band should force, and one this instrument does not yet apply. |
| WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution 2026-08-27 |
arXiv:2608.27454 | Prior art on skills as retrievable text carried in context, and on skills transferring across models and families. The boundary this report's scope sentence draws sits inside that picture: Driftproof measures the text in context and does not measure routing, progressive disclosure or tool execution. |
Each identifier above was resolved and its abstract read before publication, on the precedent Report 004's methodology section set. One further work carried through this report's authoring material had no identifier on file and none was found; it is dropped rather than cited, and no claim on this page rests on it.
Limits
- This report has not corrected Report 005. Two of the three archive comparisons were refused on baseline non-reproduction, so for those two cells nothing here is comparable to what that page published.
- The skill text changed between the two runs on all three cells. Every 005-to-007 pair differs in
skill.content_hash, so even the one comparison that passed its control is a revision pair and not a straight re-measurement of the same text. - The across-draw bands are small-budget estimates, uncorrected. Three to six draws per arm is a finite repeat budget, and the literature cited below shows estimators in this regime carry bias whose direction depends on what is being estimated. No debiasing is applied on this page; the bands are the plain sample spreads the receipts record.
- Cost means are over unequal n. Draw counts differ per arm, so a cell's two per-call means are averages over differently sized call sets.
- One absorbed draw remains in the published set, in
code-review-and-quality@claude-fable-5, disclosed above. - The judge remains single-model (
claude-haiku-4-5, five samples), and judge affinity is not measured here. - Three cells is three cells. Twenty-one cases on two substrates is not a corpus, and nothing on this page generalises past the skills named on it.
- Every verdict is verifiable from the receipts, which are linked below rather than described.
Amendments
v1.1 · 2026-09-02. Two sentences in this page's economics section read a mechanism out of a basis this page did not control for, and are qualified here. They are quoted rather than paraphrased. The first: “2 of 3 cells are cheaper with the skill than without it.” The second: “The mechanism is visible in the tokens rather than asserted: mean input falls by 8,557 and 17,843 tokens per call, which is the baseline's sprawl becoming unnecessary.” Both figures are correct as arithmetic over the receipts. Neither sentence is a safe reading of them.
What the basis actually was. The incremental column is one per-call mean minus another, and the two means are taken over call sets that differ in size and in composition: 21 draws with the skill against 22 without, allocated 3,3,3,3,3,3,3 against 4,3,3,3,3,3,3 on code-review-and-quality@claude-fable-5; 23 draws with the skill against 22 without, allocated 3,3,3,3,3,5,3 against 3,3,3,3,3,4,3 on git-workflow-and-versioning@claude-sonnet-5; 24 draws with the skill against 31 without, allocated 3,3,3,3,6,3,3 against 4,5,3,4,6,6,3 on writing-plans@claude-fable-5. Adaptive stopping chose those allocations per arm, so the two arms of a cell are averages over different case mixes. On top of that, every input token in both means is priced at the list input rate, cached tokens included, which is what the tables' own caption says and what computeEconomics does. On this surface cached context runs 89.8% to 95.1% of each call's input. The largest term in each mean is therefore harness context that neither arm chose and that the skill cannot move, priced at full rate, and the difference of two such means over unequal call sets is what the headline sentence read a mechanism out of. This page's Limits already stated the first half of that, “Cost means are over unequal n”; what it did not state is that the headline sentence rested on it.
The same draws, re-priced with cached input excluded. Nothing is re-run and no receipt is touched: cached_tokens is recorded per draw so that a reader can do exactly this, and the receipts say so in economics.notes.cache_pricing. Fresh input means input minus cached, at the same frozen rates, over the same measured draws.
| cell | as published (cached at list rate) | fresh input only | on the fresh basis | fresh, case-balanced |
|---|---|---|---|---|
code-review-and-quality@claude-fable-5 | -$0.110884 | -$0.028848 | sign survives | -$0.030390 |
git-workflow-and-versioning@claude-sonnet-5 | +$0.012264 | -$0.002212 | sign does not survive | -$0.002040 |
writing-plans@claude-fable-5 | -$0.117997 | +$0.040498 | sign does not survive | +$0.077634 |
2 of the 3 cells change sign, and they are not the ones the sentence would lead a reader to expect. On the fresh basis the cheaper cells are code-review-and-quality@claude-fable-5 and git-workflow-and-versioning@claude-sonnet-5, and the dearer one is writing-plans@claude-fable-5. The count of two survives; the identities do not. writing-plans@claude-fable-5, named in the sentence as one of the two cheaper cells at -$0.117997 per call, is +$0.040498 on fresh input, the dearest of the three. git-workflow-and-versioning@claude-sonnet-5, named as the one cell that costs more, is -$0.002212, cheaper. Only code-review-and-quality@claude-fable-5 keeps its sign, and its magnitude falls to 26.0% of the published figure.
The mechanism claim is retracted as written. “The baseline's sprawl becoming unnecessary” was offered as the explanation of an input-token fall, and the fall is not in the tokens the model generated. Of the 8,557-token fall on code-review-and-quality@claude-fable-5, 8,204 tokens, 95.9%, are cached; of the 17,843-token fall on writing-plans@claude-fable-5, 15,850 tokens, 88.8%, are cached. On fresh input the two falls are 354 and 1,994 tokens per call. A skill's text cannot make a harness preamble shorter, and an arm without the skill prepended consuming more input than the arm with it is a fact about cache warmth and call ordering, not about sprawl.
What survives, restated to its actual axis. On the fresh basis every cell's incremental is carried by output length, not input. code-review-and-quality@claude-fable-5 is cheaper because the skilled arm emits 506 fewer output tokens per call, worth -$0.025312 against -$0.003537 from fresh input. git-workflow-and-versioning@claude-sonnet-5 is cheaper on the same axis, -$0.002671 from 178 fewer output tokens. writing-plans@claude-fable-5 is dearer on that axis for the same reason in reverse: +1,209 output tokens per call, worth +$0.060434, which more than repays its fresh-input fall. So the defensible sentence is narrower than the published one and points somewhere else: where this skill changes cost at all, it changes it by changing how much the model writes, and the direction is not the same in every cell. Two cells shorter, one cell longer, on differences that are small in absolute terms and were never metered.
Neither basis is the true one, and that is the point. Cached input is billed below list rate on metered surfaces and at nothing at all here, where the metered spend is $0.00; pricing it at list is an over-estimate the receipts declare, and excluding it entirely is an under-estimate. The published tables use the first, this amendment adds the second, and the honest statement is that the sign of this metric on this surface is a property of which one a reader picks. That is a fact about the measurement and not about the skill, and it should have been on the page before a mechanism was asserted from either.
No figure above has been edited and no table is withdrawn. Every number in the economics tables re-derives from the committed receipts through computeEconomics at each receipt's own frozen snapshot, and the receipts are byte-unchanged. The two bases are computed from the same recorded fields, so this amendment adds a column beside the published one rather than replacing it. Filed by Report 008, which found the same two sentences waiting to be written about a second model pair and did not write them.
Amendments filed by this report
Report 005 to v1.2. The three cells' published lifts are named as single-draw, judge-spread figures, re-measured here with no cell-level separation. No figure on that page is edited.
Report 006 to v1.1. That page's writing-plans cell is corrected under pairwise exclusion, from +0.055 to +0.031, and its 6 against 6 aggregate is restated as the 5 against 5 it should have been. The cell's verdict does not change and its baseline-reproduction control does not move.
Both amendments are versioned records appended to the pages they concern. Constitution invariant 4: a published report is amended visibly, never edited silently.
Run record
Run 1: 2 published cells plus one evidence cell, 2026-08-31 07:59 UTC, committed at 7468e9e. Run 2: writing-plans@claude-fable-5 re-run under the corrected timeout, 2026-08-31 19:52 UTC, committed at 7050fd8. Published set: 144 draws, 143 measured, 1 absorbed. Judge claude-haiku-4-5 at five samples per draw. Surfaces: claude-fable-5 to claude-cli, claude-sonnet-5 to claude-cli. Metered spend $0.00 on vendor subscription CLIs; every dollar on this page is the if-it-had-been-metered equivalent at each receipt's own frozen snapshot. Verification level TESTED on all four receipts.
Receipts
All four receipts for this report, each one resolvable. Every number on this page re-derives from these files. They are the evidence, not a description of it: open any one and check it against a table above.
code-review-and-quality@claude-fable-5(run 1, published):receipts/report-007/code-review-and-quality-claude-fable-5-2026-08-31.jsongit-workflow-and-versioning@claude-sonnet-5(run 1, published):receipts/report-007/git-workflow-and-versioning-claude-sonnet-5-2026-08-31.jsonwriting-plans@claude-fable-5(run 2, published):receipts/report-007-rerun/writing-plans-claude-fable-5-2026-08-31.jsonwriting-plans@claude-fable-5(run 1, defect evidence, NOT published):receipts/report-007/writing-plans-claude-fable-5-2026-08-31.json
Directories: receipts/report-007/ and receipts/report-007-rerun/. Validate any of them with npx driftproof validate <file>. Run 1 landed at 7468e9e and run 2 at 7050fd8; no receipt has been modified since it was sealed.