Driftproof

Report 007: the baseline is the noisy arm

Instrument re-measurement report. Three cells Report 005 already published, re-run on the same suites and the same substrates with the generation sampled n times per arm instead of once. Nothing under the skill moves. The instrument does.

Three skills re-measured under generation sampling, with a corrected instrument. At the cell level no effect separates from noise: each of the three lifts is smaller than its own band (+0.055 ± 0.111 on code-review-and-quality@claude-fable-5, -0.002 ± 0.167 on git-workflow-and-versioning@claude-sonnet-5, +0.131 ± 0.157 on writing-plans@claude-fable-5). Per case across all 21: 3 improved · 0 regressed · 18 no effect · 0 not measured.

The cell band is suite dispersion: the sample standard deviation of the per-case means across a suite's seven cases, with the two arms' dispersions combined in quadrature. It is not the standard error of the mean, and the difference decides this page. A standard error shrinks as cases and draws are added, which makes a headline hypersensitive and turns trivial movement into an apparent result; aggregateBands() in lib/stats.js records that rationale beside the code that computes it.

Report 005 published lifts of +0.103, +0.116, +0.177 for these same three cells, each measured against a baseline arm generated once. All three are lower here. None of the three is a correction: two of the three archive comparisons were refused before any verdict was formed, and the skill text changed upstream between the two runs on every cell. What this report establishes is about the instrument, not about the skills.

What this report measures, and what it does not.

The set

cellrunreceiptsha256 of filereceipt_hash
code-review-and-quality@claude-fable-5run 1code-review-and-quality-claude-fable-5-2026-08-31.json11e1a9f9d199ad54…7d940158ee7f65b3…
git-workflow-and-versioning@claude-sonnet-5run 1git-workflow-and-versioning-claude-sonnet-5-2026-08-31.json69ded5000415053c…abec2aed0bc78b80…
writing-plans@claude-fable-5run 2writing-plans-claude-fable-5-2026-08-31.json280b15259b9faa91…f0dd7f573ee59593…

All three are schema_version 0.5, generation_sampled: true, verification level TESTED, surface claude-cli, provider anthropic. writing-plans@claude-fable-5 is run 2, and the next section is why. The other two cells were never re-run.

Result: no cell separates

cellwith_skill (mean ± band)baseline (mean ± band)liftlift bandseparates?
code-review-and-quality@claude-fable-50.885 ± 0.0170.830 ± 0.109+0.055± 0.111no
git-workflow-and-versioning@claude-sonnet-50.824 ± 0.1220.826 ± 0.115-0.002± 0.167no
writing-plans@claude-fable-50.833 ± 0.0740.701 ± 0.139+0.131± 0.157no

Every band above is suite dispersion. All three cells are paired 7 against 7 with an empty exclusion list, so run.failed_case_count is absent from all three published receipts: that field appears only when a case was excluded from the aggregate.

Per-case verdicts

A case counts only when the two bands do not overlap and the mean moves at least 0.05. Every band below carries its provenance: (generation) is an across-draw spread over n generation draws, which is what receipt spec v0.5 records, and it is a different statistic from the judge-sample spread earlier receipts carried.

code-review-and-quality@claude-fable-5 (run 1)

casebaseline bandwith_skill bandΔverdict
severity-labeled-findings0.585 ± 0.071 (generation)0.917 ± 0.026 (generation)+0.332improved
commit-message-imperative-body0.872 ± 0.003 (generation)0.868 ± 0.007 (generation)-0.004no effect (bands overlap)
security-axis-sql-injection0.885 ± 0.009 (generation)0.881 ± 0.003 (generation)-0.003no effect (bands overlap)
propose-structural-remedy0.885 ± 0.007 (generation)0.897 ± 0.011 (generation)+0.013no effect (bands overlap)
oversized-change-split0.869 ± 0.021 (generation)0.882 ± 0.006 (generation)+0.013no effect (bands overlap)
approve-when-improves-health0.877 ± 0.018 (generation)0.877 ± 0.002 (generation)+0.000no effect (bands overlap)
reject-clean-it-up-later0.836 ± 0.023 (generation)0.872 ± 0.014 (generation)+0.036no effect (bands overlap)

1 improved · 0 regressed · 6 no effect · 0 not measured.

git-workflow-and-versioning@claude-sonnet-5 (run 1)

casebaseline bandwith_skill bandΔverdict
commit-message-conventional-type0.861 ± 0.011 (generation)0.863 ± 0.004 (generation)+0.002no effect (bands overlap)
semver-clean-bump0.869 ± 0.004 (generation)0.867 ± 0.011 (generation)-0.003no effect (bands overlap)
split-into-atomic-commits0.870 ± 0.018 (generation)0.863 ± 0.012 (generation)-0.007no effect (bands overlap)
changelog-curated-by-impact0.867 ± 0.002 (generation)0.872 ± 0.007 (generation)+0.005no effect (bands overlap)
trunk-based-short-lived-branches0.858 ± 0.009 (generation)0.865 ± 0.004 (generation)+0.007no effect (bands overlap)
semver-hidden-breaking-change0.567 ± 0.355 (generation)0.548 ± 0.298 (generation)-0.019no effect (bands overlap)
release-cut-version-tag-changelog0.888 ± 0.012 (generation)0.890 ± 0.005 (generation)+0.002no effect (bands overlap)

0 improved · 0 regressed · 7 no effect · 0 not measured.

writing-plans@claude-fable-5 (run 2)

casebaseline bandwith_skill bandΔverdict
bite-sized-tdd-steps0.752 ± 0.155 (generation)0.872 ± 0.008 (generation)+0.120no effect (bands overlap)
repair-placeholder-steps0.740 ± 0.094 (generation)0.877 ± 0.004 (generation)+0.138improved
file-structure-by-responsibility0.821 ± 0.010 (generation)0.831 ± 0.006 (generation)+0.009no effect (bands overlap)
interfaces-exact-signatures0.725 ± 0.144 (generation)0.829 ± 0.006 (generation)+0.104no effect (bands overlap)
task-right-sizing-testable-deliverable0.637 ± 0.133 (generation)0.676 ± 0.169 (generation)+0.039no effect (bands overlap)
full-small-plan-header-and-tasks0.420 ± 0.084 (generation)0.903 ± 0.013 (generation)+0.482improved
self-review-spec-coverage0.815 ± 0.005 (generation)0.843 ± 0.020 (generation)+0.028no effect (below floor)

2 improved · 0 regressed · 5 no effect · 0 not measured.

3 improved · 0 regressed · 18 no effect · 0 not measured, over 21 published cases.

Two of the eighteen are the ones a reader will argue with. bite-sized-tdd-steps lifts +0.120 and interfaces-exact-signatures lifts +0.104, each more than twice the 0.05 floor, and each reads no effect. The reason sits on the same row: their baseline bands are ± 0.155 and ± 0.144. A lift twice the effect floor is not a finding when the baseline it is measured against will not sit still. Exactly one of the eighteen fails on the floor rather than on overlap: self-review-spec-coverage, at +0.028.

The instrument, and what it cost this study

A declared 300 s timeout had never executed. lib/provider.js declares a per-surface retry policy, and for a claude-cli surface it declares 300000 ms, with a written rationale about cold-start-dominated subprocesses. lib/run.js then set its own default as a numeric literal, 120000, and the provider layer documents that an explicit caller value wins. So every run this project has ever made used 120 s on every surface, and the CLI policy was dead text from the day it was written.

The dates are the uncomfortable part. The literals landed on 2026-07-27 with the runner skeleton. The policy that they shadowed was written on 2026-07-31, four days later, and read as authoritative from then on. Nothing failed, nothing logged, and no gate reached it until preparing this report did. It was fixed on 2026-08-31.

Run 1 lost 25 draws to it, 24 of them in writing-plans@claude-fable-5. Every loss carries the same recorded reason, and the string names the number: provider(claude-cli) timed out after 120000ms. Run 2 re-ran that one cell under the corrected timeout and measured 55 of 55, with nothing lost.

Both runs are published. Run 1's three receipts are committed at 7468e9e and run 2's at 7050fd8. Run 1's writing-plans@claude-fable-5 receipt is retained as defect evidence and is not part of the published set; the other two cells of run 1 are the published receipts for their skills, and were never re-run.

What the broken run measured, beside what the clean one measured

run 1 (truncated at 120 s)run 2 (300 s, clean)
draws71 drawn, 47 measured, 24 lost55 drawn, 55 measured, 0 lost
mean draws per arm5.073.93
arms with a variance ratio13 of 1414 of 14
maximum variance ratio1.55×5.88×
arms above 1×1 of 135 of 14
cell lift+0.116 ± 0.115+0.131 ± 0.157
pairing7 against 6, unpaired7 against 7, paired

The truncated run drew more, measured less, and looked calmer. 71 draws against 55, 5.07 per arm against 3.93, and a maximum variance ratio of 1.55× where the clean run reports 5.88×. 1 of 13 arms sat above 1× in run 1; 5 of 14 do in run 2.

The mechanism is the ordering. A 120 s ceiling removes the long generations first, and the long scattered generations are the ones carrying the across-draw spread. So the timeout did not merely lose data at random: it truncated the distribution from above, and biased the variance estimate downward. That is the direction that makes an instrument look more precise than it is. Silent truncation flatters stability.

It also cost this study its one apparent result. Run 1's writing-plans@claude-fable-5 lift of +0.116 ± 0.115 sat outside its own band by about a thousandth: the only separation anywhere in Report 007. Run 2's +0.131 ± 0.157 does not. Run 1 reported three improvements out of six measurable cases and could not measure the seventh at all; run 2, with every draw measured, reports two out of seven, and the case run 1 lost, full-small-plan-header-and-tasks, turns out to carry the largest case lift in the whole study at +0.482. The one result the study appeared to have was an artifact of the draws it lost.

One arm of run 1 lost every draw it took, so its variance ratio is null with variance_ratio_unavailable: "no_measured_draws" rather than a fabricated zero. That is why run 1 reports 13 ratios over 14 arms and run 2 reports 14.

Against Report 005

cellReport 005 lift and bandReport 007 lift and banddiffer verdict, 005 to 007
code-review-and-quality@claude-fable-5 +0.103 ± 0.217 (legacy) +0.055 ± 0.111 (generation) REFUSED
baseline did not reproduce: severity-labeled-findings
git-workflow-and-versioning@claude-sonnet-5 +0.116 ± 0.224 (legacy) -0.002 ± 0.167 (generation) REFUSED
baseline did not reproduce: commit-message-conventional-type, changelog-curated-by-impact
writing-plans@claude-fable-5 +0.177 ± 0.194 (legacy) +0.131 ± 0.157 (generation) WITHIN NOISE
precondition passed; 0 regressions, 0 improvements, 7 within noise

Band provenance, which the tool does not print on this row and which the report therefore supplies. Both columns are comparison.delta_uncertainty, the two arms' suite dispersions combined in quadrature, but they are built from different per-case statistics. The Report 005 column rests on legacy bands: a judge-sample spread over a single generation, which is what receipt spec v0.4 and earlier recorded. The Report 007 column rests on generation bands: an across-draw spread. They are different statistics, and the comparison is not like for like. It is also the only comparison the older receipt admits.

Five of six comparisons refuse

Six comparisons were run, one per cell against each of the two archives. Five of the six refused, and every refusal is a precondition failure rather than a computed result: four on baseline non-reproduction, one on a baseline with no measured draws. A refusal is an outcome this instrument publishes, not an error it recovers from. The reason strings below are the ones the tool emits.

Against Report 005, --mode revision

Against Report 006, --mode release

Why release mode and not revision. All three Report 006 pairs refuse in revision mode before any comparison is attempted, on a different precondition: the two receipts carry the SAME skill.content_hash: there is no revision between them to measure. Compare these two with the default release mode, or supply a pair that differs only in skill.content_hash. Release is therefore the only mode these three pairs admit, and it is the mode the list above reports.

Two of these refusals deserve a second sentence. On git-workflow-and-versioning@claude-sonnet-5 against Report 005, two cases moved beyond their bands, and the rendered reason names only the first: baselineReproduces returns bad[0] into the message while carrying the full list in cases, so the refusal a reader sees understates how many cases moved. It is true and incomplete, and it is filed. On writing-plans@claude-fable-5 against Report 006 the refusal is structural rather than empirical: Report 006's own receipt lost two arms to the same 120 s timeout, so a case on the Report 007 side has no band to compare against and the control returns baseline_unmeasured. Report 006's timeout losses are what stop Report 007 from comparing against it.

The one comparison that passed its precondition. writing-plans@claude-fable-5, Report 005 to Report 007, --mode revision: WITHIN NOISE — the revision moved no case beyond its confidence band; the pinned text and the current text measure the same.

with_skill mean moved +0.002 (0.831 ± 0.092 → 0.833 ± 0.074; band = suite dispersion). Per-case band-overlap verdicts: 0 regression(s), 0 improvement(s), 7 within noise.

So on the one cell where the control held, a lift of +0.177 and a lift of +0.131 are the same measurement as far as this instrument can tell. That is the honest reading of the mapping table above, and it is why none of the three lower lifts is presented as a correction.

What the mapping shows, and what it cannot. Every lift fell and every band narrowed. All three Report 005 lifts sat inside their own bands too, so nothing here reverses a separated finding. The direction is consistent: Report 005's single-draw baselines were 0.786, 0.746, 0.654, and Report 007's multi-draw baselines are 0.830, 0.826, 0.701, higher on two of the three. A baseline generated once was, on this evidence, a low estimate, and the lift measured against it was correspondingly high. That is a claim about the instrument. It is not a claim about the skills, and this report cannot make one: two of the three comparisons never got past their control, and the skill text differs between the two runs on all three cells.

Why the two band types are different statistics, and why an instrument that samples one of them is not measuring the other: a variance decomposition that partitions benchmark score variance into scenario, generation, judge and residual components is given by CyclicJudge, arXiv:2603.01865. Driftproof sampled the judge five times and the generation once until receipt spec v0.5, which is to say it had been putting its error bars on one component and reading them as the whole. Second-Order Response Laws for LLM Judges, arXiv:2608.16253, which cites CyclicJudge for that decomposition, goes on to the estimator: it shows that a plug-in estimate of prompt instability is biased upward at finite repeat budgets, because it confounds within-prompt noise with between-prompt variation, and derives an unbiased estimator from the difference between within- and across-prompt agreement. Report 007 does not apply that correction, and says so in its limits.

Economics

No figure in this section comes from a projection. estimateRunCostUSD, the projection path, was not called. Each table is computed at build time from the draws the receipt records, priced at the snapshot the receipt froze at run time.

code-review-and-quality@claude-fable-5

Basis. Subscription claude-cli surface; metered spend $0.00. Every dollar is a metered-equivalent derived at build time from draws[].usage at this receipt's own frozen snapshot (input $10 / Mtok, output $50 / Mtok, frozen 2026-08-31T07:59:57.284Z). Read the incremental row, not the absolute ones. Judge cost is measurement overhead this project imposes and is excluded from every field.
armcallsmean input tokmean output tok$ / callmedian wallp25 / p75
with_skill2124,184.951,251.81$0.30444021,635 ms16,705 / 23,055.5
baseline2232,742.181,758.05$0.41532423,689 ms19,855 / 40,100
incrementaln/a-8,557.23-506.24-$0.110884 (-$110.884 per 1k calls)-2054.0 msn/a

Judge overhead, excluded: $2.601445 over 14 case rows. Dollars re-derive from the recorded tokens and the frozen rates: dollarsTraceable returns {"traceable":true,"mismatches":[]}. Source block: the receipt's recorded economics block is null; this table is a recomputation from the same committed receipt.

git-workflow-and-versioning@claude-sonnet-5

Basis. Subscription claude-cli surface; metered spend $0.00. Every dollar is a metered-equivalent derived at build time from draws[].usage at this receipt's own frozen snapshot (input $3 / Mtok, output $15 / Mtok, frozen 2026-08-31T05:42:11.959Z). Read the incremental row, not the absolute ones. Judge cost is measurement overhead this project imposes and is excluded from every field.
armcallsmean input tokmean output tok$ / callmedian wallp25 / p75
with_skill2333,082.39300.22$0.1037506,250 ms4,485 / 7,029
baseline2228,104.09478.27$0.0914867,510.5 ms5,051 / 10,389
incrementaln/a+4,978.30-178.05+$0.012264 (+$12.264 per 1k calls)-1260.5 msn/a

Judge overhead, excluded: $2.313420 over 14 case rows. Dollars re-derive from the recorded tokens and the frozen rates: dollarsTraceable returns {"traceable":true,"mismatches":[]}. Source block: the receipt's recorded economics block is null; this table is a recomputation from the same committed receipt.

writing-plans@claude-fable-5

Basis. Subscription claude-cli surface; metered spend $0.00. Every dollar is a metered-equivalent derived at build time from draws[].usage at this receipt's own frozen snapshot (input $10 / Mtok, output $50 / Mtok, frozen 2026-08-31T19:52:29.216Z). Read the incremental row, not the absolute ones. Judge cost is measurement overhead this project imposes and is excluded from every field.
armcallsmean input tokmean output tok$ / callmedian wallp25 / p75
with_skill2437,688.215,206.83$0.63722443,164.5 ms21,832.5 / 95,723.5
baseline3155,531.293,998.16$0.75522152,638 ms37,538 / 60,930
incrementaln/a-17,843.08+1,208.67-$0.117997 (-$117.997 per 1k calls)-9473.5 msn/a

Judge overhead, excluded: $3.253265 over 14 case rows. Dollars re-derive from the recorded tokens and the frozen rates: dollarsTraceable returns {"traceable":true,"mismatches":[]}. Source block: present and verified, and the recomputation matches it field for field across both arms and the incremental.

2 of 3 cells are cheaper with the skill than without it, at -$110.884 per thousand calls on code-review-and-quality@claude-fable-5 and -$117.997 per thousand calls on writing-plans@claude-fable-5. The mechanism is visible in the tokens rather than asserted: mean input falls by 8,557 and 17,843 tokens per call, which is the baseline's sprawl becoming unnecessary. The one cell that costs more, git-workflow-and-versioning@claude-sonnet-5, costs +$12.264 per thousand calls, and it is also the cell with no measured benefit at all: a lift of -0.002 and zero improved cases. Latency falls in 3 of 3.

A caution the tables carry and a reader should not skip: call counts are draw counts, and draw counts differ per arm. writing-plans@claude-fable-5's baseline drew 31 calls against with_skill's 24, because the baseline was unstable enough to keep drawing. Per-call means are therefore computed over differently sized draw sets in the two arms. That is correct, since a draw is a call and each measured draw contributes one usage row, but it means the cost delta and the score delta are averages over unequal n, and a reader who assumes matched n will misread both.

Two mechanical notes on the absolute columns. A CLI call carries a fixed harness preamble of roughly 25k input tokens that is identical in both arms and cancels in the delta, which is why the incremental row is the one to read. Cached input tokens are costed at the list input rate, which over-estimates; cached_tokens is recorded per call so a reader can recompute on other assumptions.

Disclosures

The published set is not fully clean, and one absorbed draw sits inside it. code-review-and-quality@claude-fable-5's severity-labeled-findings baseline arm drew 5 and measured 4: one generation call was lost to the 120 s timeout, and the band for that arm was computed from the 4 that survived. It produced no failed_timeout, no exclusion and no failed_case_count, because 4 measured draws is a legal draw set. That is silent absorption, and it is the same failure mode this report's instrument section describes, sitting in a cell this report publishes. It is one draw of 144 and it does not move the verdict, which reads improved at +0.332 on a wide baseline band either way. But "run 2 fixed it" is true of writing-plans@claude-fable-5 only. The other two cells were never re-run, and this one still carries the absorbed draw.

Run 1's receipts were sealed with null economics, and the reader was fixed after the sealing. Both run-1 receipts record call_count: 0 and null in every arm field. generationUsages read only the case-level usage, and a v0.5 receipt puts usage on the draws, so the arm figures came out empty while the draws carried complete usage the whole time. The judge block was never affected, because judge_usage stayed at case level, and that asymmetry is what located the defect. The economics tables above are recomputed from the same committed receipts through the corrected armEconomics, with no re-run and no edit to any receipt. writing-plans@claude-fable-5's run-2 receipt was sealed after the fix, and its recorded block reproduces field for field under recomputation, which is the control on the other two.

The per-case table above is computed by a path with no command behind it. driftproof diff compares two receipts. There is no shipped command that renders within-report, per-case, skill-against-baseline verdicts: a receipt carries one aggregate comparison.delta, and lib/verdict.js reads that aggregate through the effect floor with no band test at all. This page therefore calls the shipped rule functions directly, bandOf from lib/reuse.js and the band-and-floor rule from lib/diff.js, over each receipt's own case rows. No number is re-derived by hand, but the headline per-case table has no command a reader can run to get it, and no assertion over that command. It is filed as a gap in the tool, disclosed here rather than after someone notices.

Surface. The surface is claude-cli with local machine context, one box, held constant across all cells and all reports.

Scope. The instrument measures skill text in context on the append-system-prompt surface. Out of scope: routing via description:, progressive disclosure, and tool execution.

Two observations, offered as observations

Neither of the two sections below is a finding. Nothing in them was put to a band-separation test, and this report claims no verdict from them. They are stated because they are the most legible things in the clean run, and because leaving them out would be a choice about what a reader gets to see.

Time to settle

Draws are adaptive: a minimum of three per arm, and more until the across-draw spread settles. stopping_reason reads min_reached when three sufficed and stabilised when the arm needed more.

cellarmssettled at the 3-draw minimumneeded extra drawstotal drawnmean draws / arm
code-review-and-quality@claude-fable-5 (run 1)14131443.14
git-workflow-and-versioning@claude-sonnet-5 (run 1)14122453.21
writing-plans@claude-fable-5 (run 2)1486553.93

The arms that needed extra draws are, with two exceptions, baseline arms. In writing-plans@claude-fable-5 five of the six unstable arms are baselines. The with-skill arm settles at the minimum almost everywhere; the baseline is what will not sit still.

Variance asymmetry

The ratio below is an arm's across-draw spread over its mean judge-sample spread, the variance_ratio the receipt records. Above 1× means generation-level noise exceeds judge-level noise: the axis this instrument sampled once, against the axis it sampled five times.

cellarms with a ratiominmedianmaxarms above 1×
code-review-and-quality@claude-fable-5 (run 1)140.19×0.50×1.14×2 of 14
git-workflow-and-versioning@claude-sonnet-5 (run 1)140.18×0.63×10.26×5 of 14
writing-plans@claude-fable-5 (run 2)140.17×0.73×5.88×5 of 14

The distribution is not symmetric and the median misleads on its own: most arms are quiet at generation level and a few are extremely loud. git-workflow-and-versioning@claude-sonnet-5's semver-hidden-breaking-change runs 10.26× on baseline and 6.06× on with_skill, one case supplying most of that cell's dispersion. Medians here are median() from lib/value.js, which averages the two middle values on an even count.

Baseline spread against with-skill spread, writing-plans@claude-fable-5 run 2

casebaseline sdwith_skill sdratio
bite-sized-tdd-steps0.15490.008019.36×
repair-placeholder-steps0.09430.004222.66×
file-structure-by-responsibility0.01030.00611.68×
interfaces-exact-signatures0.14440.006422.46×
task-right-sizing-testable-deliverable0.13350.16870.79×
full-small-plan-header-and-tasks0.08430.01336.33×
self-review-spec-coverage0.00500.02000.25×

5 of the seven cases have a baseline wider than the with-skill arm, three of them by more than nineteen times. code-review-and-quality@claude-fable-5 shows the same shape more weakly (5 of seven wider, maximum 8.00×) and git-workflow-and-versioning@claude-sonnet-5 weakly again (5 of seven, maximum 2.73×).

What the shape suggests, without this report having tested it, is that the skill's most visible effect on this corpus is on consistency rather than on the mean. That would also explain why so few cases separate: a wide baseline band is precisely what prevents a real mean difference from clearing an overlap test. This is a hypothesis the design of Report 007 cannot settle. A mean-difference test is not an instrument for a variance-reduction effect, and nothing here was pre-registered as a variance claim, so the shape is recorded and left open.

Why we looked, not what we confirmed. Two lines of prior work motivated these two tables and neither is evidence for them. Prompt rankings are unstable under ordinary evaluation variability, and ranking by a lower confidence bound rather than by a point estimate is a documented response: On the Stability of Prompt Ranking in Large Language Model Evaluation, arXiv:2606.24381. A wide baseline band is exactly the condition under which that selection rule matters, and Driftproof does not currently apply one. The variance-component framing above is CyclicJudge, arXiv:2603.01865, cited again here for the same decomposition.

Related work

workidentifiercited for
CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation
2026-03-02
arXiv:2603.01865 The variance decomposition that partitions benchmark score variance into scenario, generation, judge and residual components. It is why a judge-sample band and an across-draw band are different statistics, and why an instrument that samples one of them is not measuring the other.
Second-Order Response Laws for LLM Judges: Debiased Estimation of Prompt Instability
2026-08-17
arXiv:2608.16253 The estimator, rather than the decomposition. It formalises the split between sampling noise within a prompt and systematic difference across prompts as a second-order response law, and shows that the usual plug-in measure of prompt instability is biased upward at finite repeat budgets because it confounds the two; unbiased estimators follow from the difference between within- and across-prompt agreement. It cites CyclicJudge for the variance decomposition. Cited here as the nearest published treatment of the small-repeat-budget estimator problem this report's three-to-six draws per arm sit inside, and as a bias pointing the opposite way from the truncation bias measured above. This report applies no such correction.
On the Stability of Prompt Ranking in Large Language Model Evaluation
2026-06-23
arXiv:2606.24381 A stability-aware selection rule based on a lower confidence bound, which accounts for performance and variance together instead of ranking on a point estimate. The selection rule a wide baseline band should force, and one this instrument does not yet apply.
WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution
2026-08-27
arXiv:2608.27454 Prior art on skills as retrievable text carried in context, and on skills transferring across models and families. The boundary this report's scope sentence draws sits inside that picture: Driftproof measures the text in context and does not measure routing, progressive disclosure or tool execution.

Each identifier above was resolved and its abstract read before publication, on the precedent Report 004's methodology section set. One further work carried through this report's authoring material had no identifier on file and none was found; it is dropped rather than cited, and no claim on this page rests on it.

Limits

Amendments

v1.1 · 2026-09-02. Two sentences in this page's economics section read a mechanism out of a basis this page did not control for, and are qualified here. They are quoted rather than paraphrased. The first: “2 of 3 cells are cheaper with the skill than without it.” The second: “The mechanism is visible in the tokens rather than asserted: mean input falls by 8,557 and 17,843 tokens per call, which is the baseline's sprawl becoming unnecessary.” Both figures are correct as arithmetic over the receipts. Neither sentence is a safe reading of them.

What the basis actually was. The incremental column is one per-call mean minus another, and the two means are taken over call sets that differ in size and in composition: 21 draws with the skill against 22 without, allocated 3,3,3,3,3,3,3 against 4,3,3,3,3,3,3 on code-review-and-quality@claude-fable-5; 23 draws with the skill against 22 without, allocated 3,3,3,3,3,5,3 against 3,3,3,3,3,4,3 on git-workflow-and-versioning@claude-sonnet-5; 24 draws with the skill against 31 without, allocated 3,3,3,3,6,3,3 against 4,5,3,4,6,6,3 on writing-plans@claude-fable-5. Adaptive stopping chose those allocations per arm, so the two arms of a cell are averages over different case mixes. On top of that, every input token in both means is priced at the list input rate, cached tokens included, which is what the tables' own caption says and what computeEconomics does. On this surface cached context runs 89.8% to 95.1% of each call's input. The largest term in each mean is therefore harness context that neither arm chose and that the skill cannot move, priced at full rate, and the difference of two such means over unequal call sets is what the headline sentence read a mechanism out of. This page's Limits already stated the first half of that, “Cost means are over unequal n”; what it did not state is that the headline sentence rested on it.

The same draws, re-priced with cached input excluded. Nothing is re-run and no receipt is touched: cached_tokens is recorded per draw so that a reader can do exactly this, and the receipts say so in economics.notes.cache_pricing. Fresh input means input minus cached, at the same frozen rates, over the same measured draws.

Per-call incremental, with skill minus without, on two bases over one set of recorded draws. The last column weights every case equally in both arms instead of weighting it by how many draws adaptive stopping gave it, and is the control on whether a sign belongs to the draw allocation.
cellas published (cached at list rate)fresh input onlyon the fresh basisfresh, case-balanced
code-review-and-quality@claude-fable-5-$0.110884-$0.028848sign survives-$0.030390
git-workflow-and-versioning@claude-sonnet-5+$0.012264-$0.002212sign does not survive-$0.002040
writing-plans@claude-fable-5-$0.117997+$0.040498sign does not survive+$0.077634

2 of the 3 cells change sign, and they are not the ones the sentence would lead a reader to expect. On the fresh basis the cheaper cells are code-review-and-quality@claude-fable-5 and git-workflow-and-versioning@claude-sonnet-5, and the dearer one is writing-plans@claude-fable-5. The count of two survives; the identities do not. writing-plans@claude-fable-5, named in the sentence as one of the two cheaper cells at -$0.117997 per call, is +$0.040498 on fresh input, the dearest of the three. git-workflow-and-versioning@claude-sonnet-5, named as the one cell that costs more, is -$0.002212, cheaper. Only code-review-and-quality@claude-fable-5 keeps its sign, and its magnitude falls to 26.0% of the published figure.

The mechanism claim is retracted as written. “The baseline's sprawl becoming unnecessary” was offered as the explanation of an input-token fall, and the fall is not in the tokens the model generated. Of the 8,557-token fall on code-review-and-quality@claude-fable-5, 8,204 tokens, 95.9%, are cached; of the 17,843-token fall on writing-plans@claude-fable-5, 15,850 tokens, 88.8%, are cached. On fresh input the two falls are 354 and 1,994 tokens per call. A skill's text cannot make a harness preamble shorter, and an arm without the skill prepended consuming more input than the arm with it is a fact about cache warmth and call ordering, not about sprawl.

What survives, restated to its actual axis. On the fresh basis every cell's incremental is carried by output length, not input. code-review-and-quality@claude-fable-5 is cheaper because the skilled arm emits 506 fewer output tokens per call, worth -$0.025312 against -$0.003537 from fresh input. git-workflow-and-versioning@claude-sonnet-5 is cheaper on the same axis, -$0.002671 from 178 fewer output tokens. writing-plans@claude-fable-5 is dearer on that axis for the same reason in reverse: +1,209 output tokens per call, worth +$0.060434, which more than repays its fresh-input fall. So the defensible sentence is narrower than the published one and points somewhere else: where this skill changes cost at all, it changes it by changing how much the model writes, and the direction is not the same in every cell. Two cells shorter, one cell longer, on differences that are small in absolute terms and were never metered.

Neither basis is the true one, and that is the point. Cached input is billed below list rate on metered surfaces and at nothing at all here, where the metered spend is $0.00; pricing it at list is an over-estimate the receipts declare, and excluding it entirely is an under-estimate. The published tables use the first, this amendment adds the second, and the honest statement is that the sign of this metric on this surface is a property of which one a reader picks. That is a fact about the measurement and not about the skill, and it should have been on the page before a mechanism was asserted from either.

No figure above has been edited and no table is withdrawn. Every number in the economics tables re-derives from the committed receipts through computeEconomics at each receipt's own frozen snapshot, and the receipts are byte-unchanged. The two bases are computed from the same recorded fields, so this amendment adds a column beside the published one rather than replacing it. Filed by Report 008, which found the same two sentences waiting to be written about a second model pair and did not write them.

Amendments filed by this report

Report 005 to v1.2. The three cells' published lifts are named as single-draw, judge-spread figures, re-measured here with no cell-level separation. No figure on that page is edited.

Report 006 to v1.1. That page's writing-plans cell is corrected under pairwise exclusion, from +0.055 to +0.031, and its 6 against 6 aggregate is restated as the 5 against 5 it should have been. The cell's verdict does not change and its baseline-reproduction control does not move.

Both amendments are versioned records appended to the pages they concern. Constitution invariant 4: a published report is amended visibly, never edited silently.

Run record

Run 1: 2 published cells plus one evidence cell, 2026-08-31 07:59 UTC, committed at 7468e9e. Run 2: writing-plans@claude-fable-5 re-run under the corrected timeout, 2026-08-31 19:52 UTC, committed at 7050fd8. Published set: 144 draws, 143 measured, 1 absorbed. Judge claude-haiku-4-5 at five samples per draw. Surfaces: claude-fable-5 to claude-cli, claude-sonnet-5 to claude-cli. Metered spend $0.00 on vendor subscription CLIs; every dollar on this page is the if-it-had-been-metered equivalent at each receipt's own frozen snapshot. Verification level TESTED on all four receipts.

Receipts

All four receipts for this report, each one resolvable. Every number on this page re-derives from these files. They are the evidence, not a description of it: open any one and check it against a table above.

Directories: receipts/report-007/ and receipts/report-007-rerun/. Validate any of them with npx driftproof validate <file>. Run 1 landed at 7468e9e and run 2 at 7050fd8; no receipt has been modified since it was sealed.