Driftproof

Report 008: the skill stabilises the floor, not the ceiling

Release drift report. Two cells re-measured on claude-fable-5-1 against claude-fable-5, with the skill text and the suite asserted byte-identical before launch. The first release pair in this project where both sides are generation-sampled receipts, which is what makes the delta attributable to the model.

Both cells came back WITHIN NOISE. Not one of the fourteen cases moved beyond its confidence band when the same skill text, over the same suite, on the same surface, was measured on claude-fable-5-1 instead of claude-fable-5. Per-case regressions, improvements and within-noise counts are 0/0/7 and 0/0/7. The skill holds up on the new model.

code-review-and-quality@claude-fable-5-1: WITHIN NOISE — no case moved beyond its confidence band; the skill holds up.
with_skill mean moved -0.020 (0.885 ± 0.017 → 0.865 ± 0.036; band = suite dispersion). Per-case band-overlap verdicts: 0 regression(s), 0 improvement(s), 7 within noise.

writing-plans@claude-fable-5-1: WITHIN NOISE — no case moved beyond its confidence band; the skill holds up.
with_skill mean moved +0.011 (0.833 ± 0.074 → 0.844 ± 0.048; band = suite dispersion). Per-case band-overlap verdicts: 0 regression(s), 0 improvement(s), 7 within noise (1 of them band-separated but below the 0.05 effect floor).

Both lines above are driftproof diff's own output, quoted out of the markdown it emits rather than restated here. Neither invocation passed a flag that relaxes epistemics, both exited 0, and neither cell produced a refusal: the cross-version comparability refusal did not fire because the receipts are the same schema version, and the below-TESTED refusal did not fire because all four receipts are TESTED.

Verdict

HOLDS ACROSS THE RELEASE, ON A PAIR WHERE THE MODEL IS THE ONLY THING THAT MOVED.

What that does and does not say. It says the two skills' measured behaviour on claude-fable-5-1 is indistinguishable, by this instrument, from their behaviour on claude-fable-5, and that the attribution is clean: skill content_hash and suite_hash were asserted identical against the Report 007 baselines before the first call, and surface, judge and judge samples per case are held on both sides of both pairs. It does not say the skills got better, it does not say they got worse, and it does not say nothing changed underneath. A within-noise verdict is a statement about what this instrument can resolve at this sample size, which is 92 generation draws across 28 arms.

Where the verdict is closest to its own edge. Exactly one case in the run is held inside the verdict by the effect floor rather than by band overlap: writing-plans@claude-fable-5-1's file-structure-by-responsibility, whose bands ARE separated (0.831 ± 0.006 (generation) against 0.853 ± 0.013 (generation)) and whose move of +0.023 is below the 0.05 floor. It is named here rather than left in a table, because it is the one row where a reader who set a lower floor would read a different verdict. Every other case in the run fails the separation test outright, so the verdict does not rest on the floor.

The finding: Report 007's thesis, recurring on a second model pair

Report 007 found that across three cells the arm that would not sit still was the baseline, and that what the skill did was stabilise the floor rather than raise the ceiling. That report moved the instrument and held the substrate. This one moves the substrate and holds the instrument, and the same shape comes back.

The headline case of code-review-and-quality@claude-fable-5-1 went pass to borderline, and the reason is not the skill. severity-labeled-findings carries 96.2% of that cell's entire lift: its +0.488 against a sum of +0.507 across all 7 cases, with the remaining six summing to +0.019. Take it out and the cell's lift is +0.003, indistinguishable from zero and far below the 0.05 floor. On the release axis that one case moved like this:

armon claude-fable-5on claude-fable-5-1Δ
baseline0.585 ± 0.071 (generation)0.308 ± 0.019 (generation)-0.277
with_skill0.917 ± 0.026 (generation)0.796 ± 0.147 (generation)-0.121

The baseline arm fell 0.585 to 0.308. The with_skill arm held, moving -0.121 and staying inside its own band. The cell's lift on this case therefore grew, +0.332 to +0.488, and almost all of that growth is the unskilled arm getting worse rather than the skilled arm getting better. A report that quoted the lift without this sentence would credit the skill with something the model did to the arm that does not have it.

The stability moved the other way, and that is the cost of the same fact. The with_skill band on that case widened from ± 0.026 to ± 0.147, a factor of 5.7, which is what took its outcome label from pass to borderline and what makes the release-axis verdict on it read within noise despite a move of -0.121. The receipt says where the width came from: the arm drew a fourth generation with stopping_reason: "stabilised", and one of the four scored far below the others. A one-case pillar that has become unstable is a fragility this page discloses whatever the aggregate verdict says.

The same pattern carries writing-plans@claude-fable-5-1, with two pillars instead of one. The 2 cases whose baseline arm outright fails carry 88.6% of that cell's lift: full-small-plan-header-and-tasks at +0.480 (baseline 0.415) and task-right-sizing-testable-deliverable at +0.302 (baseline 0.449). The other 5 cases sum to +0.101. In both cells, on both models, the lift is not spread across the suite: it is a handful of cases where the unskilled arm fails and the skilled arm does not.

Which arm actually moved between the two models. The differ compares the with_skill arm only, so baseline drift is invisible in its output. Computed here over the same receipts, the largest movements in the run are baseline movements:

The five cases whose baseline arm moved most between the two models, with the same case's with_skill movement beside it. Bands carry their source.
cellcasebaseline, claude-fable-5 to claude-fable-5-1Δ baselineΔ with_skill
code-review-and-quality@claude-fable-5-1severity-labeled-findings0.585 ± 0.071 (generation) to 0.308 ± 0.019 (generation)-0.277-0.121
writing-plans@claude-fable-5-1task-right-sizing-testable-deliverable0.637 ± 0.133 (generation) to 0.449 ± 0.092 (generation)-0.187+0.076
writing-plans@claude-fable-5-1bite-sized-tdd-steps0.752 ± 0.155 (generation) to 0.860 ± 0.021 (generation)+0.108+0.017
writing-plans@claude-fable-5-1interfaces-exact-signatures0.725 ± 0.144 (generation) to 0.796 ± 0.054 (generation)+0.072-0.006
writing-plans@claude-fable-5-1repair-placeholder-steps0.740 ± 0.094 (generation) to 0.790 ± 0.033 (generation)+0.050-0.029

Across all 14 cases the widest baseline movement is 0.277 and the widest with_skill movement is 0.121. The floor moved further than the ceiling did, on a release the skill text did not know about. That is Report 007's finding, reproduced on a substrate change rather than an instrument change, and it is the reason this report presents the aggregate lift with its per-case table attached rather than on its own.

One honest caution about the sentence above. The baseline arm is where a model's own default behaviour shows through, so it is the arm most exposed to a release, and finding the larger movement there is not by itself surprising. What the two reports together support is narrower than a law: on five cells across two model pairs, the arm with the skill text moved less than the arm without it, and the skill's measured value came from the cases where the unskilled arm failed. Five cells is five cells.

The measurement, per cell

code-review-and-quality@claude-fable-5-1

Release axis, with_skill on claude-fable-5 against with_skill on claude-fable-5-1. This is the comparison the verdict rests on. Skill content_hash 13d360d7f786 and suite_hash 4729eb8fc298 are identical on both sides.

Every band below is an across-draw spread over the generation draws the receipt records, and carries (generation) as its source. No (legacy) band appears anywhere in this report: both sides of both pairs are v0.5 generation-sampled receipts, which is what makes this the first release pair the project can compare like for like.
case2026-08-31 (claude-fable-5)2026-09-01 (claude-fable-5-1)Δverdict
severity-labeled-findings0.917 ± 0.026 (generation)0.796 ± 0.147 (generation)-0.121within noise
approve-when-improves-health0.877 ± 0.002 (generation)0.847 ± 0.034 (generation)-0.031within noise
reject-clean-it-up-later0.872 ± 0.014 (generation)0.861 ± 0.022 (generation)-0.011within noise
commit-message-imperative-body0.868 ± 0.007 (generation)0.869 ± 0.006 (generation)+0.001within noise
propose-structural-remedy0.897 ± 0.011 (generation)0.903 ± 0.010 (generation)+0.005within noise
oversized-change-split0.882 ± 0.006 (generation)0.890 ± 0.012 (generation)+0.008within noise
security-axis-sql-injection0.881 ± 0.003 (generation)0.891 ± 0.008 (generation)+0.010within noise

Cell aggregates: with_skill 0.865 ± 0.036, baseline 0.793, lift +0.072 ± 0.217. An aggregate band carries no source label because it is not a band over draws: it is the dispersion of the seven case means across the suite, a different statistic from the per-case bands above, and the two are not interchangeable.

The lift inside code-review-and-quality@claude-fable-5-1, per case

Within-report, with_skill against baseline on claude-fable-5-1, through the same band-and-floor rule. Share is the case's Δ over the sum of all 7 Δs; the sum over 7 reproduces the receipt's own comparison.delta.
casewith_skillbaselineΔshare of liftverdict
severity-labeled-findings0.796 ± 0.147 (generation)0.308 ± 0.019 (generation)+0.48896.2%improved
oversized-change-split0.890 ± 0.012 (generation)0.864 ± 0.037 (generation)+0.0265.1%no effect (bands overlap)
security-axis-sql-injection0.891 ± 0.008 (generation)0.878 ± 0.007 (generation)+0.0132.6%no effect (bands overlap)
propose-structural-remedy0.903 ± 0.010 (generation)0.900 ± 0.014 (generation)+0.0030.5%no effect (bands overlap)
commit-message-imperative-body0.869 ± 0.006 (generation)0.872 ± 0.005 (generation)-0.003-0.7%no effect (bands overlap)
reject-clean-it-up-later0.861 ± 0.022 (generation)0.868 ± 0.007 (generation)-0.007-1.3%no effect (bands overlap)
approve-when-improves-health0.847 ± 0.034 (generation)0.859 ± 0.017 (generation)-0.013-2.5%no effect (bands overlap)

The aggregate lift of +0.072 ± 0.217 is not a band-verified improvement claim, and this page does not make one. The lift clears the 0.05 floor and fails the separation test, because the baseline aggregate band (± 0.214) overlaps the with_skill one. That width is produced by the 1 failing baseline case in the arm the skill is meant to fix, not by instability in the with_skill arm (± 0.036). Reading +0.072 ± 0.217 as "the lift might be negative" is the reading that width invites and the per-case table above contradicts.

writing-plans@claude-fable-5-1

Release axis, with_skill on claude-fable-5 against with_skill on claude-fable-5-1. This is the comparison the verdict rests on. Skill content_hash 5325ab3d55d5 and suite_hash 37cb3c341572 are identical on both sides.

Every band below is an across-draw spread over the generation draws the receipt records, and carries (generation) as its source. No (legacy) band appears anywhere in this report: both sides of both pairs are v0.5 generation-sampled receipts, which is what makes this the first release pair the project can compare like for like.
case2026-08-31 (claude-fable-5)2026-09-02 (claude-fable-5-1)Δverdict
repair-placeholder-steps0.877 ± 0.004 (generation)0.848 ± 0.031 (generation)-0.029within noise
full-small-plan-header-and-tasks0.903 ± 0.013 (generation)0.895 ± 0.019 (generation)-0.008within noise
interfaces-exact-signatures0.829 ± 0.006 (generation)0.823 ± 0.033 (generation)-0.006within noise
self-review-spec-coverage0.843 ± 0.020 (generation)0.850 ± 0.011 (generation)+0.007within noise
bite-sized-tdd-steps0.872 ± 0.008 (generation)0.889 ± 0.013 (generation)+0.017within noise
file-structure-by-responsibility0.831 ± 0.006 (generation)0.853 ± 0.013 (generation)+0.023within noise (below effect floor)
task-right-sizing-testable-deliverable0.676 ± 0.169 (generation)0.752 ± 0.152 (generation)+0.076within noise

Cell aggregates: with_skill 0.844 ± 0.048, baseline 0.718, lift +0.126 ± 0.203. An aggregate band carries no source label because it is not a band over draws: it is the dispersion of the seven case means across the suite, a different statistic from the per-case bands above, and the two are not interchangeable.

The lift inside writing-plans@claude-fable-5-1, per case

Within-report, with_skill against baseline on claude-fable-5-1, through the same band-and-floor rule. Share is the case's Δ over the sum of all 7 Δs; the sum over 7 reproduces the receipt's own comparison.delta.
casewith_skillbaselineΔshare of liftverdict
full-small-plan-header-and-tasks0.895 ± 0.019 (generation)0.415 ± 0.094 (generation)+0.48054.3%improved
task-right-sizing-testable-deliverable0.752 ± 0.152 (generation)0.449 ± 0.092 (generation)+0.30234.3%improved
repair-placeholder-steps0.848 ± 0.031 (generation)0.790 ± 0.033 (generation)+0.0586.6%no effect (bands overlap)
bite-sized-tdd-steps0.889 ± 0.013 (generation)0.860 ± 0.021 (generation)+0.0293.2%no effect (bands overlap)
interfaces-exact-signatures0.823 ± 0.033 (generation)0.796 ± 0.054 (generation)+0.0263.0%no effect (bands overlap)
self-review-spec-coverage0.850 ± 0.011 (generation)0.855 ± 0.009 (generation)-0.005-0.5%no effect (bands overlap)
file-structure-by-responsibility0.853 ± 0.013 (generation)0.861 ± 0.012 (generation)-0.007-0.8%no effect (bands overlap)

The aggregate lift of +0.126 ± 0.203 is not a band-verified improvement claim, and this page does not make one. The lift clears the 0.05 floor and fails the separation test, because the baseline aggregate band (± 0.198) overlaps the with_skill one. That width is produced by the 2 failing baseline cases in the arm the skill is meant to fix, not by instability in the with_skill arm (± 0.048). Reading +0.126 ± 0.203 as "the lift might be negative" is the reading that width invites and the per-case table above contradicts.

Across all 14 published cases on the release axis: 0 improved · 0 regressed · 14 within noise · 0 not measured.

Economics, on fresh input

Every dollar below excludes cached input tokens. The receipts price all input at the list input rate, cached included, and record cached_tokens per draw precisely so a reader can re-price. On this surface cached context is 90.1% to 92.2% of each call's input here, so on the list-rate basis the largest term in every mean is harness preamble that neither arm chose and the skill cannot move. Report 007 v1.1 is what happened when a mechanism was read out of that basis: two of its three cells changed the sign of their incremental under this same re-pricing. This page therefore prints one basis and does not claim the other.

Per-call incremental, with_skill minus baseline, WITHIN each report, on fresh input at each receipt's own frozen rates. Both figures in a row are the same arithmetic over different receipts, so the row is a like-for-like comparison; neither figure is an absolute cost and neither was metered.
skillon claude-fable-5 (Report 007)on claude-fable-5-1 (Report 008)direction
code-review-and-quality-$0.028848+$0.012699changes sign between the reports
writing-plans+$0.040498+$0.123954positive in both reports

The incremental is positive in both reports on writing-plans@claude-fable-5-1 and only on that cell. code-review-and-quality@claude-fable-5-1 is -$0.028848 on the older model and +$0.012699 on the newer one, so it changes sign between the reports on the fresh basis. A single sentence covering both cells would be wrong about one of them, and the difference is small enough in absolute terms that neither sign should be leaned on: the whole column is hundredths of a cent per call.

What moved in the tokens, writing-plans@claude-fable-5-1. The with_skill arm's per-call cost rises between the two models, and 77.3% of that rise is cached context priced at the list rate. 22.7% of it survives on fresh input, and it is chiefly output: +1,635 mean output tokens per call, at an output rate 5 times the input rate, against +2,177 fresh input tokens. That comparison is at an identical draw allocation: 24 draws in both reports, allocated 3,3,3,3,6,3,3 both times, with the same case in the wide slot. Adaptive stopping made the same decisions on the new model in this arm, so the movement is per-draw token growth and not a change in what was averaged. claude-fable-5-1 writes materially longer plans for the same seven cases.

And on code-review-and-quality@claude-fable-5-1, the same arithmetic points the other way. Its with_skill arm emits -264 output tokens per call on the new model and takes on +244 fresh input tokens, so on fresh input that arm gets cheaper per call across the release while the list-rate basis says it gets dearer. Longer outputs are a property of one of these two skills' generations on claude-fable-5-1, not of the release.

The caveat this whole section rests on, stated plainly. Cached input is not free to the vendor and is not billed at the list rate either; the receipts over-estimate it and this page under-estimates it by excluding it entirely, and neither basis is the true one. Actual metered spend for this run was $0.00: the surface is a subscription CLI and every figure here is the if-it-had-been-metered equivalent at each receipt's own frozen pricing snapshot. Prices are identical across all four receipts (10 in / 50 out per Mtok for both Fable models), so nothing above is price-driven. Judge cost is measurement overhead this project imposes and is excluded from every field. No figure on this page comes from the run's pre-flight projection or from any cap.

Latency, observed on subscription CLI surface, indicative: inside this run a call is slower with the skill than without it, by +1262 ms of median wall-clock on code-review-and-quality@claude-fable-5-1 and +2740 ms of median wall-clock on writing-plans@claude-fable-5-1. Wall-clock here includes cold starts and the vendor's own harness and is not a measurement of serving latency.

Run record

Launched against claude-fable-5-1 on 2026-09-01, from an isolated git worktree at 1b7b151 on spec/019b-fable-5-1-registration. code-review-and-quality ran 2026-09-01T18:45:54Z to 2026-09-01T20:44:42Z, writing-plans ran 2026-09-01T20:44:42Z to 2026-09-02T01:14:17Z, 258 calls and 294 calls, both exit 0. The receipts landed at a0aecb0.

The hashes were asserted before the first call, not after the last one. A pre-launch assertion exited 0 having compared, per cell, the skill content_hash and the suite_hash the run was about to use against the Report 007 baseline receipt it would be diffed against: code-review-and-quality (13d360d7f786, 4729eb8fc298, 7 cases) and writing-plans (5325ab3d55d5, 37cb3c341572, 7 cases). Both matched on both fields. That ordering is the whole attribution claim. Checking afterwards tells a reader what was measured; checking first is what makes it a controlled comparison, because a mismatch would have stopped the run before it spent anything. The differ re-confirms the same equality from the receipts themselves, so the assertion has a second, independent witness.

Zero draws were lost to timeout. 92 generation draws across 28 arms, 92 measured, 0 unmeasured. The per-call timeout was 300,000 ms, resolved from the surface policy rather than passed as a flag (lib/provider.js retryPolicyForSurface). Report 007's subject was that this policy had been declared and not executed since 2026-07-27, shadowed by a numeric literal at the call site, so every run before the fix ran at 120 s; this is the first release-drift run to execute under the policy as written, and the loss count is the control on that fix rather than a claim about this run's difficulty. A restart script was prepared before launch and never invoked.

Caps per cell: --max-cases 7, --max-calls 840, --max-usd 11. Judge claude-haiku-4-5 at 5 samples per draw on both cells. Surface claude-cli throughout, metered spend $0.00. Verification level TESTED on all four receipts, none carrying a run.source, so none is an imported declaration and the differ's NOT MEASURED path was not reached. The timings, the assertion result and the timeout policy above are transcribed from the run log into reports/report-008-run-record.json, which records that log's sha256 so the transcription is checkable; the log is an operator-home file and is not committed.

Limits

Receipts

All four receipts this report rests on, each one resolvable. Two measured here, two the Report 007 baselines they are diffed against. Every number on this page re-derives from these files.

Validate any of them with npx driftproof validate <file>, and reproduce either verdict with npx driftproof diff <baseline> <receipt>. The two receipts measured here landed at a0aecb0 and neither has been modified since it was sealed. receipts/report-007/writing-plans-claude-fable-5-2026-08-31.json also exists; it is Report 007's unpublished defect evidence for that cell, and this report neither reads nor modifies it.