Driftproof

Report #006 — the skill text moves, the substrate holds still

Revision drift report — a pinned skill revision against the revision upstream ships today, on a held substrate. The fifth report type: #001 and #003 moved the model release, #002 the vendor surface, #004 the capability tier, #005 the axes and the price. Here the skill's own text is the variable and everything underneath it is the control.

Three cells, one per skill upstream had revised since Report #005 pinned it, and none of the three returned a verdict. Each was refused by its own baseline control: the no-skill arm, which contains no skill text and which a revision cannot touch, failed to reproduce the arm #005 measured on the same model id, the same surface and the same suite. A 120-call stability probe then explained the failure: the baseline measurement is a single draw from a wide distribution, and generation-level noise — which this instrument never sampled — dominates judge-level noise, which it samples five times. No substrate movement is needed to account for any of it. This report publishes that instead of a verdict it cannot support.

Report #005 published a fairness disclosure saying that two of its ten pinned skills had already been revised upstream at the time it ran. This report is that disclosure measured rather than repeated. Where a revision improved a skill, #005's published figure understates the pack a reader can install today, and this report says so in the cell's own row and amends #005 rather than editing it. Where a cell lands within noise, the pin was still representative, and that is the finding — there is no salvage framing for a null.

How to read a revision-drift cell — and what it does NOT mean.
Disclosure.

Per-cell verdicts

cellrevision verdict#005 lift (pinned text)#006 lift (current text)Δ liftwhat this cell says
code-review-and-quality
claude-fable-5
⃠ NOT MEASURED
largest #005 effect (+0.103; runner-up sonnet-5 +0.035)
+0.103
pinned text · 534b6663e6ff…
+0.029
current text · 13d360d7f786…
n/a
no verdict emitted

Pre-registered: WITHIN NOISE — null control. The revision repoints two references/ link paths, and references/ is outside the measured surface by the manifest's own scoping rule.

Outcome: the cell returned no verdict, so the prediction is neither confirmed nor refuted.

Not measured: the fresh baseline does not reproduce the reused receipt's baseline (1 case(s) moved beyond the band and the 0.05 floor) — the reused pinned arm is not comparable to the fresh arm, so revision drift cannot be separated from whatever else changed in this cell; the control establishes non-reproduction and does not identify a cause

git-workflow-and-versioning
claude-sonnet-5
⃠ NOT MEASURED
largest #005 effect (+0.116; runner-up gpt-5.6-sol -0.002)
+0.116
pinned text · c0e9cf58fddf…
-0.001
current text · 2ed0afa4e21f…
n/a
no verdict emitted

Pre-registered: WITHIN NOISE — description-only revision, measurable as context and not as a trigger.

Outcome: the cell returned no verdict, so the prediction is neither confirmed nor refuted.

Scope of this cell: This revision changes the frontmatter description: line. In a skill runtime a description is a routing trigger: it decides whether the skill loads, and never reaches the model as guidance. Driftproof makes no routing decision — it always injects the skill, and passes the whole file, frontmatter included, as the system prompt. This cell therefore measures the revision as added context and cannot measure it as a trigger.

Not measured: the fresh baseline does not reproduce the reused receipt's baseline (3 case(s) moved beyond the band and the 0.05 floor) — the reused pinned arm is not comparable to the fresh arm, so revision drift cannot be separated from whatever else changed in this cell; the control establishes non-reproduction and does not identify a cause

writing-plans
claude-fable-5
⃠ NOT MEASURED
largest #005 effect (+0.177), but a TIE — sonnet-5 at +0.176, inside its own band
+0.177
pinned text · 5ac443162b52…
+0.055
current text · 5325ab3d55d5…
n/a
no verdict emitted

Pre-registered: the one cell with a plausible signal — a new required field in the plan-header template.

Outcome: the cell returned no verdict, so the prediction is neither confirmed nor refuted.

Not measured: the fresh baseline does not reproduce the reused receipt's baseline (2 case(s) moved beyond the band and the 0.05 floor) — the reused pinned arm is not comparable to the fresh arm, so revision drift cannot be separated from whatever else changed in this cell; the control establishes non-reproduction and does not identify a cause

Per cell: the classification under the shared verdict rule; the with/without lift #005 measured on the text it had pinned; the lift this report measures on the current upstream text, on the same substrate and the same suite; and the change between them. The Δ lift column is the difference of two aggregate lifts and is shown for orientation — the verdict is decided per case on band separation and the 0.05 floor, never on this column. Every number re-derives from the receipts under receipts/report-006/ and receipts/report-005/; nothing is hand-entered.

Verdict basis — the per-case drivers in each cell. A case drives a cell only when its with_skill bands under the pinned and current text do not overlap AND the mean moves at least 0.05. Cases marked †point-band rest on a zero-width band (all 5 judge samples identical).
The baseline-reproduction control, per cell. The check that makes the reused pinned arm honest: the fresh no-skill baseline re-measured against #005's, on the same rule as everything else.

Per-case detail

Why the baselines did not reproduce — a stability probe

The control proves non-reproduction. It cannot say why. That distinction is the whole of this section. An earlier draft of this report asserted that the substrate had moved — an explanation invented to fit the shape of the failure, with nothing measuring it. This replaces that assertion with 120 model calls.

Design. Two of the six non-reproducing cases, 10 fresh draws each, 5 judge samples per draw — 120 calls, $0 metered on the same subscription surface. Each draw is a new generation of the baseline arm, which carries no skill text. Same judge (claude-haiku-4-5, n=5), same surface (claude-cli), and the rubric hashes are asserted equal to the report's: a953d281f6a9… and c17f2f214cec…. Held in evidence/probe-baseline-stability-20260828.json, with per-draw rows in evidence/probe-baseline-stability-20260828.draws.jsonl. It is marked not_a_receipt: true and is exploratory: it grades no skill and attests no verdict.

severity-labeled-findings on claude-fable-5 — 10 draws

drawmean ± judge sdjudge samplesgeneration chars
1 (salvaged)0.626 ± 0.037[0.6,0.6,0.65,0.68,0.6]n/a
2 (salvaged)0.732 ± 0.091[0.85,0.63,0.76,0.77,0.65]n/a
3 (salvaged)0.690 ± 0.105[0.65,0.68,0.6,0.87,0.65]n/a
4 (salvaged)0.300 ± 0.000[0.3,0.3,0.3,0.3,0.3]n/a
50.506 ± 0.083[0.48,0.5,0.45,0.45,0.65]3842
60.686 ± 0.037[0.7,0.71,0.62,0.7,0.7]4332
70.774 ± 0.064[0.7,0.8,0.87,0.75,0.75]3691
80.808 ± 0.101[0.68,0.85,0.85,0.93,0.73]4070
90.734 ± 0.068[0.7,0.85,0.74,0.68,0.7]3697
100.300 ± 0.000[0.3,0.3,0.3,0.3,0.3]2897

Spread 0.300–0.808, mean 0.616, across-draw sd 0.186. Report #005's pinned value 0.300 lies inside this spread (z = -1.70); this report's re-measured value 0.696 lies inside it (z = 0.43). Verdict rule (b): UNSTABLE SINGLE DRAW — generation-level noise explains the gap.

commit-message-conventional-type on claude-sonnet-5 — 10 draws

drawmean ± judge sdjudge samplesgeneration chars
10.300 ± 0.000[0.3,0.3,0.3,0.3,0.3]311
20.290 ± 0.022[0.25,0.3,0.3,0.3,0.3]294
30.292 ± 0.011[0.3,0.3,0.28,0.3,0.28]242
40.286 ± 0.022[0.28,0.3,0.3,0.25,0.3]258
50.294 ± 0.013[0.3,0.3,0.3,0.27,0.3]259
60.290 ± 0.022[0.3,0.3,0.3,0.3,0.25]304
70.296 ± 0.009[0.3,0.3,0.3,0.28,0.3]253
80.214 ± 0.121[0.3,0.25,0.25,0.27,0]344
90.856 ± 0.005[0.86,0.86,0.85,0.85,0.86]343
100.278 ± 0.018[0.3,0.28,0.28,0.25,0.28]390

Spread 0.214–0.856, mean 0.340, across-draw sd 0.183. Report #005's pinned value 0.282 lies inside this spread (z = -0.31); this report's re-measured value 0.852 lies inside it (z = 2.80). Verdict rule (b): UNSTABLE SINGLE DRAW — generation-level noise explains the gap.

Where the variance actually lives

caseacross-draw sd (generation)mean within-draw sd (judge)ratio
severity-labeled-findings0.1860.0593.2×
commit-message-conventional-type0.1830.0247.5×

Generation-level noise dominates judge-level noise by 3.2× and 7.5×. Driftproof samples the judge five times and the generation once — so it has been putting its error bars on the smaller of the two sources. That is the finding, and it is about this instrument, not about any skill.

The explanation ladder, as it actually happened.
  1. "The substrate moved." The first reading, from the control alone. It fit every observation and rested on nothing.
  2. "The scores are bimodal." After the first draws showed a low mode and a high mode, this looked like a clean two-state story.
  3. "Case 1 is a broad continuous spread; case 2 is tight with a rare high mode." Where 20 draws actually land. Case 1 runs 0.300–0.808 with draws at 0.506, 0.626, 0.686, 0.690, 0.732, 0.734, 0.774 — not two modes, a spread. Case 2 sits between 0.214 and 0.300 in nine draws and jumped to 0.856 in one.

The probe session retracted its own claim mid-run. With eight tight draws on case 2 in hand, it computed this report's re-measured 0.852 as roughly 122 standard deviations from the mean — a number that reads as impossible. Draw 9 then returned 0.856, landing squarely on the high mode and demolishing that statistic. The "122 sigma" figure was an artifact of estimating a standard deviation from a sample that had not yet met the tail. It is recorded here because a retraction that leaves no trace is not a retraction; the draw-by-draw sequence is in evidence/probe-baseline-stability-20260828.progress.txt.

What this does and does not license us to say.
Prior work, and what is actually new here.

Sampling variance in LLM evaluation is established ground, not our discovery. Miller (2024), arXiv:2411.00640, sets out resampling and error bars for exactly this class of measurement; arXiv:2502.08943 quantifies how badly single-sample evaluation understates variance in benchmark scores. Both describe what the probe re-derived the hard way.

What we claim is narrower and, as far as we can tell, unoccupied: applying it to agent-skill verification, and turning it on our own published numbers. The literature says single-draw benchmarks are unreliable. This report is a verification project discovering that its own receipts inherited the defect, publishing the draws that show it, and changing the receipt specification because of it — rather than leaving the finding as a caveat about someone else's work.

Run record

Current-text arms run 2026-08-28 19:38 UTC, receipt-attested. 3 fresh receipts (3 skills × 1 substrate each, with/without × n=5, judged by claude-haiku-4-5), 252 model calls. Pinned-text arms reused from Report #005 at no cost. Surfaces: claude-fable-5 → claude-cli, claude-sonnet-5 → claude-cli. Metered spend $0 (vendor subscription CLIs); metered-equivalent $18.50 against a pre-registered basis of $26.88. Verification level: TESTED, the weakest among the receipts summed. Upstream revision scan: state/skill-version-check.json, the 2026-08-19 record #005's disclosure was written from — cited by path and date, not linked: it is kept in the source repository and build-public.sh excludes it from the published tree.

Relation to Report #005

This report exists because #005's fairness disclosure named the exposure before anything had measured it: eight of ten skills were still byte-identical to their pins on 2026-08-19, and two were not. A third, git-workflow-and-versioning, was revised upstream after that check was made — #005's check was complete when it was made, and the timing is disclosed here rather than being allowed to look like an omission.

No cell in this report was measured, so nothing here amends a #005 figure — and nothing here confirms one either. A within-noise result would have said the pins were still representative. That is not what happened: every cell was refused by its own baseline control before a verdict existed, so this report has no evidence about the revisions in either direction.

It does raise a question about #005, and the probe above sharpens it without answering it. #005's per-cell lifts rest on baselines measured with one generation each, and two of those baselines are now shown to sit in spreads wide enough that a single draw does not pin them. This report's re-measured lifts are single draws from the same spreads and are not replacements for #005's figures. On commit-message-conventional-type the direction of the doubt actually runs toward this report: #005's 0.282 is the modal value across ten draws, and the 0.852 measured here was the roughly 1-in-10 outcome. The question of what #005's lifts would be under proper generation sampling is left open, because nothing run so far can close it.

Amendments

v1.1 · 2026-09-01. The writing-plans cell's aggregate was computed over two different case sets, and is corrected here. This page publishes that cell at +0.055. Under correct pairing it is +0.031, and its band moves from ± 0.194 to ± 0.038.

The mechanism, because the counts hid it. The two arms were reported as 6 against 6, which looks paired and is not: two different cases each lost one arm to the same 120 s call timeout, task-right-sizing-testable-deliverable (baseline) and full-small-plan-header-and-tasks (with_skill). Filtering each arm's failures independently left six rows on each side over different case sets, so the aggregate subtracted a mean of one set of cases from a mean of another. Excluding a case from both arms when either arm is unmeasured gives 5 against 5, genuinely paired. The counts matching at 6 against 6 was a coincidence, which is why nothing on this page or in the receipt suggested anything was wrong. run.failed_case_count reads 2 before and after; what the receipt did not carry was the list, and excluded_cases now names it.

The band moves further than the lift does: ± 0.194 to ± 0.038, a fivefold narrowing, because the two cases that dropped out were the two widest. The corrected lift of +0.031 against ± 0.038 is still inside its own band, narrowly, as +0.055 against ± 0.194 was comfortably.

The verdict does not change, and neither does the control. This cell's classification stays NOT MEASURED. The baseline-reproduction control that produced it is driven by per-case movement, and it is recorded in a separate control record whose per-case table pairwise exclusion does not touch: moved_cases is still bite-sized-tdd-steps, interfaces-exact-signatures, two cases past the band and the 0.05 floor, and the reason string is unchanged. Exclusion changes aggregates only, and results.cases is left intact, so no per-case comparison moves.

One disclosed diagnostic inside the control block does move, and is stated rather than left to be found. aggregate_baseline_delta is computed over the baseline arm's aggregate, and that aggregate goes from 0.754 over 6 cases to 0.820 over 5. The figure therefore moves from +0.099 to +0.165. It is a diagnostic and not the verdict basis, since the reason string counts cases rather than reading the aggregate, but it is a published-adjacent number that this correction touches, and leaving it unmentioned would be the same omission in miniature.

No figure on this page has been edited, and the archived receipt is byte-unchanged. The corrected figures above are rebuilt from that receipt's own case rows through the current aggregate path. Filed by Report 007, which measured the timeout that caused both losses.

v1.0 · 2026-08-29 — First publication.

Receipts

Receipts

All 3 receipts for this report, each one resolvable. Every number on this page re-derives from these files. They are the evidence, not a description of it — open any one and check it against the table above.

Directory: receipts/report-006/. Validate any of them with npx driftproof validate <file>.

Receipts are named as bin/driftproof run writes them, <slug>-<model>-<date>.json. Beside each one sits a control record, <same base>.control.json, holding that cell's baseline-reproduction result and its revision-pair preconditions — the persisted proof that no verdict was emitted past a failed control. A cell that passed its control would also carry a rendered drift report; none of the three did, and the absence of those files is itself the evidence. The pinned-text arm of every cell is a Report #005 receipt, linked from that report's own receipts block.

The three control records. Each one names the cases whose no-skill baseline failed to reproduce, and is re-readable rather than logged.