Report #003 — do skill verdicts hold across a model release?
What did Report 003 find?
- What we tested. Ten skills, seven tasks each, on claude-opus-4-8 and the newer claude-opus-5, with and without the skill. Another model scored each answer five times.
- What we found. With the skill, four skills had at least one task clearly higher on the newer model, two had one task clearly lower, and four showed no clear difference.
- One gain rests on coarse scores. For writing-plans, the older model's five scores on the deciding task were identical, too coarse to show finer differences.
- What it doesn't show. Each task was answered once each way, with one grader. No clear difference does not show nothing changed, and a clear difference does not prove the model's behaviour moved.
Release drift report
2 regressed · 4 improved · 0 mixed · 4 within noise — over 10 skills.
The same 10 suites (Report #001 v1.2 rubrics) and the same fixed judge (claude-haiku-4-5, n=5) run on two versions of one model: claude-opus-4-8 (older) → claude-opus-5 (newer). Per skill, the verdict compares the two versions' with_skill bands per case; a regression/improvement is claimed only when BOTH (1) the two bands (mean ± stddev over the 5 judge samples) do not overlap AND (2) the mean moves at least the 0.05 effect floor. Bands that overlap, or separate by less than the floor, are within noise — never drift. See the methodology.
Note, 2026-10-03. The headline counts on this page include receipts whose receipt pages mark them Not measured. Corrected counts follow in spec 159.
Surface (disclosed): both versions → claude-cli (subscription — claude -p -m; metered spend $0, the $ figure is the estimated metered-equivalent). Same surface, same fixed judge, same suites for both columns — the only thing that moves is the model version, which is the point.
Per-skill drift
| skill | with_skill (old → new) | baseline (old → new) | verdict |
|---|---|---|---|
code-review-and-quality |
0.887 ± 0.024 → 0.886 ± 0.016 Δ -0.000 | 0.819 → 0.833 | WITHIN NOISE |
git-workflow-and-versioning |
0.879 ± 0.009 → 0.877 ± 0.011 Δ -0.002 | 0.839 → 0.878 | WITHIN NOISE |
documentation-and-adrs |
0.769 ± 0.232 → 0.867 ± 0.032 Δ +0.098 | 0.795 → 0.839 | IMPROVED (1) |
commit-work |
0.882 ± 0.015 → 0.882 ± 0.014 Δ -0.000 | 0.847 → 0.886 | WITHIN NOISE |
writing-clearly-and-concisely |
0.876 ± 0.024 → 0.845 ± 0.033 Δ -0.030 | 0.874 → 0.861 | REGRESSED (1) |
crafting-effective-readmes |
0.760 ± 0.213 → 0.871 ± 0.032 Δ +0.111 | 0.870 → 0.795 | IMPROVED (2) |
naming-analyzer |
0.874 ± 0.012 → 0.866 ± 0.054 Δ -0.008 | 0.797 → 0.798 | REGRESSED (1) |
requesting-code-review |
0.762 ± 0.214 → 0.871 ± 0.025 Δ +0.109 | 0.544 → 0.671 | IMPROVED (1) |
writing-plans |
0.864 ± 0.033 → 0.883 ± 0.028 Δ +0.019 | 0.578 → 0.690 | IMPROVED (1) † low-res |
skill-creator |
0.887 ± 0.007 → 0.885 ± 0.008 Δ -0.001 | 0.849 → 0.889 | WITHIN NOISE |
Columns: each skill's with_skill band (mean ± stddev over 5 judge samples) on the old version → the new version, with the aggregate Δ shown as context; the baseline (no-skill) mean old → new; and the drift verdict. The verdict does not rest on the aggregate Δ — following Report #001's anti-cry-wolf discipline it is the tally of per-case band-separated verdicts that clear the 0.05 floor. A † low-res mark means a driving case rests on a zero-width judge point band. Every number is re-derived from the receipts under receipts/report-003/; nothing here is hand-entered.
Verdict basis — the per-case band-separated drivers behind each label. A case drives a verdict only when its two with_skill bands (old vs new, mean ± stddev, n=5) do not overlap AND the mean moves ≥ 0.05. Cases marked †point-band rest on a zero-width band (all 5 judge samples identical) — see the low-resolution note.
documentation-and-adrs: 🔼surface-conflicting-adr-conventions(0.246 ± 0.058 → 0.890 ± 0.012, Δ+0.644)writing-clearly-and-concisely: 🔻positive-form-status-sentences(0.916 ± 0.032 → 0.866 ± 0.011, Δ-0.050)crafting-effective-readmes: 🔼categorize-task-before-writing(0.588 ± 0.096 → 0.828 ± 0.050, Δ+0.240), 🔼lead-with-one-sentence-problem(0.346 ± 0.082 → 0.870 ± 0.012, Δ+0.524)naming-analyzer: 🔻abbreviations-wellknown-js(0.872 ± 0.013 → 0.744 ± 0.096, Δ-0.128)requesting-code-review: 🔼crafted-context-not-session-history(0.292 ± 0.011 → 0.906 ± 0.027, Δ+0.614)writing-plans: 🔼interfaces-exact-signatures(0.800 ± 0.000 → 0.886 ± 0.030, Δ+0.086) †point-band
with_skill band is zero-width: all 5 judge samples returned the identical score. The judge grades on a coarse quantization grid, so a zero-width band is a clean grid-step effect the 0.05 floor still gates, but one whose finer structure the judge cannot resolve. The verdict stands; its confidence is grid-limited: writing-plans → IMPROVED (1) (interfaces-exact-signatures).Run record
Completed 2026-08-12 on the claude-cli surface — 20 receipts (10 skills × 2 versions, with/without × n=5, judged by claude-haiku-4-5). Metered spend $0 (vendor subscription CLI); estimated metered-equivalent ~$11.79, under the $40 guard. All 20 receipts complete; zero timeouts, zero cases excluded.
Amendments
v1.1 · 2026-09-14. This entry corrects how the page’s band wording is read; no earlier text is changed, and no figure, verdict token, table value or receipt reference changes. The headline says “2 regressed · 4 improved · 0 mixed · 4 within noise”, the table labels four skills WITHIN NOISE, and the introduction says “Bands that overlap, or separate by less than the floor, are within noise” and “never drift”. None of those four skills had a with_skill case separated under the rule between the two models: no separation detected at the sample size used, which is not evidence of equivalence, not evidence that nothing changed, and not a finding of no drift. The labels are the verdict values this report recorded, and are left as published. As amended:
2 regressed · 4 improved · 0 mixed · 4 with no separation detected over 10 skills.
How to read this page. A case is separated under the rule when its two bands do not overlap and its mean moved by at least the 0.05 effect floor; bands that do not overlap by a smaller move are not separated. What the introduction calls bands that separate by less than the floor are bands that do not overlap. Wherever the page reads a result that was not separated as settled, as within noise or never drift, read it as no separation detected at the sample size used. Wherever it reads a separation as settled, as the low-resolution note’s The verdict stands, read it as a separation detected under the rule, which is not proof that the model’s behaviour moved; that note’s confidence is not a property of any band on this page. Labels and headings are read the same way and are left as published.
The bands. Each band on this page is a mean plus or minus one sample standard deviation, of two kinds: a case’s band, over its five judge samples, and a skill’s with_skill band, over its seven per-case means, which is suite dispersion. The column note’s band (mean ± stddev over 5 judge samples), said of the skill bands, reads as suite dispersion. Each band is a descriptive spread, not a confidence interval, with no coverage probability. The site’s report index and feed quote the amended headline; this page’s summary card, description and social card image keep the headline as published. Filed under the wording rules of the repository's spec 031, amendment A-031-20.
v1.2 · 2026-10-03. This entry adds a dated note under the headline; no earlier text is changed, and no figure, verdict token, table value or receipt reference changes. The note says the headline counts on this page include receipts whose receipt pages mark them Not measured, and that corrected counts follow in spec 159. It was written after each of the 20 receipts this page links was read with the badge's own rule, and all 20 read Not measured.
Receipts
All 20 receipts for this report, each one resolvable. Every number on this page re-derives from these files. They are the evidence, not a description of it — open any one and check it against the table above.
code-review-and-quality— claude-opus-4-8 · claude-opus-5commit-work— claude-opus-4-8 · claude-opus-5crafting-effective-readmes— claude-opus-4-8 · claude-opus-5documentation-and-adrs— claude-opus-4-8 · claude-opus-5git-workflow-and-versioning— claude-opus-4-8 · claude-opus-5naming-analyzer— claude-opus-4-8 · claude-opus-5requesting-code-review— claude-opus-4-8 · claude-opus-5skill-creator— claude-opus-4-8 · claude-opus-5writing-clearly-and-concisely— claude-opus-4-8 · claude-opus-5writing-plans— claude-opus-4-8 · claude-opus-5
Directory: receipts/report-003/. Validate any of them with npx driftproof validate <file>.