Report #003 — do skill verdicts hold across a model release?
Release drift report
2 regressed · 4 improved · 0 mixed · 4 within noise — over 10 skills.
The same 10 suites (Report #001 v1.2 rubrics) and the same fixed judge (claude-haiku-4-5, n=5) run on two versions of one model: claude-opus-4-8 (older) → claude-opus-5 (newer). Per skill, the verdict compares the two versions' with_skill bands per case; a regression/improvement is claimed only when BOTH (1) the two bands (mean ± stddev over the 5 judge samples) do not overlap AND (2) the mean moves at least the 0.05 effect floor. Bands that overlap, or separate by less than the floor, are within noise — never drift. See the methodology.
Surface (disclosed): both versions → claude-cli (subscription — claude -p -m; metered spend $0, the $ figure is the estimated metered-equivalent). Same surface, same fixed judge, same suites for both columns — the only thing that moves is the model version, which is the point.
Per-skill drift
| skill | with_skill (old → new) | baseline (old → new) | verdict |
|---|---|---|---|
code-review-and-quality |
0.887 ± 0.024 → 0.886 ± 0.016 Δ -0.000 | 0.819 → 0.833 | WITHIN NOISE |
git-workflow-and-versioning |
0.879 ± 0.009 → 0.877 ± 0.011 Δ -0.002 | 0.839 → 0.878 | WITHIN NOISE |
documentation-and-adrs |
0.769 ± 0.232 → 0.867 ± 0.032 Δ +0.098 | 0.795 → 0.839 | IMPROVED (1) |
commit-work |
0.882 ± 0.015 → 0.882 ± 0.014 Δ -0.000 | 0.847 → 0.886 | WITHIN NOISE |
writing-clearly-and-concisely |
0.876 ± 0.024 → 0.845 ± 0.033 Δ -0.030 | 0.874 → 0.861 | REGRESSED (1) |
crafting-effective-readmes |
0.760 ± 0.213 → 0.871 ± 0.032 Δ +0.111 | 0.870 → 0.795 | IMPROVED (2) |
naming-analyzer |
0.874 ± 0.012 → 0.866 ± 0.054 Δ -0.008 | 0.797 → 0.798 | REGRESSED (1) |
requesting-code-review |
0.762 ± 0.214 → 0.871 ± 0.025 Δ +0.109 | 0.544 → 0.671 | IMPROVED (1) |
writing-plans |
0.864 ± 0.033 → 0.883 ± 0.028 Δ +0.019 | 0.578 → 0.690 | IMPROVED (1) † low-res |
skill-creator |
0.887 ± 0.007 → 0.885 ± 0.008 Δ -0.001 | 0.849 → 0.889 | WITHIN NOISE |
Columns: each skill's with_skill band (mean ± stddev over 5 judge samples) on the old version → the new version, with the aggregate Δ shown as context; the baseline (no-skill) mean old → new; and the drift verdict. The verdict does not rest on the aggregate Δ — following Report #001's anti-cry-wolf discipline it is the tally of per-case band-separated verdicts that clear the 0.05 floor. A † low-res mark means a driving case rests on a zero-width judge point band. Every number is re-derived from the receipts under receipts/report-003/; nothing here is hand-entered.
Verdict basis — the per-case band-separated drivers behind each label. A case drives a verdict only when its two with_skill bands (old vs new, mean ± stddev, n=5) do not overlap AND the mean moves ≥ 0.05. Cases marked †point-band rest on a zero-width band (all 5 judge samples identical) — see the low-resolution note.
documentation-and-adrs: 🔼surface-conflicting-adr-conventions(0.246 ± 0.058 → 0.890 ± 0.012, Δ+0.644)writing-clearly-and-concisely: 🔻positive-form-status-sentences(0.916 ± 0.032 → 0.866 ± 0.011, Δ-0.050)crafting-effective-readmes: 🔼categorize-task-before-writing(0.588 ± 0.096 → 0.828 ± 0.050, Δ+0.240), 🔼lead-with-one-sentence-problem(0.346 ± 0.082 → 0.870 ± 0.012, Δ+0.524)naming-analyzer: 🔻abbreviations-wellknown-js(0.872 ± 0.013 → 0.744 ± 0.096, Δ-0.128)requesting-code-review: 🔼crafted-context-not-session-history(0.292 ± 0.011 → 0.906 ± 0.027, Δ+0.614)writing-plans: 🔼interfaces-exact-signatures(0.800 ± 0.000 → 0.886 ± 0.030, Δ+0.086) †point-band
with_skill band is zero-width: all 5 judge samples returned the identical score. The judge grades on a coarse quantization grid, so a zero-width band is a clean grid-step effect the 0.05 floor still gates, but one whose finer structure the judge cannot resolve. The verdict stands; its confidence is grid-limited: writing-plans → IMPROVED (1) (interfaces-exact-signatures).Run record
Completed 2026-08-12 on the claude-cli surface — 20 receipts (10 skills × 2 versions, with/without × n=5, judged by claude-haiku-4-5). Metered spend $0 (vendor subscription CLI); estimated metered-equivalent ~$11.79, under the $40 guard. All 20 receipts complete; zero timeouts, zero cases excluded.