Driftproof

Report #003 — do skill verdicts hold across a model release?

Release drift report

2 regressed · 4 improved · 0 mixed · 4 within noise — over 10 skills.

The same 10 suites (Report #001 v1.2 rubrics) and the same fixed judge (claude-haiku-4-5, n=5) run on two versions of one model: claude-opus-4-8 (older) → claude-opus-5 (newer). Per skill, the verdict compares the two versions' with_skill bands per case; a regression/improvement is claimed only when BOTH (1) the two bands (mean ± stddev over the 5 judge samples) do not overlap AND (2) the mean moves at least the 0.05 effect floor. Bands that overlap, or separate by less than the floor, are within noise — never drift. See the methodology.

Surface (disclosed): both versions → claude-cli (subscription — claude -p -m; metered spend $0, the $ figure is the estimated metered-equivalent). Same surface, same fixed judge, same suites for both columns — the only thing that moves is the model version, which is the point.

Per-skill drift

skill with_skill (old → new) baseline (old → new) verdict
code-review-and-quality 0.887 ± 0.024 → 0.886 ± 0.016 Δ -0.000 0.819 → 0.833 WITHIN NOISE
git-workflow-and-versioning 0.879 ± 0.009 → 0.877 ± 0.011 Δ -0.002 0.839 → 0.878 WITHIN NOISE
documentation-and-adrs 0.769 ± 0.232 → 0.867 ± 0.032 Δ +0.098 0.795 → 0.839 IMPROVED (1)
commit-work 0.882 ± 0.015 → 0.882 ± 0.014 Δ -0.000 0.847 → 0.886 WITHIN NOISE
writing-clearly-and-concisely 0.876 ± 0.024 → 0.845 ± 0.033 Δ -0.030 0.874 → 0.861 REGRESSED (1)
crafting-effective-readmes 0.760 ± 0.213 → 0.871 ± 0.032 Δ +0.111 0.870 → 0.795 IMPROVED (2)
naming-analyzer 0.874 ± 0.012 → 0.866 ± 0.054 Δ -0.008 0.797 → 0.798 REGRESSED (1)
requesting-code-review 0.762 ± 0.214 → 0.871 ± 0.025 Δ +0.109 0.544 → 0.671 IMPROVED (1)
writing-plans 0.864 ± 0.033 → 0.883 ± 0.028 Δ +0.019 0.578 → 0.690 IMPROVED (1) † low-res
skill-creator 0.887 ± 0.007 → 0.885 ± 0.008 Δ -0.001 0.849 → 0.889 WITHIN NOISE

Columns: each skill's with_skill band (mean ± stddev over 5 judge samples) on the old version → the new version, with the aggregate Δ shown as context; the baseline (no-skill) mean old → new; and the drift verdict. The verdict does not rest on the aggregate Δ — following Report #001's anti-cry-wolf discipline it is the tally of per-case band-separated verdicts that clear the 0.05 floor. A † low-res mark means a driving case rests on a zero-width judge point band. Every number is re-derived from the receipts under receipts/report-003/; nothing here is hand-entered.

Verdict basis — the per-case band-separated drivers behind each label. A case drives a verdict only when its two with_skill bands (old vs new, mean ± stddev, n=5) do not overlap AND the mean moves ≥ 0.05. Cases marked †point-band rest on a zero-width band (all 5 judge samples identical) — see the low-resolution note.
Low-resolution: judge quantization (1). These verdicts are supported — in whole or in part — by a case whose old or new with_skill band is zero-width: all 5 judge samples returned the identical score. The judge grades on a coarse quantization grid, so a zero-width band is a clean grid-step effect the 0.05 floor still gates, but one whose finer structure the judge cannot resolve. The verdict stands; its confidence is grid-limited: writing-plans → IMPROVED (1) (interfaces-exact-signatures).

Run record

Completed 2026-08-12 on the claude-cli surface — 20 receipts (10 skills × 2 versions, with/without × n=5, judged by claude-haiku-4-5). Metered spend $0 (vendor subscription CLI); estimated metered-equivalent ~$11.79, under the $40 guard. All 20 receipts complete; zero timeouts, zero cases excluded.