Driftproof

Report #004 — does encoded expertise still lift output on the frontier tier?

Capability-gap report — cross-family, one provider, two tiers. NOT release drift: claude-fable-5 has no family predecessor.

3 durable · 5 tier-dependent · 0 regresses · 2 no effect — over 10 skills.

Encoded expertise survives the frontier tier — 3 durable, 5 tier-dependent, 0 regressions; where lift had collapsed, it had collapsed on the flagship first.

First cross-report repeatability check: all 10 opus-5 verdicts independently reproduce Report #003's character, including naming-analyzer's regression at case level.

The same 10 suites (Report #001 v1.2 rubrics) and the same fixed judge (claude-haiku-4-5, n=5) run on one provider's two tiers: claude-opus-5 (flagship) and claude-fable-5 (frontier). Per skill, the verdict rests on the with/without-skill delta on EACH tier — per-case band separation plus the 0.05 effect floor, the same anti-cry-wolf rule the whole series uses. The cross-tier comparison is context, never a ranking. The question this report asks is whether the frontier tier's baselines catch up — whether skills that earn their keep on the flagship become ceremonial at the edge.

Surface (disclosed): both tiers → claude-cli (subscription — claude -p -m; metered spend $0, the $ figure is the estimated metered-equivalent). Same surface, same fixed judge, same suites on both tiers — only the model changes.

Per-skill capability gap

skill claude-opus-5 with (± band, baseline, Δ lift) claude-fable-5 with (± band, baseline, Δ lift) lift shift (frontier − flagship) verdict
code-review-and-quality 0.875 ± 0.022 base 0.845 · Δ +0.030 0.880 ± 0.028 base 0.798 · Δ +0.082 +0.051 DURABLE
git-workflow-and-versioning 0.879 ± 0.011 base 0.878 · Δ +0.001 0.875 ± 0.011 base 0.876 · Δ -0.001 -0.002 NO EFFECT
documentation-and-adrs 0.843 ± 0.083 base 0.839 · Δ +0.004 0.862 ± 0.036 base 0.798 · Δ +0.064 +0.060 TIER-DEPENDENT † low-res
commit-work 0.878 ± 0.019 base 0.877 · Δ +0.001 0.877 ± 0.017 base 0.874 · Δ +0.003 +0.002 TIER-DEPENDENT
writing-clearly-and-concisely 0.856 ± 0.021 base 0.853 · Δ +0.003 0.870 ± 0.024 base 0.858 · Δ +0.013 +0.009 TIER-DEPENDENT
crafting-effective-readmes 0.881 ± 0.011 base 0.860 · Δ +0.021 0.886 ± 0.021 base 0.859 · Δ +0.027 +0.006 NO EFFECT
naming-analyzer 0.851 ± 0.108 base 0.820 · Δ +0.030 0.833 ± 0.069 base 0.757 · Δ +0.075 +0.045 TIER-DEPENDENT
requesting-code-review 0.835 ± 0.124 base 0.550 · Δ +0.285 0.855 ± 0.042 base 0.715 · Δ +0.140 -0.146 DURABLE † low-res
writing-plans 0.866 ± 0.028 base 0.689 · Δ +0.177 0.827 ± 0.109 base 0.705 · Δ +0.122 -0.055 DURABLE
skill-creator 0.893 ± 0.013 base 0.864 · Δ +0.029 0.889 ± 0.009 base 0.872 · Δ +0.017 -0.013 TIER-DEPENDENT

Columns: each tier's with_skill band (mean ± stddev over 5 judge samples) with its no-skill baseline and lift Δ; the lift shift (how much the skill's benefit changed moving flagship → frontier, context only); and the capability-gap verdict from the per-case with/without drivers on each tier. A shrinking lift with a RISING baseline means the model caught up — not that the skill decayed. Every number is re-derived from the receipts under receipts/report-004/; nothing is hand-entered.

Verdict basis — the per-case with/without drivers on each tier. A case drives a verdict only when its with_skill and baseline bands (mean ± stddev, n=5) do not overlap AND the mean moves ≥ 0.05. Cases marked †point-band rest on a zero-width band (all 5 judge samples identical).
Low-resolution: judge quantization (2). These verdicts rest — in whole or in part — on a zero-width judge point band (all 5 samples identical). The verdict stands; its confidence is grid-limited: documentation-and-adrs → TIER-DEPENDENT; requesting-code-review → DURABLE.

The documentation-and-adrs flag is the fragile one, stated plainly: its claude-opus-5 side reads MIXED (one hurting case, match-existing-adr-convention −0.206, against one lifting case, document-public-api-function +0.200) — and that lifting case rests on a zero-width point band. Discount point-band drivers and the opus-5 side becomes purely hurt, which would read REGRESSES on claude-opus-5 (the flagship — the claude-fable-5 side is two clean lifts either way). Flagged, not suppressed, per policy. requesting-code-review cannot flip: six of its seven drivers are clean bands.

Run record

Completed 2026-08-16 on the claude-cli surface — 20 receipts (10 skills × 2 tiers, with/without × n=5, judged by claude-haiku-4-5). Metered spend $0 (vendor subscription CLI); estimated metered-equivalent ~$16.08, under the $40 guard. All receipts complete. [resume: 0 restored, 0 skipped]

Registry note: claude-fable-5 registry released field updated null → 2026-06-09 for this report (verified from Anthropic announcements; launched 2026-06-09, export-control pause 06-12→30, redeployed 2026-07-01).

Receipts

All 20 receipts for this report, each one resolvable. Every number on this page re-derives from these files. They are the evidence, not a description of it — open any one and check it against the table above.

Directory: receipts/report-004/. Validate any of them with npx driftproof validate <file>.