Report #004 — does encoded expertise still lift output on the frontier tier?
Capability-gap report — cross-family, one provider, two tiers. NOT release drift: claude-fable-5 has no family predecessor.
3 durable · 5 tier-dependent · 0 regresses · 2 no effect — over 10 skills.
Encoded expertise survives the frontier tier — 3 durable, 5 tier-dependent, 0 regressions; where lift had collapsed, it had collapsed on the flagship first.
First cross-report repeatability check: all 10 opus-5 verdicts independently reproduce Report #003's character, including naming-analyzer's regression at case level.
The same 10 suites (Report #001 v1.2 rubrics) and the same fixed judge (claude-haiku-4-5, n=5) run on one provider's two tiers: claude-opus-5 (flagship) and claude-fable-5 (frontier). Per skill, the verdict rests on the with/without-skill delta on EACH tier — per-case band separation plus the 0.05 effect floor, the same anti-cry-wolf rule the whole series uses. The cross-tier comparison is context, never a ranking. The question this report asks is whether the frontier tier's baselines catch up — whether skills that earn their keep on the flagship become ceremonial at the edge.
Surface (disclosed): both tiers → claude-cli (subscription — claude -p -m; metered spend $0, the $ figure is the estimated metered-equivalent). Same surface, same fixed judge, same suites on both tiers — only the model changes.
Per-skill capability gap
| skill | claude-opus-5 with (± band, baseline, Δ lift) | claude-fable-5 with (± band, baseline, Δ lift) | lift shift (frontier − flagship) | verdict |
|---|---|---|---|---|
code-review-and-quality |
0.875 ± 0.022 base 0.845 · Δ +0.030 | 0.880 ± 0.028 base 0.798 · Δ +0.082 | +0.051 | DURABLE |
git-workflow-and-versioning |
0.879 ± 0.011 base 0.878 · Δ +0.001 | 0.875 ± 0.011 base 0.876 · Δ -0.001 | -0.002 | NO EFFECT |
documentation-and-adrs |
0.843 ± 0.083 base 0.839 · Δ +0.004 | 0.862 ± 0.036 base 0.798 · Δ +0.064 | +0.060 | TIER-DEPENDENT † low-res |
commit-work |
0.878 ± 0.019 base 0.877 · Δ +0.001 | 0.877 ± 0.017 base 0.874 · Δ +0.003 | +0.002 | TIER-DEPENDENT |
writing-clearly-and-concisely |
0.856 ± 0.021 base 0.853 · Δ +0.003 | 0.870 ± 0.024 base 0.858 · Δ +0.013 | +0.009 | TIER-DEPENDENT |
crafting-effective-readmes |
0.881 ± 0.011 base 0.860 · Δ +0.021 | 0.886 ± 0.021 base 0.859 · Δ +0.027 | +0.006 | NO EFFECT |
naming-analyzer |
0.851 ± 0.108 base 0.820 · Δ +0.030 | 0.833 ± 0.069 base 0.757 · Δ +0.075 | +0.045 | TIER-DEPENDENT |
requesting-code-review |
0.835 ± 0.124 base 0.550 · Δ +0.285 | 0.855 ± 0.042 base 0.715 · Δ +0.140 | -0.146 | DURABLE † low-res |
writing-plans |
0.866 ± 0.028 base 0.689 · Δ +0.177 | 0.827 ± 0.109 base 0.705 · Δ +0.122 | -0.055 | DURABLE |
skill-creator |
0.893 ± 0.013 base 0.864 · Δ +0.029 | 0.889 ± 0.009 base 0.872 · Δ +0.017 | -0.013 | TIER-DEPENDENT |
Columns: each tier's with_skill band (mean ± stddev over 5 judge samples) with its no-skill baseline and lift Δ; the lift shift (how much the skill's benefit changed moving flagship → frontier, context only); and the capability-gap verdict from the per-case with/without drivers on each tier. A shrinking lift with a RISING baseline means the model caught up — not that the skill decayed. Every number is re-derived from the receipts under receipts/report-004/; nothing is hand-entered.
Verdict basis — the per-case with/without drivers on each tier. A case drives a verdict only when its with_skill and baseline bands (mean ± stddev, n=5) do not overlap AND the mean moves ≥ 0.05. Cases marked †point-band rest on a zero-width band (all 5 judge samples identical).
code-review-and-quality— claude-opus-5: 🔼severity-labeled-findings(Δ+0.182) · claude-fable-5: 🔼severity-labeled-findings(Δ+0.562)documentation-and-adrs— claude-opus-5: 🔼document-public-api-function(Δ+0.200) †point-band, 🔻match-existing-adr-convention(Δ-0.206) · claude-fable-5: 🔼comment-intent-not-implementation(Δ+0.232) †point-band, 🔼document-public-api-function(Δ+0.202) †point-bandcommit-work— claude-opus-5: 🔼conventional-commit-single-change(Δ+0.062)writing-clearly-and-concisely— claude-fable-5: 🔼emphatic-word-at-end(Δ+0.076)naming-analyzer— claude-opus-5: 🔼boolean-prefixes-js(Δ+0.054), 🔻abbreviations-wellknown-js(Δ-0.140), 🔼language-casing-python(Δ+0.276) · claude-fable-5: 🔼boolean-prefixes-js(Δ+0.160), 🔼abbreviations-wellknown-js(Δ+0.332)requesting-code-review— claude-opus-5: 🔼crafted-context-not-session-history(Δ+0.606), 🔼mandatory-vs-optional-triggers(Δ+0.576), 🔼resist-simple-self-review(Δ+0.234), 🔼full-handoff-before-merge(Δ+0.410) †point-band · claude-fable-5: 🔼triage-review-findings(Δ+0.178), 🔼resist-simple-self-review(Δ+0.212), 🔼full-handoff-before-merge(Δ+0.358)writing-plans— claude-opus-5: 🔼bite-sized-tdd-steps(Δ+0.240), 🔼file-structure-by-responsibility(Δ+0.054), 🔼task-right-sizing-testable-deliverable(Δ+0.342), 🔼full-small-plan-header-and-tasks(Δ+0.518) · claude-fable-5: 🔼bite-sized-tdd-steps(Δ+0.240), 🔼full-small-plan-header-and-tasks(Δ+0.494)skill-creator— claude-opus-5: 🔼domain-variant-organization(Δ+0.134)
documentation-and-adrs → TIER-DEPENDENT; requesting-code-review → DURABLE.
The documentation-and-adrs flag is the fragile one, stated plainly: its claude-opus-5 side reads MIXED (one hurting case, match-existing-adr-convention −0.206, against one lifting case, document-public-api-function +0.200) — and that lifting case rests on a zero-width point band. Discount point-band drivers and the opus-5 side becomes purely hurt, which would read REGRESSES on claude-opus-5 (the flagship — the claude-fable-5 side is two clean lifts either way). Flagged, not suppressed, per policy. requesting-code-review cannot flip: six of its seven drivers are clean bands.
Run record
Completed 2026-08-16 on the claude-cli surface — 20 receipts (10 skills × 2 tiers, with/without × n=5, judged by claude-haiku-4-5). Metered spend $0 (vendor subscription CLI); estimated metered-equivalent ~$16.08, under the $40 guard. All receipts complete. [resume: 0 restored, 0 skipped]
Registry note: claude-fable-5 registry released field updated null → 2026-06-09 for this report (verified from Anthropic announcements; launched 2026-06-09, export-control pause 06-12→30, redeployed 2026-07-01).
Receipts
All 20 receipts for this report, each one resolvable. Every number on this page re-derives from these files. They are the evidence, not a description of it — open any one and check it against the table above.
code-review-and-quality— claude-fable-5 · claude-opus-5commit-work— claude-fable-5 · claude-opus-5crafting-effective-readmes— claude-fable-5 · claude-opus-5documentation-and-adrs— claude-fable-5 · claude-opus-5git-workflow-and-versioning— claude-fable-5 · claude-opus-5naming-analyzer— claude-fable-5 · claude-opus-5requesting-code-review— claude-fable-5 · claude-opus-5skill-creator— claude-fable-5 · claude-opus-5writing-clearly-and-concisely— claude-fable-5 · claude-opus-5writing-plans— claude-fable-5 · claude-opus-5
Directory: receipts/report-004/. Validate any of them with npx driftproof validate <file>.