Driftproof

Report #001 — do skill verdicts hold across a model release?

What did Report 001 find?

  • What we tested. Ten skills, seven tasks each, on claude-sonnet-4-6 and the newer claude-sonnet-5, with and without the skill. Another model scored each answer five times.
  • What we found. With the skill, nine of ten skills had at least one task scoring clearly higher or clearly lower on the newer model. No skill's average over its seven tasks, with the skill, clearly differed.
  • Skills went different ways. Two skills had only clearly lower tasks, three only clearly higher, four had both, and one showed no clear difference.
  • What it doesn't show. Tasks we wrote, each answered once each way, with one grader. A clear difference does not prove the model changed; no clear difference does not show nothing changed.

Release drift report

v1.2 (2026-07-30) — this edition supersedes v1.1. An in-text grounding policy was codified and applied across all 10 suites; imported rubric criteria in git-workflow-and-versioning and commit-work were amended and those two suites re-run. See Amendments.

Model pair: claude-sonnet-5 (released 2026-06-30) vs claude-sonnet-4-6 (released 2026-02-17)
Judge: claude-haiku-4-5, 5 samples/case (surface-controlled) · Surface: claude-cli · Skills: 10

9 of 10 skills moved beyond noise on this model pair.

2 regressed · 3 improved · 4 mixed · 1 within noise — across 9 case regression(s) and 8 case improvement(s). A verdict is claimed only when BOTH (1) two with_skill bands (mean ± stddev over judge samples) do not overlap AND (2) the mean moves at least the 0.05 effect floor; bands that overlap, or separate by less than the floor, are reported as within noise, never as drift.

Note, 2026-10-03. The headline counts on this page include receipts whose receipt pages mark them Not measured. Corrected counts follow in spec 159.

Per-skill summary

skillsourcewith_skill (old → new)baseline (old → new)verdict
code-review-and-quality addyosmani · MIT 0.869 ± 0.025 → 0.854 ± 0.031 0.824 → 0.836 WITHIN NOISE
git-workflow-and-versioning addyosmani · MIT 0.868 ± 0.025 → 0.834 ± 0.104 0.788 → 0.859 MIXED
documentation-and-adrs addyosmani · MIT 0.825 ± 0.100 → 0.706 ± 0.297 0.809 → 0.720 MIXED
commit-work Softaworks · MIT 0.825 ± 0.082 → 0.849 ± 0.028 0.818 → 0.766 IMPROVED (1)
writing-clearly-and-concisely Softaworks · MIT 0.864 ± 0.016 → 0.841 ± 0.064 0.839 → 0.813 REGRESSED (1)
crafting-effective-readmes Softaworks · MIT 0.870 ± 0.015 → 0.804 ± 0.192 0.849 → 0.811 REGRESSED (2)
naming-analyzer Softaworks · MIT 0.817 ± 0.077 → 0.839 ± 0.026 0.703 → 0.757 MIXED
requesting-code-review Jesse Vincent (obra) · MIT 0.668 ± 0.249 → 0.719 ± 0.226 0.540 → 0.617 IMPROVED (1)
writing-plans Jesse Vincent (obra) · MIT 0.811 ± 0.089 → 0.785 ± 0.145 0.576 → 0.675 MIXED
skill-creator Anthropic, PBC · Apache-2.0 0.842 ± 0.095 → 0.875 ± 0.012 0.828 → 0.774 IMPROVED (1)

code-review-and-quality — Code Review and Quality WITHIN NOISE

A five-axis code-review framework and severity-labeling discipline for assessing changes before merge.

Source: https://github.com/addyosmani/agent-skills @ 7829ffd90d · License: MIT · Author: addyosmani
Receipts: old · new · drift

claude-sonnet-4-6 (old)claude-sonnet-5 (new)
with_skill (mean ± sd)0.869 ± 0.0250.854 ± 0.031
baseline (mean ± sd)0.824 ± 0.0810.836 ± 0.051
caseold (mean ± sd)new (mean ± sd)Δverdict
severity-labeled-findings0.908 ± 0.0420.794 ± 0.091-0.114within noise
security-axis-sql-injection0.882 ± 0.0080.858 ± 0.008-0.024within noise (below floor)
commit-message-imperative-body0.862 ± 0.0130.856 ± 0.009-0.006within noise
reject-clean-it-up-later0.866 ± 0.0150.868 ± 0.008+0.002within noise
oversized-change-split0.856 ± 0.0150.860 ± 0.010+0.004within noise
propose-structural-remedy0.878 ± 0.0190.896 ± 0.022+0.018within noise
approve-when-improves-health0.828 ± 0.0280.848 ± 0.026+0.020within noise

git-workflow-and-versioning — Git Workflow and Versioning MIXED

Trunk-based git discipline, atomic commits, conventional commit types, and semantic-versioning / changelog conventions.

Source: https://github.com/addyosmani/agent-skills @ 7829ffd90d · License: MIT · Author: addyosmani
Receipts: old · new · drift

claude-sonnet-4-6 (old)claude-sonnet-5 (new)
with_skill (mean ± sd)0.868 ± 0.0250.834 ± 0.104
baseline (mean ± sd)0.788 ± 0.1940.859 ± 0.024
caseold (mean ± sd)new (mean ± sd)Δverdict
commit-message-conventional-type0.856 ± 0.0130.600 ± 0.000-0.256regressed
release-cut-version-tag-changelog0.910 ± 0.0170.888 ± 0.011-0.022within noise
changelog-curated-by-impact0.888 ± 0.0290.876 ± 0.011-0.012within noise
split-into-atomic-commits0.868 ± 0.0080.858 ± 0.011-0.010within noise
semver-clean-bump0.860 ± 0.0120.856 ± 0.009-0.004within noise
trunk-based-short-lived-branches0.864 ± 0.0230.868 ± 0.013+0.004within noise
semver-hidden-breaking-change0.832 ± 0.0200.894 ± 0.022+0.062improved

documentation-and-adrs — Documentation and ADRs MIXED

When and how to write Architecture Decision Records, inline comments, and API docs; document the why.

Source: https://github.com/addyosmani/agent-skills @ 7829ffd90d · License: MIT · Author: addyosmani
Receipts: old · new · drift

claude-sonnet-4-6 (old)claude-sonnet-5 (new)
with_skill (mean ± sd)0.825 ± 0.1000.706 ± 0.297
baseline (mean ± sd)0.809 ± 0.1000.720 ± 0.293
caseold (mean ± sd)new (mean ± sd)Δverdict
match-existing-adr-convention0.858 ± 0.0160.076 ± 0.043-0.782regressed
surface-conflicting-adr-conventions0.850 ± 0.0250.582 ± 0.149-0.268regressed
supersede-not-delete-old-adr0.886 ± 0.0130.876 ± 0.005-0.010within noise
comment-intent-not-implementation0.860 ± 0.0140.864 ± 0.017+0.004within noise
document-why-not-what-rewrite0.860 ± 0.0100.868 ± 0.013+0.008within noise
adr-for-costly-to-reverse-decision0.864 ± 0.0180.876 ± 0.005+0.012within noise
document-public-api-function0.600 ± 0.0000.800 ± 0.000+0.200improved

commit-work — Commit Work IMPROVED (1)

Producing high-quality git commits: staging discipline, logical splitting, Conventional Commits messages.

Source: https://github.com/softaworks/agent-toolkit @ 3027f20f31 · License: MIT · Author: Softaworks
Receipts: old · new · drift

claude-sonnet-4-6 (old)claude-sonnet-5 (new)
with_skill (mean ± sd)0.825 ± 0.0820.849 ± 0.028
baseline (mean ± sd)0.818 ± 0.0630.766 ± 0.222
caseold (mean ± sd)new (mean ± sd)Δverdict
full-workflow-multi-concern-diff0.906 ± 0.0340.876 ± 0.011-0.030within noise
patch-stage-mixed-single-file0.872 ± 0.0080.860 ± 0.010-0.012within noise
review-cached-catch-secret-and-debug0.852 ± 0.0540.844 ± 0.092-0.008within noise
split-dependency-bump-vs-behavior0.872 ± 0.0040.866 ± 0.015-0.006within noise
split-feature-vs-refactor0.852 ± 0.0110.854 ± 0.011+0.002within noise
conventional-commit-single-change0.738 ± 0.0690.790 ± 0.041+0.052within noise
two-sentence-describability-test0.682 ± 0.0540.850 ± 0.023+0.168improved

writing-clearly-and-concisely — Writing Clearly and Concisely REGRESSED (1)

Strunk-style prose rules for docs, commit / error messages, reports, and UI text.

Source: https://github.com/softaworks/agent-toolkit @ 3027f20f31 · License: MIT · Author: Softaworks
Receipts: old · new · drift

claude-sonnet-4-6 (old)claude-sonnet-5 (new)
with_skill (mean ± sd)0.864 ± 0.0160.841 ± 0.064
baseline (mean ± sd)0.839 ± 0.0440.813 ± 0.100
caseold (mean ± sd)new (mean ± sd)Δverdict
concrete-language-incident-summary0.862 ± 0.0220.710 ± 0.073-0.152regressed
emphatic-word-at-end0.848 ± 0.0230.806 ± 0.013-0.042within noise (below floor)
tighten-puffy-release-note0.874 ± 0.0320.862 ± 0.022-0.012within noise
positive-form-status-sentences0.886 ± 0.0380.884 ± 0.015-0.002within noise
active-voice-logging-passage0.852 ± 0.0040.862 ± 0.011+0.010within noise
omit-needless-words-cache-note0.878 ± 0.0330.894 ± 0.039+0.016within noise
keep-related-words-together-modifiers0.846 ± 0.0110.872 ± 0.023+0.026within noise

crafting-effective-readmes — Crafting Effective READMEs REGRESSED (2)

Audience-first README structure with mandatory sections and project-type templates.

Source: https://github.com/softaworks/agent-toolkit @ 3027f20f31 · License: MIT · Author: Softaworks
Receipts: old · new · drift

claude-sonnet-4-6 (old)claude-sonnet-5 (new)
with_skill (mean ± sd)0.870 ± 0.0150.804 ± 0.192
baseline (mean ± sd)0.849 ± 0.0670.811 ± 0.144
caseold (mean ± sd)new (mean ± sd)Δverdict
lead-with-one-sentence-problem0.874 ± 0.0050.374 ± 0.136-0.500regressed
categorize-task-before-writing0.868 ± 0.0080.818 ± 0.034-0.050regressed
audience-config-folder-future-you0.876 ± 0.0150.870 ± 0.012-0.006within noise
review-validate-against-project-files0.888 ± 0.0290.890 ± 0.024+0.002within noise
three-mandatory-sections-cli-tool0.872 ± 0.0080.876 ± 0.013+0.004within noise
audience-internal-service-runbook0.872 ± 0.0040.894 ± 0.028+0.022within noise
oss-project-type-sections0.840 ± 0.0900.908 ± 0.041+0.068within noise

naming-analyzer — Naming Analyzer MIXED

Variable / function / class naming conventions and how to suggest clearer names.

Source: https://github.com/softaworks/agent-toolkit @ 3027f20f31 · License: MIT · Author: Softaworks
Receipts: old · new · drift

claude-sonnet-4-6 (old)claude-sonnet-5 (new)
with_skill (mean ± sd)0.817 ± 0.0770.839 ± 0.026
baseline (mean ± sd)0.703 ± 0.2140.757 ± 0.141
caseold (mean ± sd)new (mean ± sd)Δverdict
go-acronym-casing0.868 ± 0.0160.814 ± 0.022-0.054regressed
constants-include-units-js0.866 ± 0.0110.814 ± 0.022-0.052regressed
vague-names-js0.854 ± 0.0450.816 ± 0.042-0.038within noise
language-casing-python0.886 ± 0.0240.882 ± 0.013-0.004within noise
misleading-name-mutation-js0.828 ± 0.0080.842 ± 0.004+0.014within noise (below floor)
boolean-prefixes-js0.720 ± 0.1150.860 ± 0.010+0.140improved
abbreviations-wellknown-js0.696 ± 0.0990.846 ± 0.027+0.150improved

requesting-code-review — Requesting Code Review IMPROVED (1)

How to prepare and package a change for review: context, base/head SHAs, and severity triage of feedback.

Source: https://github.com/obra/superpowers @ 3dcbd5c4b4 · License: MIT · Author: Jesse Vincent (obra)
Receipts: old · new · drift

claude-sonnet-4-6 (old)claude-sonnet-5 (new)
with_skill (mean ± sd)0.668 ± 0.2490.719 ± 0.226
baseline (mean ± sd)0.540 ± 0.2620.617 ± 0.200
caseold (mean ± sd)new (mean ± sd)Δverdict
request-package-basic0.898 ± 0.0300.866 ± 0.021-0.032within noise
full-handoff-before-merge0.400 ± 0.0000.396 ± 0.009-0.004within noise
mandatory-vs-optional-triggers0.862 ± 0.0130.864 ± 0.013+0.002within noise
resist-simple-self-review0.864 ± 0.0090.882 ± 0.022+0.018within noise
identify-base-head-shas0.800 ± 0.0000.824 ± 0.018+0.024within noise (below floor)
crafted-context-not-session-history0.290 ± 0.0220.384 ± 0.188+0.094within noise
triage-review-findings0.560 ± 0.0290.820 ± 0.109+0.260improved

writing-plans — Writing Plans MIXED

How to author a concrete, self-contained implementation plan before coding.

Source: https://github.com/obra/superpowers @ 3dcbd5c4b4 · License: MIT · Author: Jesse Vincent (obra)
Receipts: old · new · drift

claude-sonnet-4-6 (old)claude-sonnet-5 (new)
with_skill (mean ± sd)0.811 ± 0.0890.785 ± 0.145
baseline (mean ± sd)0.576 ± 0.2300.675 ± 0.199
caseold (mean ± sd)new (mean ± sd)Δverdict
task-right-sizing-testable-deliverable0.816 ± 0.0940.460 ± 0.150-0.356regressed
self-review-spec-coverage0.882 ± 0.0110.844 ± 0.029-0.038within noise
full-small-plan-header-and-tasks0.882 ± 0.0300.850 ± 0.041-0.032within noise
bite-sized-tdd-steps0.864 ± 0.0130.876 ± 0.005+0.012within noise
file-structure-by-responsibility0.802 ± 0.0040.826 ± 0.037+0.024within noise
repair-placeholder-steps0.802 ± 0.0040.836 ± 0.035+0.034within noise
interfaces-exact-signatures0.626 ± 0.1500.804 ± 0.009+0.178improved

skill-creator — Skill Creator IMPROVED (1)

Conventions for authoring a SKILL.md: description-as-trigger, progressive disclosure, instructional tone.

Source: https://github.com/anthropics/skills @ b29e7cf65e · License: Apache-2.0 · Author: Anthropic, PBC
Receipts: old · new · drift

claude-sonnet-4-6 (old)claude-sonnet-5 (new)
with_skill (mean ± sd)0.842 ± 0.0950.875 ± 0.012
baseline (mean ± sd)0.828 ± 0.0700.774 ± 0.115
caseold (mean ± sd)new (mean ± sd)Δverdict
restructure-oversized-skill0.882 ± 0.0180.854 ± 0.032-0.028within noise
reframe-rigid-musts0.876 ± 0.0050.872 ± 0.004-0.004within noise
domain-variant-organization0.876 ± 0.0150.872 ± 0.008-0.004within noise
write-triggering-description0.874 ± 0.0130.874 ± 0.011+0.000within noise
critique-and-fix-skillmd0.890 ± 0.0350.894 ± 0.025+0.004within noise
output-format-template0.868 ± 0.0110.880 ± 0.012+0.012within noise
draft-full-skillmd0.628 ± 0.0740.882 ± 0.022+0.254improved

Methodology

Reproduce

node scripts/fetch-skills.js
node scripts/run-report-001.js --concurrency 5
node scripts/build-report-001.js

Amendments

v1.2 · 2026-07-30 — In-text grounding policy adopted and applied across all 10 suites. Headline: v1.1 read 3 regressed / 3 improved / 3 mixed / 1 within noise; v1.2 reads 2 regressed / 3 improved / 4 mixed / 1 within noise. This supersedes v1.1.

  • Policy adopted (codified, applied unilaterally). Every gradable rubric criterion must trace to text present in the skill's SKILL.md at the pinned SHA. Criteria that import domain knowledge the skill does not state are not gradable, regardless of how standard that knowledge is. Claims may summarize; rubrics may not extrapolate. The grading standard does not depend on the audited party's tolerance. Full sweep of all 10 suites (70 cases): reports/rubric-sweep.md.
  • Criteria amended (2 of 10 suites; the other 8 were grounded and unchanged). git-workflow-and-versioning/semver-clean-bump required "the highest-precedence bump governs when changes are combined / reset PATCH to 0" — a rule the SKILL.md never states (it defines only MAJOR/MINOR/PATCH). That criterion was removed; the case now grades only the stated MAJOR/MINOR/PATCH definitions. commit-work (all 7 cases) required a specific Conventional Commits type token (e.g. chore(deps) vs feat) and, in one case, a subject <=72-chars / no-trailing-period / imperative-mood rule — none of which its SKILL.md states (it names Conventional Commits and shows the type(scope): summary shape but does not enumerate the type vocabulary). Type grading is now shape-only; the length/period/mood rule was removed. The logical-split, patch-staging, and secret/debug checks the SKILL.md does state remain fully graded.
  • Re-run. Only the two amended suites were re-run (both models, n=5, claude-cli, transcripts retained); the amendment changes only those two suite_hashes. The grounding reference now recorded on every case is excluded from suite_hash, so annotating the other eight suites did not disturb their receipts.
  • Verdict changes (stated plainly, not softened). The amended case git-workflow-and-versioning/semver-clean-bump flips regression → within noise (v1.1 Δ−0.054 on separated bands; v1.2 Δ−0.004). The v1.1 "regression" was driven by the imported precedence/reset criterion; grading only the MAJOR/MINOR/PATCH definitions the SKILL.md actually states, the new model shows no drift on this case. But re-running a suite re-rolls all seven of its cases (one receipt covers the whole suite), and the fresh n=5 measurement surfaced, on two unamended, well-grounded cases, a real regression and a real improvement: commit-message-conventional-type now regresses (Δ−0.256 — the new model wrote feat: instead of fix: for a bug fix; all five judge samples scored 0.60 with that same reason, transcript-verified) and semver-hidden-breaking-change now improves (Δ+0.062). Net, the skill moves REGRESSED(1) → MIXED. commit-work is unchanged at IMPROVED(1): its driver (two-sentence-describability-test, Δ+0.168) held, and its amended cases (specific type token → shape, and the removed ≤72-char/no-period/imperative-mood rule) remain within noise. The other eight skills were not re-run and keep their v1.1 labels.

v1.1 · 2026-07-28 — Pre-launch QA corrections. Net headline change: v1.0 read 4 regressed / 3 improved / 3 mixed / 0 within noise; v1.1 reads 3 regressed / 3 improved / 3 mixed / 1 within noise. Two causes, below.

  • (1) Minimum-effect floor of 0.05. A case is now called regressed/improved only when its two with_skill bands do not overlap AND the mean moved ≥ 0.05 (one judge-quantization step) — band separation alone is no longer sufficient, because the judge quantizes to a coarse grid and a confident grade often collapses to a zero-width point band. This reclassified 3 hair-trigger cases (|Δ| 0.014–0.024) to “within noise (below effect floor)”: code-review-and-quality/security-axis-sql-injection (which flips that whole skill REGRESSED → WITHIN NOISE), naming-analyzer/misleading-name-mutation-js (MIXED 2r/3i → 2r/2i), and requesting-code-review/identify-base-head-shas (IMPROVED 2 → 1).
  • (2) Two artifact-contaminated receipts re-run. A transcript audit found new-model bands containing a judge parse-fallback score of exactly 0.0 (an unparseable judge reply defaults to 0, per lib/judge.js). Both __claude-sonnet-5 receipts were re-run on the same claude-cli surface, n=5 (old-model receipts, which were clean, were kept as-is): writing-clearly-and-concisely and documentation-and-adrs.
  • What the re-run showed. active-voice-logging-passage’s v1.0 “regression” was entirely the artifact — the clean re-run scores 0.862 (vs old 0.852), i.e. within noise. match-existing-adr-convention reconfirmed as a genuine new-model failure (~0.08 again on an independent run). Because a receipt covers all 7 of a skill’s cases, re-running also refreshed the other cases honestly: comment-intent-not-implementation’s v1.0 regression did not reproduce (now within noise), while two different, real regressions surfaced — documentation-and-adrs/surface-conflicting-adr-conventions (Δ−0.268) and writing-clearly-and-concisely/concrete-language-incident-summary (Δ−0.152). Both skills keep their v1.0 labels (MIXED and REGRESSED-1 respectively); only the driving cases changed.
  • Full per-case transcript evidence and REAL-vs-artifact classifications are in reports/report-001-audit.md.

v1.3 (2026-09-14): the entry below corrects how this report’s band wording is read.

v1.3 · 2026-09-14. This entry corrects how the report’s band wording is read. It follows the entries recorded before it, no earlier text in the report is changed, and no figure, verdict token, table value or receipt reference changes.

  • The headline. This page says “9 of 10 skills moved beyond noise on this model pair.” Each of those nine skills has at least one case whose two bands do not overlap and whose mean moved by at least the 0.05 effect floor: a separation detected under the rule, which is not proof that the skill’s behaviour moved. No skill’s aggregate with_skill band separated. As amended:

9 of 10 skills showed a separation detected under the rule on this model pair.

  • How to read this page. A case is separated under the rule when its two bands do not overlap and its mean moved by at least the 0.05 effect floor; bands that do not overlap by a smaller move are not separated. The four rows marked within noise (below floor) are of that kind, and what the headline rule calls bands that separate by less than the floor are bands that do not overlap. Wherever the page reads a result that was not separated as settled, as within noise, never as drift, no drift on this case, hair-trigger, a regression that did not reproduce, or a cause confirmed because a later run was not separated, as the entry of 2026-07-30 does of the imported criterion and the entry of 2026-07-28 of the artifact, read it as no separation detected at the sample size used, which is not evidence of equivalence, not evidence that nothing changed, and not proof of the cause. Wherever it reads a separation as settled, as a real regression and a real improvement, real regressions surfaced, reconfirmed as a genuine new-model failure, or a driver that held, read it as a separation detected under the rule, which is not proof that the model’s behaviour moved. The Methodology’s no meaningful change is a hypothetical that explains the floor, and the footer’s Verdicts age because the substrate moves names a cause that no separation here proves. Labels, tallies and headings, the summary card’s What moved among them, are read the same way and are left as published.
  • The bands. Each band on this page is a mean plus or minus one sample standard deviation, of two kinds: a case’s band, over its five judge samples, and a skill’s with_skill and baseline bands, over its seven per-case means, which is suite dispersion. Each band is a descriptive spread, not a confidence interval, with no coverage probability.
  • Where the amended headline appears. The site’s report index and feed quote the amended headline; this page’s summary card, description and social card image keep the headline as published. Filed under the wording rules of the repository's spec 031, amendment A-031-20.

v1.4 · 2026-10-03. This entry adds a dated note under the headline; no earlier text is changed, and no figure, verdict token, table value or receipt reference changes. The note says the headline counts on this page include receipts whose receipt pages mark them Not measured, and that corrected counts follow in spec 159. It was written after each of the 20 receipts this page links was read with the badge's own rule, and all 20 read Not measured.

Receipts validate against the Driftproof receipt spec (v0.2; the two suites re-run in v1.2 are v0.3, a superset the loader also accepts). Every verdict above is a function of the two with_skill bands recorded in the receipts. Verdicts age because the substrate moves — that is the phenomenon this report measures.