Driftproof

Report 011: Claude Opus 5.5 on release day, three skills

Release drift report. The first table is the type’s question: the model moves under the skill and the harness stays fixed. The second holds the model and moves the harness, which the type does not describe; the limits say how it is read.

Claude Opus 5.5 on the three skills and cases of Report 009, one case each, beside a fresh Claude Opus 5 arm on the same harness. Read by the runner’s own comparison of the with-skill arms, code-review-and-quality reads no separation detected, git-workflow-and-versioning reads no separation detected; not enough draws to conclude at this effect floor, documentation-and-adrs reads no separation detected; not enough draws to conclude at this effect floor. Separately, Claude Opus 5 re-run on the newer Claude Code against Report 009’s own receipts: code-review-and-quality reads no separation detected, git-workflow-and-versioning reads no separation detected, documentation-and-adrs reads no separation detected; not enough draws to conclude at this effect floor.

Every figure below is read from a receipt or a file published beside it, and each block names its files. The target model of the new arm is claude-opus-5-5, and every draw of every arm was judged by claude-opus-5.

Read from: docs/reports/011/evidence/code-review-and-quality-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/code-review-and-quality-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/diff--model--code-review-and-quality.md; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/diff--model--git-workflow-and-versioning.md; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/diff--model--documentation-and-adrs.md; docs/reports/009/evidence/three-skill-comparison--driftproof--code-review-and-quality--receipts--code-review-and-quality-claude-opus-5-2026-09-15.json; docs/reports/011/evidence/diff--harness--code-review-and-quality.md; docs/reports/009/evidence/three-skill-comparison--driftproof--git-workflow-and-versioning--receipts--git-workflow-and-versioning-claude-opus-5-2026-09-15.json; docs/reports/011/evidence/diff--harness--git-workflow-and-versioning.md; docs/reports/009/evidence/three-skill-comparison--driftproof--documentation-and-adrs--receipts--documentation-and-adrs-claude-opus-5-2026-09-15.json; docs/reports/011/evidence/diff--harness--documentation-and-adrs.md.

Limits, read these first

Read from: docs/reports/011/evidence/code-review-and-quality-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/code-review-and-quality-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-2026-09-23.json; three-skill-comparison/version-preflight.json (in Report 009’s bundle, not published; its sha256 is in docs/reports/009/evidence/three-skill-comparison--SHA256SUMS); docs/reports/011/evidence/run-20260923T062806Z--run-record.json; docs/reports/011/evidence/code-review-and-quality-claude-opus-5-5-2026-09-23.surface.json; driftproof-calls/001/command.json (not published: a raw call record; its sha256 is in docs/reports/011/evidence/run-20260923T062806Z--driftproof-calls.SHA256SUMS); docs/reports/011/evidence/run-20260923T062806Z--driftproof-calls.SHA256SUMS; docs/reports/011/evidence/guard.py.

What stayed fixed and what moved

The inputs are Report 009’s. The same SKILL.md bytes (skill.content_hash, the same in this run’s two receipts and Report 009’s for each skill: code-review-and-quality 13d360d7f786, git-workflow-and-versioning 91c8c72654ee, documentation-and-adrs b67a9f07ed10) and the same suites (suite.suite_hash: code-review-and-quality 5d729f885294, git-workflow-and-versioning 4e74150753d0, documentation-and-adrs 6cce54e7d9d0).

The runner is Report 009’s. Driftproof runner 0.10.1 on the claude-cli surface, with the flags Report 009’s run used: at most 80 calls and 8.00 dollars of estimated spend per skill run. It puts the SKILL.md text in the prompt for the with-skill arm and gives the baseline arm the task alone, with no tools in either. It draws generations per arm until the spread settles or a maximum is reached, and the judge scores each draw against the case’s rubric on a continuous 0 to 1 scale.

The judge is Report 009’s. claude-opus-5, with the grading template 82586d1e44f8 in all three.

The rule. A case’s band is its mean across draws plus or minus the sample standard deviation across draws: a descriptive spread with no coverage probability. Two with-skill bands separate under the rule when they do not overlap and their means differ by at least the 0.05 effect floor. A case that does not separate is read under spec 035’s rule: when the two arms’ spreads and draws could not have resolved a shift of the floor’s size, it reads not enough draws to conclude at this effect floor, and otherwise no separation detected. Neither is evidence that nothing differs. The verdicts are the runner’s: driftproof diff over each pair, whose outputs are published below.

Read from: docs/reports/011/evidence/code-review-and-quality-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/code-review-and-quality-claude-opus-5-2026-09-23.json; docs/reports/009/evidence/three-skill-comparison--driftproof--code-review-and-quality--receipts--code-review-and-quality-claude-opus-5-2026-09-15.json; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-2026-09-23.json; docs/reports/009/evidence/three-skill-comparison--driftproof--git-workflow-and-versioning--receipts--git-workflow-and-versioning-claude-opus-5-2026-09-15.json; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-2026-09-23.json; docs/reports/009/evidence/three-skill-comparison--driftproof--documentation-and-adrs--receipts--documentation-and-adrs-claude-opus-5-2026-09-15.json; docs/reports/011/evidence/code-review-and-quality-claude-opus-5-5-2026-09-23.stdout.txt; config.js.

The low figures, read before the verdicts

Three figures on this run read low, and a low figure can be an instrument defect rather than an answer: a truncated generation, a refusal, a tool error, a timeout or a judge reply that did not parse. Report 007’s timeout defect is the precedent. Each draw behind these three figures was read, in the receipt and in the raw call record, before any verdict was written. None is an artefact, and no cell is excluded. The review is published as run-20260923T062806Z--artefact-review-20260923.md.

One more draw, outside the three. In Claude Opus 5’s baseline on git-workflow-and-versioning, draw three of ten scored 0.00. The surface gives the model no tools, and the model wrote a Glob invocation out as text, followed by the words No files found., and gave no commit message. The judge scored that absence. It ended its turn, was not truncated, ran no tool and parsed, so it is the model’s own output and it is not excluded. It sits in a baseline arm, which no verdict here reads.

Read from: docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-5-2026-09-23.json; driftproof-calls/041/stdout.jsonl (not published: a raw call record; its sha256 is in docs/reports/011/evidence/run-20260923T062806Z--driftproof-calls.SHA256SUMS); driftproof-calls/042/prompt.txt (not published: a raw call record; its sha256 is in docs/reports/011/evidence/run-20260923T062806Z--driftproof-calls.SHA256SUMS); docs/reports/011/evidence/code-review-and-quality-claude-opus-5-5-2026-09-23.json; driftproof-calls/017/stdout.jsonl (not published: a raw call record; its sha256 is in docs/reports/011/evidence/run-20260923T062806Z--driftproof-calls.SHA256SUMS); driftproof-calls/019/stdout.jsonl (not published: a raw call record; its sha256 is in docs/reports/011/evidence/run-20260923T062806Z--driftproof-calls.SHA256SUMS); docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-5-2026-09-23.json; driftproof-calls/056/stdout.jsonl (not published: a raw call record; its sha256 is in docs/reports/011/evidence/run-20260923T062806Z--driftproof-calls.SHA256SUMS); docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-2026-09-23.json; driftproof-calls/133/stdout.jsonl (not published: a raw call record; its sha256 is in docs/reports/011/evidence/run-20260923T062806Z--driftproof-calls.SHA256SUMS).

The model: Claude Opus 5.5 against Claude Opus 5, both on this run’s harness

skillClaude Opus 5, with skill, mean ± sd across drawsClaude Opus 5.5, with skill, mean ± sd across drawswith-skill deltaunder the rulebaseline mean (Claude Opus 5; Claude Opus 5.5), contextlift (Claude Opus 5; Claude Opus 5.5), context
code-review-and-quality0.901 ± 0.007 (3 draws)0.910 ± 0.015 (3 draws)+0.009no separation detected0.862; 0.680+0.039; +0.230
git-workflow-and-versioning0.861 ± 0.008 (3 draws)0.838 ± 0.027 (3 draws)-0.023no separation detected; not enough draws to conclude at this effect floor0.765; 0.300+0.096; +0.538
documentation-and-adrs0.804 ± 0.103 (7 draws)0.661 ± 0.148 (6 draws)-0.143no separation detected; not enough draws to conclude at this effect floor0.567; 0.567+0.237; +0.094

The with-skill delta is the second with-skill mean minus the first. The rule column is the runner’s comparison of the two with-skill bands; the delta, the baseline means and the lifts are shown as context, and no verdict is read from them. A lift is a receipt’s with-skill mean minus its baseline mean.

Read from: docs/reports/011/evidence/code-review-and-quality-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/code-review-and-quality-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/diff--model--code-review-and-quality.md; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/diff--model--git-workflow-and-versioning.md; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/diff--model--documentation-and-adrs.md.

The harness: Claude Opus 5 in Report 009 against Claude Opus 5 in this run

skillReport 009, with skill, mean ± sd across drawsthis run, with skill, mean ± sd across drawswith-skill deltaunder the rulebaseline mean (Report 009; this run), contextlift (Report 009; this run), context
code-review-and-quality0.918 ± 0.010 (3 draws)0.901 ± 0.007 (3 draws)-0.017no separation detected0.783; 0.862+0.134; +0.039
git-workflow-and-versioning0.862 ± 0.007 (3 draws)0.861 ± 0.008 (3 draws)-0.001no separation detected0.300; 0.765+0.562; +0.096
documentation-and-adrs0.822 ± 0.017 (3 draws)0.804 ± 0.103 (7 draws)-0.018no separation detected; not enough draws to conclude at this effect floor0.532; 0.567+0.291; +0.237

The same columns, with Report 009’s receipt first. Report 009’s receipts are published with that report and linked here, not copied.

Read from: docs/reports/009/evidence/three-skill-comparison--driftproof--code-review-and-quality--receipts--code-review-and-quality-claude-opus-5-2026-09-15.json; docs/reports/011/evidence/code-review-and-quality-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/diff--harness--code-review-and-quality.md; docs/reports/009/evidence/three-skill-comparison--driftproof--git-workflow-and-versioning--receipts--git-workflow-and-versioning-claude-opus-5-2026-09-15.json; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/diff--harness--git-workflow-and-versioning.md; docs/reports/009/evidence/three-skill-comparison--driftproof--documentation-and-adrs--receipts--documentation-and-adrs-claude-opus-5-2026-09-15.json; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/diff--harness--documentation-and-adrs.md.

Three readings

1. Between the two models, the rule detects no separation on any of the three skills, and on two of them the draws could not have told. On code-review-and-quality the with-skill means differ by +0.009 and the rule reads no separation detected. On git-workflow-and-versioning they differ by -0.023; at these spreads the runner puts the draws that would have been needed at 16 per arm, against the 3 drawn. On documentation-and-adrs the move is the largest of the three, -0.143, and the two spreads sum to 0.251, which is at or above the floor, so no draw count at these spreads resolves it. None of this is evidence that the two models score these skills alike.

2. Claude Opus 5 on this run’s harness separates from its Report 009 receipts on no skill. The with-skill means move by -0.017 on code-review-and-quality, -0.001 on git-workflow-and-versioning, -0.018 on documentation-and-adrs. The rule reads no separation detected on the first two, and on documentation-and-adrs, where this run’s spread is 0.103 against Report 009’s 0.017, no separation detected; not enough draws to conclude at this effect floor. This is a statement about these draws on two harnesses and two hosts, not about Claude Code.

3. On two skills the two models’ baselines sit further apart than their with-skill scores, and no verdict here reads them. Without the skill, git-workflow-and-versioning reads 0.765 for Claude Opus 5 and 0.300 for Claude Opus 5.5, and code-review-and-quality 0.862 and 0.680. On those two skills the lifts therefore differ more than the with-skill scores do; on documentation-and-adrs the two baselines read 0.567 and 0.567. The runner’s comparison of two receipts does not read baseline arms, so this is offered as an observation to test on more cases, not as a finding; the section on the low figures says what those draws are.

Read from: docs/reports/011/evidence/code-review-and-quality-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/code-review-and-quality-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/diff--model--code-review-and-quality.md; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-5-2026-09-23.json; docs/reports/009/evidence/three-skill-comparison--driftproof--code-review-and-quality--receipts--code-review-and-quality-claude-opus-5-2026-09-15.json; docs/reports/009/evidence/three-skill-comparison--driftproof--git-workflow-and-versioning--receipts--git-workflow-and-versioning-claude-opus-5-2026-09-15.json; docs/reports/009/evidence/three-skill-comparison--driftproof--documentation-and-adrs--receipts--documentation-and-adrs-claude-opus-5-2026-09-15.json; docs/reports/011/evidence/diff--harness--documentation-and-adrs.md.

Run record

The run. Run 20260923T062806Z, one skill at a time, the Claude Opus 5.5 arm first and then the Claude Opus 5 arm, code-review-and-quality first in each. The command file writes a status line after each skill run; there are six, each with exit status 0 and one receipt, the last at 2026-09-23T07:08:39Z. Each receipt’s run.date_utc is written by the runner and is not read here as a start or a finish.

The caps. Per skill run, the runner’s own caps, as in Report 009 (see Setup). Across the whole run, a guard in front of Claude Code allows 240 calls in all and 240 seconds per call. The run used 204 calls, and the six receipts record 0 unmeasured draws.

The harness. Claude Code 2.1.280, installed at the npm integrity the run record states, with runner 0.10.1 at Report 009’s integrity. The runner writes no Claude Code version into a receipt, and a receipt is sealed, so each receipt has a sidecar, <receipt>.surface.json, bound to it by its receipt_hash and recording the version each of its calls reported; the table gives it.

The price. The runner’s registry had no row for claude-opus-5-5, so this run used a copy of it with one row added, at 4 and 20 dollars per million input and output tokens, read from the vendor’s pricing page on the day of the run (docs-pricing-snapshot-2026-09-23.json). The row is this run’s only; the product’s registry does not carry it.

The cost. Metered spend was 0.00 dollars: the claude-cli surface runs on a subscription. The estimated API-equivalent below is every draw’s generation and judge token counts, as each receipt records them, at the prices each receipt froze, with cached input counted at the full input price, as the runner counts it.

armskillcallsClaude Code reported by every callgeneration, estimated USDjudge, estimated USDtotal, estimated USD
claude-opus-5-5code-review-and-quality282.1.2800.421.171.58
claude-opus-5-5git-workflow-and-versioning242.1.2800.190.630.82
claude-opus-5-5documentation-and-adrs362.1.2800.361.291.65
claude-opus-5code-review-and-quality242.1.2800.670.961.63
claude-opus-5git-workflow-and-versioning522.1.2800.871.492.36
claude-opus-5documentation-and-adrs402.1.2800.751.432.17
all six3.256.9710.21

Integrity. Receipt hashes: claude-opus-5-5 code-review-and-quality 5c3ad19b40dd4b39, claude-opus-5-5 git-workflow-and-versioning ca69c3436d0dbfa5, claude-opus-5-5 documentation-and-adrs b924a435c70e0092, claude-opus-5 code-review-and-quality d7bd2fcd3dbea881, claude-opus-5 git-workflow-and-versioning 2ba4ef175d7f1caf, claude-opus-5 documentation-and-adrs 34107e18c90f3b20. Each validates with its receipt hash verified. The raw call records carry the SKILL.md text and stay unpublished; their sha256 are published.

Read from: docs/reports/011/evidence/run-20260923T062806Z--run-record.json; docs/reports/011/evidence/run-20260923T062806Z--status.jsonl; docs/reports/011/evidence/guard.py; docs/reports/011/evidence/code-review-and-quality-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/code-review-and-quality-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/code-review-and-quality-claude-opus-5-5-2026-09-23.surface.json; docs/reports/011/evidence/code-review-and-quality-claude-opus-5-5-2026-09-23.stdout.txt; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-5-2026-09-23.surface.json; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-5-2026-09-23.surface.json; docs/reports/011/evidence/code-review-and-quality-claude-opus-5-2026-09-23.surface.json; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-2026-09-23.surface.json; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-2026-09-23.surface.json; docs/reports/011/evidence/run-20260923T062806Z--driftproof-calls.SHA256SUMS.

Published evidence

The files this report makes public. The six receipts with their summaries, surface sidecars and the runner’s console output; the six driftproof diff outputs the verdicts are read from; the run record, its registry copy, the status lines, the review of the low figures and the sha256 of every raw call record; the pricing snapshot; and the call guard.

Report 009’s three receipts, the other side of the second table, are published with that report: code-review-and-quality, git-workflow-and-versioning, documentation-and-adrs. Validate a receipt with npx driftproof validate <file>. Each copy here is byte-identical to its twin under specs/042-opus-5-5-release-receipts/, and each sidecar names the receipt it belongs to and that receipt’s hash.