Report 011: Claude Opus 5.5 on release day, three skills
Release drift report. The first table is the type’s question: the model moves under the skill and the harness stays fixed. The second holds the model and moves the harness, which the type does not describe; the limits say how it is read.
Claude Opus 5.5 on the three skills and cases of Report 009, one case each, beside a fresh Claude Opus 5 arm on the same harness. Read by the runner’s own comparison of the with-skill arms, code-review-and-quality reads no separation detected, git-workflow-and-versioning reads no separation detected; not enough draws to conclude at this effect floor, documentation-and-adrs reads no separation detected; not enough draws to conclude at this effect floor. Separately, Claude Opus 5 re-run on the newer Claude Code against Report 009’s own receipts: code-review-and-quality reads no separation detected, git-workflow-and-versioning reads no separation detected, documentation-and-adrs reads no separation detected; not enough draws to conclude at this effect floor.
Every figure below is read from a receipt or a file published beside it, and each block names its files. The target model of the new arm is claude-opus-5-5, and every draw of every arm was judged by claude-opus-5.
Read from: docs/reports/011/evidence/code-review-and-quality-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/code-review-and-quality-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/diff--model--code-review-and-quality.md; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/diff--model--git-workflow-and-versioning.md; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/diff--model--documentation-and-adrs.md; docs/reports/009/evidence/three-skill-comparison--driftproof--code-review-and-quality--receipts--code-review-and-quality-claude-opus-5-2026-09-15.json; docs/reports/011/evidence/diff--harness--code-review-and-quality.md; docs/reports/009/evidence/three-skill-comparison--driftproof--git-workflow-and-versioning--receipts--git-workflow-and-versioning-claude-opus-5-2026-09-15.json; docs/reports/011/evidence/diff--harness--git-workflow-and-versioning.md; docs/reports/009/evidence/three-skill-comparison--driftproof--documentation-and-adrs--receipts--documentation-and-adrs-claude-opus-5-2026-09-15.json; docs/reports/011/evidence/diff--harness--documentation-and-adrs.md.
Limits, read these first
- One case per skill. Each receipt carries a suite of 1, 1, 1 case, in the order
code-review-and-quality,git-workflow-and-versioning,documentation-and-adrs, the cases Report 009 used. Nothing on this page describes how any of the three skills behaves on any other task, or how Claude Opus 5.5 behaves in general. - Few draws. Each arm drew between 3 and 10 generations, each judged 3 times, and each case’s band is the spread across its draws. With one case there is no spread across cases, and each receipt records its combined uncertainty on the lift as absent,
single_case. - Not Report 009’s harness, for either arm. Report 009 ran on Claude Code
2.1.272. That version cannot call this model: an earlier start of this run stopped at its first call with the API’s answerversion 2.1.280 or newer is required
, and produced no receipt. Both arms of this run therefore used Claude Code2.1.280, and each receipt’s sidecar records the version every one of its calls reported (see Run record). The first table holds that harness fixed and moves the model. It is not Report 009’s harness, so neither table puts a Claude Opus 5.5 figure beside a Report 009 figure. - The harness comparison moves more than Claude Code. The second table sets Report 009’s Claude Opus 5 receipts beside this run’s Claude Opus 5 receipts. Between them the Claude Code version changed, and so did the host: Report 009’s Claude Code binary was the
darwin-arm64build, and this run’s thelinux-x64build. So did the date, and every generation is a fresh draw. The runner, the judge, the grading template, the SKILL.md bytes, the cases, the draw rule and the caps are the same (see Setup). A separation in that table could not be put down to the Claude Code version alone. - The judge is one of the two models compared. Every draw in both tables was judged by
claude-opus-5, Report 009’s judge, which is also the model of the older arm. This report does not measure whether that judge scores its own model’s text differently from another model’s. It is also not the judge the judge policy fixes for Driftproof’s reports, the departure Report 009 made, so these figures are comparable with Report 009 and not with Reports 001 to 008. - No reasoning effort was set. None of the 204 raw call records’ argv carries an effort setting, the guard adds none (guard.py), and no receipt records one (none recorded). Whatever each model’s default was on this surface applied. The tables read as the same surface with a different model, not as the same effort.
- A verdict reads the with-skill arms only. The runner’s comparison of two receipts sets one receipt’s with-skill band against the other’s, per case. The baseline and lift columns are context: no verdict on this page is read from them, however far apart they are.
- Verification levels. All six receipts of this run are
TESTED,TESTED,TESTED,TESTED,TESTED,TESTED, each answered by a model the surface attested (answered_by.kindmodel,model,model,model,model,model).
Read from: docs/reports/011/evidence/code-review-and-quality-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/code-review-and-quality-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-2026-09-23.json; three-skill-comparison/version-preflight.json (in Report 009’s bundle, not published; its sha256 is in docs/reports/009/evidence/three-skill-comparison--SHA256SUMS); docs/reports/011/evidence/run-20260923T062806Z--run-record.json; docs/reports/011/evidence/code-review-and-quality-claude-opus-5-5-2026-09-23.surface.json; driftproof-calls/001/command.json (not published: a raw call record; its sha256 is in docs/reports/011/evidence/run-20260923T062806Z--driftproof-calls.SHA256SUMS); docs/reports/011/evidence/run-20260923T062806Z--driftproof-calls.SHA256SUMS; docs/reports/011/evidence/guard.py.
What stayed fixed and what moved
The inputs are Report 009’s. The same SKILL.md bytes (skill.content_hash, the same in this run’s two receipts and Report 009’s for each skill: code-review-and-quality 13d360d7f786, git-workflow-and-versioning 91c8c72654ee, documentation-and-adrs b67a9f07ed10) and the same suites (suite.suite_hash: code-review-and-quality 5d729f885294, git-workflow-and-versioning 4e74150753d0, documentation-and-adrs 6cce54e7d9d0).
The runner is Report 009’s. Driftproof runner 0.10.1 on the claude-cli surface, with the flags Report 009’s run used: at most 80 calls and 8.00 dollars of estimated spend per skill run. It puts the SKILL.md text in the prompt for the with-skill arm and gives the baseline arm the task alone, with no tools in either. It draws generations per arm until the spread settles or a maximum is reached, and the judge scores each draw against the case’s rubric on a continuous 0 to 1 scale.
The judge is Report 009’s. claude-opus-5, with the grading template 82586d1e44f8 in all three.
The rule. A case’s band is its mean across draws plus or minus the sample standard deviation across draws: a descriptive spread with no coverage probability. Two with-skill bands separate under the rule when they do not overlap and their means differ by at least the 0.05 effect floor. A case that does not separate is read under spec 035’s rule: when the two arms’ spreads and draws could not have resolved a shift of the floor’s size, it reads not enough draws to conclude at this effect floor
, and otherwise no separation detected
. Neither is evidence that nothing differs. The verdicts are the runner’s: driftproof diff over each pair, whose outputs are published below.
Read from: docs/reports/011/evidence/code-review-and-quality-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/code-review-and-quality-claude-opus-5-2026-09-23.json; docs/reports/009/evidence/three-skill-comparison--driftproof--code-review-and-quality--receipts--code-review-and-quality-claude-opus-5-2026-09-15.json; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-2026-09-23.json; docs/reports/009/evidence/three-skill-comparison--driftproof--git-workflow-and-versioning--receipts--git-workflow-and-versioning-claude-opus-5-2026-09-15.json; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-2026-09-23.json; docs/reports/009/evidence/three-skill-comparison--driftproof--documentation-and-adrs--receipts--documentation-and-adrs-claude-opus-5-2026-09-15.json; docs/reports/011/evidence/code-review-and-quality-claude-opus-5-5-2026-09-23.stdout.txt; config.js.
The low figures, read before the verdicts
Three figures on this run read low, and a low figure can be an instrument defect rather than an answer: a truncated generation, a refusal, a tool error, a timeout or a judge reply that did not parse. Report 007’s timeout defect is the precedent. Each draw behind these three figures was read, in the receipt and in the raw call record, before any verdict was written. None is an artefact, and no cell is excluded. The review is published as run-20260923T062806Z--artefact-review-20260923.md.
- Claude Opus 5.5 without the skill on
git-workflow-and-versioning, 0.300. Draws that ended their turn, untruncated and measured: three of three. Every judge sample reads 0.3. Each draw is a complete commit message whose subject line beginsFix password-reset linkswith no conventional type prefix, and the rubric readsif the subject line does not use a conventional type prefix, cap at 0.3
. The judge’s reason:Body strongly explains cause and user-facing impact, and subject is specific, but subject lacks the conventional `fix:` type prefix, triggering the 0.3 cap.
A real answer, judged low. - Claude Opus 5.5 without the skill on
code-review-and-quality, 0.680. Draws that ended their turn, untruncated and measured: four of four, with draw means 0.860, 0.483, 0.500, 0.877. The two low draws are complete reviews that label the leaked key and the logged customer dataCriticaland put the other findings under category headings,Correctnessamong them, with no severity label; the judge says so, for exampleare categories, not severity labels
. A real answer, judged low, with the spread between draws. - Claude Opus 5.5 with the skill on
documentation-and-adrs, 0.661. Draws that ended their turn, untruncated and measured: six of six, with draw means 0.483, 0.800, 0.600, 0.853, 0.697, 0.533. Each is a complete revised function. The low draws keep the expired-couponTODOreworded rather than resolving or removing it, which the judge reads as a miss, for exampleleft a reworded but still-bare TODO
. A real answer, judged low.
One more draw, outside the three. In Claude Opus 5’s baseline on git-workflow-and-versioning, draw three of ten scored 0.00. The surface gives the model no tools, and the model wrote a Glob invocation out as text, followed by the words No files found.
, and gave no commit message. The judge scored that absence. It ended its turn, was not truncated, ran no tool and parsed, so it is the model’s own output and it is not excluded. It sits in a baseline arm, which no verdict here reads.
Read from: docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-5-2026-09-23.json; driftproof-calls/041/stdout.jsonl (not published: a raw call record; its sha256 is in docs/reports/011/evidence/run-20260923T062806Z--driftproof-calls.SHA256SUMS); driftproof-calls/042/prompt.txt (not published: a raw call record; its sha256 is in docs/reports/011/evidence/run-20260923T062806Z--driftproof-calls.SHA256SUMS); docs/reports/011/evidence/code-review-and-quality-claude-opus-5-5-2026-09-23.json; driftproof-calls/017/stdout.jsonl (not published: a raw call record; its sha256 is in docs/reports/011/evidence/run-20260923T062806Z--driftproof-calls.SHA256SUMS); driftproof-calls/019/stdout.jsonl (not published: a raw call record; its sha256 is in docs/reports/011/evidence/run-20260923T062806Z--driftproof-calls.SHA256SUMS); docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-5-2026-09-23.json; driftproof-calls/056/stdout.jsonl (not published: a raw call record; its sha256 is in docs/reports/011/evidence/run-20260923T062806Z--driftproof-calls.SHA256SUMS); docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-2026-09-23.json; driftproof-calls/133/stdout.jsonl (not published: a raw call record; its sha256 is in docs/reports/011/evidence/run-20260923T062806Z--driftproof-calls.SHA256SUMS).
The model: Claude Opus 5.5 against Claude Opus 5, both on this run’s harness
| skill | Claude Opus 5, with skill, mean ± sd across draws | Claude Opus 5.5, with skill, mean ± sd across draws | with-skill delta | under the rule | baseline mean (Claude Opus 5; Claude Opus 5.5), context | lift (Claude Opus 5; Claude Opus 5.5), context |
|---|---|---|---|---|---|---|
code-review-and-quality | 0.901 ± 0.007 (3 draws) | 0.910 ± 0.015 (3 draws) | +0.009 | no separation detected | 0.862; 0.680 | +0.039; +0.230 |
git-workflow-and-versioning | 0.861 ± 0.008 (3 draws) | 0.838 ± 0.027 (3 draws) | -0.023 | no separation detected; not enough draws to conclude at this effect floor | 0.765; 0.300 | +0.096; +0.538 |
documentation-and-adrs | 0.804 ± 0.103 (7 draws) | 0.661 ± 0.148 (6 draws) | -0.143 | no separation detected; not enough draws to conclude at this effect floor | 0.567; 0.567 | +0.237; +0.094 |
The with-skill delta is the second with-skill mean minus the first. The rule column is the runner’s comparison of the two with-skill bands; the delta, the baseline means and the lifts are shown as context, and no verdict is read from them. A lift is a receipt’s with-skill mean minus its baseline mean.
Read from: docs/reports/011/evidence/code-review-and-quality-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/code-review-and-quality-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/diff--model--code-review-and-quality.md; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/diff--model--git-workflow-and-versioning.md; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/diff--model--documentation-and-adrs.md.
The harness: Claude Opus 5 in Report 009 against Claude Opus 5 in this run
| skill | Report 009, with skill, mean ± sd across draws | this run, with skill, mean ± sd across draws | with-skill delta | under the rule | baseline mean (Report 009; this run), context | lift (Report 009; this run), context |
|---|---|---|---|---|---|---|
code-review-and-quality | 0.918 ± 0.010 (3 draws) | 0.901 ± 0.007 (3 draws) | -0.017 | no separation detected | 0.783; 0.862 | +0.134; +0.039 |
git-workflow-and-versioning | 0.862 ± 0.007 (3 draws) | 0.861 ± 0.008 (3 draws) | -0.001 | no separation detected | 0.300; 0.765 | +0.562; +0.096 |
documentation-and-adrs | 0.822 ± 0.017 (3 draws) | 0.804 ± 0.103 (7 draws) | -0.018 | no separation detected; not enough draws to conclude at this effect floor | 0.532; 0.567 | +0.291; +0.237 |
The same columns, with Report 009’s receipt first. Report 009’s receipts are published with that report and linked here, not copied.
Read from: docs/reports/009/evidence/three-skill-comparison--driftproof--code-review-and-quality--receipts--code-review-and-quality-claude-opus-5-2026-09-15.json; docs/reports/011/evidence/code-review-and-quality-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/diff--harness--code-review-and-quality.md; docs/reports/009/evidence/three-skill-comparison--driftproof--git-workflow-and-versioning--receipts--git-workflow-and-versioning-claude-opus-5-2026-09-15.json; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/diff--harness--git-workflow-and-versioning.md; docs/reports/009/evidence/three-skill-comparison--driftproof--documentation-and-adrs--receipts--documentation-and-adrs-claude-opus-5-2026-09-15.json; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/diff--harness--documentation-and-adrs.md.
Three readings
1. Between the two models, the rule detects no separation on any of the three skills, and on two of them the draws could not have told. On code-review-and-quality the with-skill means differ by +0.009 and the rule reads no separation detected. On git-workflow-and-versioning they differ by -0.023; at these spreads the runner puts the draws that would have been needed at 16 per arm, against the 3 drawn. On documentation-and-adrs the move is the largest of the three, -0.143, and the two spreads sum to 0.251, which is at or above the floor, so no draw count at these spreads resolves it. None of this is evidence that the two models score these skills alike.
2. Claude Opus 5 on this run’s harness separates from its Report 009 receipts on no skill. The with-skill means move by -0.017 on code-review-and-quality, -0.001 on git-workflow-and-versioning, -0.018 on documentation-and-adrs. The rule reads no separation detected on the first two, and on documentation-and-adrs, where this run’s spread is 0.103 against Report 009’s 0.017, no separation detected; not enough draws to conclude at this effect floor. This is a statement about these draws on two harnesses and two hosts, not about Claude Code.
3. On two skills the two models’ baselines sit further apart than their with-skill scores, and no verdict here reads them. Without the skill, git-workflow-and-versioning reads 0.765 for Claude Opus 5 and 0.300 for Claude Opus 5.5, and code-review-and-quality 0.862 and 0.680. On those two skills the lifts therefore differ more than the with-skill scores do; on documentation-and-adrs the two baselines read 0.567 and 0.567. The runner’s comparison of two receipts does not read baseline arms, so this is offered as an observation to test on more cases, not as a finding; the section on the low figures says what those draws are.
Read from: docs/reports/011/evidence/code-review-and-quality-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/code-review-and-quality-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/diff--model--code-review-and-quality.md; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-5-2026-09-23.json; docs/reports/009/evidence/three-skill-comparison--driftproof--code-review-and-quality--receipts--code-review-and-quality-claude-opus-5-2026-09-15.json; docs/reports/009/evidence/three-skill-comparison--driftproof--git-workflow-and-versioning--receipts--git-workflow-and-versioning-claude-opus-5-2026-09-15.json; docs/reports/009/evidence/three-skill-comparison--driftproof--documentation-and-adrs--receipts--documentation-and-adrs-claude-opus-5-2026-09-15.json; docs/reports/011/evidence/diff--harness--documentation-and-adrs.md.
Run record
The run. Run 20260923T062806Z, one skill at a time, the Claude Opus 5.5 arm first and then the Claude Opus 5 arm, code-review-and-quality first in each. The command file writes a status line after each skill run; there are six, each with exit status 0 and one receipt, the last at 2026-09-23T07:08:39Z. Each receipt’s run.date_utc is written by the runner and is not read here as a start or a finish.
The caps. Per skill run, the runner’s own caps, as in Report 009 (see Setup). Across the whole run, a guard in front of Claude Code allows 240 calls in all and 240 seconds per call. The run used 204 calls, and the six receipts record 0 unmeasured draws.
The harness. Claude Code 2.1.280, installed at the npm integrity the run record states, with runner 0.10.1 at Report 009’s integrity. The runner writes no Claude Code version into a receipt, and a receipt is sealed, so each receipt has a sidecar, <receipt>.surface.json, bound to it by its receipt_hash and recording the version each of its calls reported; the table gives it.
The price. The runner’s registry had no row for claude-opus-5-5, so this run used a copy of it with one row added, at 4 and 20 dollars per million input and output tokens, read from the vendor’s pricing page on the day of the run (docs-pricing-snapshot-2026-09-23.json). The row is this run’s only; the product’s registry does not carry it.
The cost. Metered spend was 0.00 dollars: the claude-cli surface runs on a subscription. The estimated API-equivalent below is every draw’s generation and judge token counts, as each receipt records them, at the prices each receipt froze, with cached input counted at the full input price, as the runner counts it.
| arm | skill | calls | Claude Code reported by every call | generation, estimated USD | judge, estimated USD | total, estimated USD |
|---|---|---|---|---|---|---|
claude-opus-5-5 | code-review-and-quality | 28 | 2.1.280 | 0.42 | 1.17 | 1.58 |
claude-opus-5-5 | git-workflow-and-versioning | 24 | 2.1.280 | 0.19 | 0.63 | 0.82 |
claude-opus-5-5 | documentation-and-adrs | 36 | 2.1.280 | 0.36 | 1.29 | 1.65 |
claude-opus-5 | code-review-and-quality | 24 | 2.1.280 | 0.67 | 0.96 | 1.63 |
claude-opus-5 | git-workflow-and-versioning | 52 | 2.1.280 | 0.87 | 1.49 | 2.36 |
claude-opus-5 | documentation-and-adrs | 40 | 2.1.280 | 0.75 | 1.43 | 2.17 |
| all six | 3.25 | 6.97 | 10.21 | |||
Integrity. Receipt hashes: claude-opus-5-5 code-review-and-quality 5c3ad19b40dd4b39, claude-opus-5-5 git-workflow-and-versioning ca69c3436d0dbfa5, claude-opus-5-5 documentation-and-adrs b924a435c70e0092, claude-opus-5 code-review-and-quality d7bd2fcd3dbea881, claude-opus-5 git-workflow-and-versioning 2ba4ef175d7f1caf, claude-opus-5 documentation-and-adrs 34107e18c90f3b20. Each validates with its receipt hash verified. The raw call records carry the SKILL.md text and stay unpublished; their sha256 are published.
Read from: docs/reports/011/evidence/run-20260923T062806Z--run-record.json; docs/reports/011/evidence/run-20260923T062806Z--status.jsonl; docs/reports/011/evidence/guard.py; docs/reports/011/evidence/code-review-and-quality-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-5-2026-09-23.json; docs/reports/011/evidence/code-review-and-quality-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-2026-09-23.json; docs/reports/011/evidence/code-review-and-quality-claude-opus-5-5-2026-09-23.surface.json; docs/reports/011/evidence/code-review-and-quality-claude-opus-5-5-2026-09-23.stdout.txt; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-5-2026-09-23.surface.json; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-5-2026-09-23.surface.json; docs/reports/011/evidence/code-review-and-quality-claude-opus-5-2026-09-23.surface.json; docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-2026-09-23.surface.json; docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-2026-09-23.surface.json; docs/reports/011/evidence/run-20260923T062806Z--driftproof-calls.SHA256SUMS.
Published evidence
The files this report makes public. The six receipts with their summaries, surface sidecars and the runner’s console output; the six driftproof diff outputs the verdicts are read from; the run record, its registry copy, the status lines, the review of the low figures and the sha256 of every raw call record; the pricing snapshot; and the call guard.
docs/reports/011/evidence/code-review-and-quality-claude-opus-5-2026-09-23.json
sha25607daeb3f464c09d33318bee34921a91b84c8b789ae0d1d4a2f0caa85a095b326docs/reports/011/evidence/code-review-and-quality-claude-opus-5-2026-09-23.stdout.txt
sha256b404b6011deb84f6a37eecaad3bede694c0f6b2120bc923d8fcfdba7e98a1bccdocs/reports/011/evidence/code-review-and-quality-claude-opus-5-2026-09-23.summary.md
sha2562e5e46216bd1147daf4d92d620a60b74d6d5ff85f1c6bc40675452e25228b9fddocs/reports/011/evidence/code-review-and-quality-claude-opus-5-2026-09-23.surface.json
sha256bab7733ba5eb4c8d835d367134c513d0989670df856ce650c1061a0f6810bd7adocs/reports/011/evidence/code-review-and-quality-claude-opus-5-5-2026-09-23.json
sha25694f2eed19305f978ea87720e772ad71608e97c22407ed24ab6e6fe138deb571cdocs/reports/011/evidence/code-review-and-quality-claude-opus-5-5-2026-09-23.stdout.txt
sha2562ab1f3ea0ccd5289914dacde01068559c984aadabe64cd5e5dccab429509a27fdocs/reports/011/evidence/code-review-and-quality-claude-opus-5-5-2026-09-23.summary.md
sha25654f203b43249e2e1e524c0bacd2eb915276eaecbe3f42506a47a326084a00f52docs/reports/011/evidence/code-review-and-quality-claude-opus-5-5-2026-09-23.surface.json
sha2562b6b296ad44df26e452652b9e0e4dccecba795e348e90f5632fad2c9d962d57fdocs/reports/011/evidence/diff--harness--code-review-and-quality.md
sha256a54d66a5e425ba0130d245ccebd96d7e7efa4e1f48778b1c279b72e3feebd449docs/reports/011/evidence/diff--harness--documentation-and-adrs.md
sha256fb066d689a11c3ecd7eb572b0ff0d132fa1e495d85a2984b0635e6bee20d20e4docs/reports/011/evidence/diff--harness--git-workflow-and-versioning.md
sha256f2d1dcab5e7c9ee64104953592949cd08c5f6dd9b63721b11e39f80b729d78f4docs/reports/011/evidence/diff--model--code-review-and-quality.md
sha256f0bfab5aaf5eac0b29415e08ad442aa09bd6ba58304a8cdb529d31c4220697b0docs/reports/011/evidence/diff--model--documentation-and-adrs.md
sha2560cb80c5630cbdcbd53bf6ea3142da6b8a96df8a7f3118d39ca203b3101b2b716docs/reports/011/evidence/diff--model--git-workflow-and-versioning.md
sha256f0f6b2e80948256e2c3481704d3c06fed986221cf08dd72b7286959634ab56efdocs/reports/011/evidence/docs-pricing-snapshot-2026-09-23.json
sha256c25bc98ae2d3d9406feaccbddc36755a345aa950e7efb1d96234f910bd19fb68docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-2026-09-23.json
sha256c0412b61639f45b2a5c5ac701eba977a6d643b9a6d91e6e86e3feaf45c9e4c86docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-2026-09-23.stdout.txt
sha256c2f80a9622330045e57275178688d15aa23c433f321420523f9f62d4c6321bffdocs/reports/011/evidence/documentation-and-adrs-claude-opus-5-2026-09-23.summary.md
sha256629287738812f9882da9154447e8ed9a365ce04e7ab0acbe6ec7abb4c9dfa4e2docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-2026-09-23.surface.json
sha256a9322c335e39726f9ad7b76925add0d0cbd93b59465d818ab4de470797cfb7d2docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-5-2026-09-23.json
sha25661fe3ca86f3d41023169eea76653bac7b55318b2e1072c11bcb2a5247b6cccefdocs/reports/011/evidence/documentation-and-adrs-claude-opus-5-5-2026-09-23.stdout.txt
sha256590a76b48dfb0e3ce44cab9baf66225f6c7cbc3fd66ba8336c166f51643d6f93docs/reports/011/evidence/documentation-and-adrs-claude-opus-5-5-2026-09-23.summary.md
sha256e680cc4da944011a219b2490237f43d6c0a17ab2e1d66c69ba5f11f8c8d08d0edocs/reports/011/evidence/documentation-and-adrs-claude-opus-5-5-2026-09-23.surface.json
sha256fd26d7a011c3a7b5de41d22ba34ccb5d792816048a89f3f5d14ad364a18b2c1edocs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-2026-09-23.json
sha25627487953281201e3db25e32fdc8f868df132e2221bc01afc0e45e6f75a929c7ddocs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-2026-09-23.stdout.txt
sha256f54aac1969621e2946f1afcfecce2fb798681165f5bc8b9baf0e330bfe23c820docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-2026-09-23.summary.md
sha256757953a5b73384b27334c3c0e84b777f448d261369232e323446d4ba3fb0dcd7docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-2026-09-23.surface.json
sha2561820211b222c21efcc04755a3cc1af5632d098874a971b75f59cbe7e35d7faa7docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-5-2026-09-23.json
sha2567cd9748196d1ef1abadc1314fb5b923b6b1fd9bb930eb3b986e27bb76193936bdocs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-5-2026-09-23.stdout.txt
sha256adefcff6997f9ebf11b19f470ed4c085cfcc94df5b6c77439001c2b661110286docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-5-2026-09-23.summary.md
sha256034ecb528ee8c1254ce862afa018374b68afe2d02c5d258a876205cdcae8cb24docs/reports/011/evidence/git-workflow-and-versioning-claude-opus-5-5-2026-09-23.surface.json
sha2566e1886d1a7493bbe1fdf53077bc7ea5bb2aa9099c833637829c1d6a730686f16docs/reports/011/evidence/guard.py
sha2564be011477676abe4c97a3234ce2462f064c0b3fd7611df2871892d547e420435docs/reports/011/evidence/run-20260923T062806Z--artefact-review-20260923.md
sha256f0d9fae004c11a5d344d15a828f60bd3b601138ebcd20e6183e3d6db85bea7bbdocs/reports/011/evidence/run-20260923T062806Z--driftproof-calls.SHA256SUMS
sha2567397d5230f4881bd037b94a07571301d8f9afe86ba90911e9f3f2e2f2420cd65docs/reports/011/evidence/run-20260923T062806Z--registry.json
sha2562a0a6dd787d6a5f5255621e1eb48497eee30d31c470ed7be0c81366e08d252fddocs/reports/011/evidence/run-20260923T062806Z--run-record.json
sha2562696c287fd6115795c0fcbc4743a9be80872d5181a488bba60c33a20ad956b45docs/reports/011/evidence/run-20260923T062806Z--status.jsonl
sha256f4ba381a121115c0db26b37a85f29f9fc34c80e4fe3ac4d7cdd85722eec11bfd
Report 009’s three receipts, the other side of the second table, are published with that report: code-review-and-quality, git-workflow-and-versioning, documentation-and-adrs. Validate a receipt with npx driftproof validate <file>. Each copy here is byte-identical to its twin under specs/042-opus-5-5-release-receipts/, and each sidecar names the receipt it belongs to and that receipt’s hash.