Driftproof

Report 009: three skills under two eval harnesses

Instrument comparison report. Nothing moves under the skill, and the thing that differs is the harness that measures it. The page follows the shared chrome and states its own reading rules in its limits and setup sections.

Three skills from one plugin, one case each, measured by Claude Code’s native plugin eval and by Driftproof, with the same SKILL.md bytes, task prompts and rubrics. The two tools apply different treatments and grade differently, so this page reads their results side by side and does not rank them.

Every figure below is read from a file in the two bundles the runs were delivered in, and each block names its files. The target model and the judge were claude-opus-5 in both tools.

Read from: docs/reports/009/evidence/three-skill-comparison--driftproof--code-review-and-quality--receipts--code-review-and-quality-claude-opus-5-2026-09-15.json.

Limits, read these first

Read from: docs/reports/009/evidence/three-skill-comparison--driftproof--code-review-and-quality--receipts--code-review-and-quality-claude-opus-5-2026-09-15.json; docs/reports/009/evidence/three-skill-comparison--driftproof--git-workflow-and-versioning--receipts--git-workflow-and-versioning-claude-opus-5-2026-09-15.json; docs/reports/009/evidence/three-skill-comparison--driftproof--documentation-and-adrs--receipts--documentation-and-adrs-claude-opus-5-2026-09-15.json; docs/reports/009/evidence/three-skill-comparison--native--code-review-and-quality--results--aggregate-result.json; docs/reports/009/evidence/three-skill-comparison--native--git-workflow-and-versioning--results--aggregate-result.json; docs/reports/009/evidence/three-skill-comparison--native--documentation-and-adrs--results--aggregate-result.json; docs/reports/009/evidence/three-skill-comparison--native-trace-pass--code-review-and-quality--results--aggregate-result.json; docs/reports/009/evidence/three-skill-comparison--native-trace-pass--git-workflow-and-versioning--results--aggregate-result.json; docs/reports/009/evidence/three-skill-comparison--native-trace-pass--documentation-and-adrs--results--aggregate-result.json; three-skill-comparison/driftproof-calls/041/stdout.jsonl (in the bundle, not published; its sha256 is listed in the bundle's published SHA256SUMS); three-skill-comparison/driftproof-calls/045/stdout.jsonl (in the bundle, not published; its sha256 is listed in the bundle's published SHA256SUMS); three-skill-comparison/driftproof-calls/049/stdout.jsonl (in the bundle, not published; its sha256 is listed in the bundle's published SHA256SUMS); three-skill-comparison/inputs/native/git-workflow-and-versioning/evals/commit-message-conventional-type/graders/criteria.md (in the bundle, not published; its sha256 is listed in the bundle's published SHA256SUMS).

What each tool measures here

The inputs are shared. Three skills from addyosmani/agent-skills at commit be4e44a9fbc5e8df0beaefadbb28bd22ee61cc39: code-review-and-quality, git-workflow-and-versioning, documentation-and-adrs. The SKILL.md bytes, the task prompt and the rubric are the same in both tools’ inputs; spec 034 checks the SKILL.md bytes against the upstream commit and the prompt and rubric text between the two inputs directories.

The native eval is Claude Code 2.1.272 running each skill as a single-skill plugin, in a with-plugin arm and a without-plugin arm (with-without). Each run is a real session: the model sees the plugin, can discover the skill and call the Skill tool, and has up to 5 turns. A criteria grader applies the rubric with the judge and returns PASS when the rubric score is at least 0.7, taking 3 judge votes per run. An arm’s score is the fraction of its runs that passed. A second grader, skill-fired, records whether the Skill tool was called; it is diagnostic and is not part of the score (scored is false).

Driftproof is runner 0.10.1 on the claude-cli surface. It scores the skill text in the prompt: the with-skill arm receives the SKILL.md with the task, the baseline arm receives the task alone, and neither has tools. It draws several generations per arm and the judge scores each one 3 times on a continuous 0 to 1 scale against the same rubric. An arm’s band is the mean of its draw means plus or minus the sample standard deviation across draws: a descriptive spread with no coverage probability. Two arms separate under the rule when their bands do not overlap and their means differ by at least the 0.05 effect floor; otherwise no separation is detected at the sample size used, which is not evidence that nothing differs.

So the native judge and the Driftproof judge read the same rubric in two ways: the native grader turns it into pass or fail at 0.7, and Driftproof keeps the score.

Read from: docs/reports/009/evidence/three-skill-comparison--protocol.md; docs/reports/009/evidence/three-skill-comparison--native--code-review-and-quality--results--aggregate-result.json; docs/reports/009/evidence/three-skill-comparison--driftproof--code-review-and-quality--receipts--code-review-and-quality-claude-opus-5-2026-09-15.json; three-skill-full-plugin/driftproof-source/README.md (in the bundle, not published; its sha256 is listed in the bundle's published SHA256SUMS).

The text-only comparison

skillnative with plugin, runs passednative without, runs passednative deltanative skill-fired: comparison; supplementDriftproof with skill, mean ± sd across drawsDriftproof baseline, mean ± sd across drawsDriftproof deltaDriftproof arms under the rule
code-review-and-quality3 of 33 of 30.003 of 3; 1 of 10.918 ± 0.010 (3 draws)0.783 ± 0.094 (4 draws)0.134separated
git-workflow-and-versioning3 of 32 of 30.333 of 3; 1 of 10.862 ± 0.007 (3 draws)0.300 ± 0.000 (3 draws)0.562separated
documentation-and-adrs3 of 31 of 30.670 of 3; 0 of 10.822 ± 0.017 (3 draws)0.532 ± 0.093 (4 draws)0.291separated

A native delta is the with-plugin pass fraction minus the without-plugin pass fraction. A Driftproof delta is the with-skill mean minus the baseline mean, shown as context; the rule column is what the rule reads. The two deltas are on different scales and are not compared with each other.

Read from: docs/reports/009/evidence/three-skill-comparison--native--code-review-and-quality--results--aggregate-result.json; docs/reports/009/evidence/three-skill-comparison--native-trace-pass--code-review-and-quality--results--aggregate-result.json; docs/reports/009/evidence/three-skill-comparison--driftproof--code-review-and-quality--receipts--code-review-and-quality-claude-opus-5-2026-09-15.json; docs/reports/009/evidence/three-skill-comparison--native--git-workflow-and-versioning--results--aggregate-result.json; docs/reports/009/evidence/three-skill-comparison--native-trace-pass--git-workflow-and-versioning--results--aggregate-result.json; docs/reports/009/evidence/three-skill-comparison--driftproof--git-workflow-and-versioning--receipts--git-workflow-and-versioning-claude-opus-5-2026-09-15.json; docs/reports/009/evidence/three-skill-comparison--native--documentation-and-adrs--results--aggregate-result.json; docs/reports/009/evidence/three-skill-comparison--native-trace-pass--documentation-and-adrs--results--aggregate-result.json; docs/reports/009/evidence/three-skill-comparison--driftproof--documentation-and-adrs--receipts--documentation-and-adrs-claude-opus-5-2026-09-15.json.

Three readings

1. On code-review-and-quality the native delta is 0.00, while Driftproof’s arms separate under the rule on that one task. Both native arms passed 3 of 3 runs. Driftproof’s with-skill band runs from 0.908 to 0.928 and its baseline band from 0.689 to 0.878, so they do not overlap, and the means differ by 0.134, at least the effect floor. This is a separation detected under the rule on one task, not proof that the skill moved the score. Neither tool’s result is read against the other’s.

2. The native eval’s largest delta, 0.67 on documentation-and-adrs, is on a skill whose skill-fired grader reads 0 of 3 and 0 of 1. The Skill tool was not called in the three with-plugin runs of the comparison or in the one-run supplement. The with-plugin arm passed 3 of 3 runs and the without-plugin arm 1 of 3, so the delta is 2 baseline failures in 3 runs, and the one baseline run that passed did so on the judge votes FAIL PASS PASS. With three runs per arm each run moves a pass fraction by a third. These files do not show what produced the difference.

3. With-skill spread is well below baseline spread on two of the three skills, offered as an observation to test. On code-review-and-quality the sd across draws is 0.010 with the skill and 0.094 without, a ratio of 9.3; on documentation-and-adrs it is 0.017 and 0.093, a ratio of 5.5. On git-workflow-and-versioning the pattern does not appear: the baseline sd is 0.000, with every one of its draw means at 0.300, and the with-skill sd 0.007. Each figure is from one case and three or four draws, so this is a pattern to test on more cases and more draws, not a finding.

Read from: docs/reports/009/evidence/three-skill-comparison--native--code-review-and-quality--results--aggregate-result.json; docs/reports/009/evidence/three-skill-comparison--driftproof--code-review-and-quality--receipts--code-review-and-quality-claude-opus-5-2026-09-15.json; docs/reports/009/evidence/three-skill-comparison--native--documentation-and-adrs--results--aggregate-result.json; docs/reports/009/evidence/three-skill-comparison--native-trace-pass--documentation-and-adrs--results--aggregate-result.json; docs/reports/009/evidence/three-skill-comparison--driftproof--documentation-and-adrs--receipts--documentation-and-adrs-claude-opus-5-2026-09-15.json; docs/reports/009/evidence/three-skill-comparison--driftproof--git-workflow-and-versioning--receipts--git-workflow-and-versioning-claude-opus-5-2026-09-15.json.

A separate run: the full plugin on tool-using tasks, native only

The second bundle records a different run with different tasks, reported here on its own and as a native-only functional check. Its tasks are not the tasks above and its figures are not comparable with them. The whole plugin was installed, version 0.6.9, including its default SessionStart hook, and each of three tasks asked for real tool use in a small repository: review a working-tree change and write review.md, make atomic Git commits, and write an ADR in the repository’s reStructuredText convention. Each task ran 3 times per arm. The graders were deterministic checks with no model judge, and a separate verification script read each session’s Git history and files afterwards.

Every session passed. 18 of 18 sessions passed the native graders, 9 of 9 with the plugin and 9 of 9 without, and 18 of 18 passed the post-session verification. These checks record no difference between the arms on these tasks, and they were written as coarse functional checks, not as a measure of how much a skill helps.

The plugin’s hook context is in every with-plugin trace. The verification records SessionStart hook context in 9 of 9 with-plugin sessions and in 0 of 9 without.

One ADR passed every structural check while asserting a history the fixture never supplied. In with-plugin session 3 of the ADR task the verification passed 10 of 10 checks, and the ADR it wrote says at line 18: Both have been observed in practice with in-request sends. The decision brief the fixture writes into the repository states no such observation, and the bundle’s manual inspection records: Structural and stated design checks do not validate this assertion.

Read from: docs/reports/009/evidence/three-skill-full-plugin--native--results--aggregate-result.json; three-skill-full-plugin/verification.json (in the bundle, not published; its sha256 is listed in the bundle's published SHA256SUMS); three-skill-full-plugin/workspaces/documentation-and-adrs/with-3/cwd/Documentation/Decisions/ADR-003-Email-Outbox.rst (in the bundle, not published; its sha256 is listed in the bundle's published SHA256SUMS); three-skill-full-plugin/suite/documentation-and-adrs/fixture.sh (in the bundle, not published; its sha256 is listed in the bundle's published SHA256SUMS); three-skill-full-plugin/manual-inspection.json (in the bundle, not published; its sha256 is listed in the bundle's published SHA256SUMS); docs/reports/009/evidence/three-skill-full-plugin--protocol.md.

Run record

The comparison ran on 2026-09-15 by each receipt’s run.date_utc (2026-09-15T16:04:34.736Z, 2026-09-15T16:06:45.501Z, 2026-09-15T16:11:24.010Z) and each native result’s startedAt (2026-09-15T15:57:22.219Z, 2026-09-15T16:01:49.644Z, 2026-09-15T16:04:35.426Z). Those are start times; neither file records a finish. The full-plugin run started at 2026-09-15T17:26:26.349Z on Claude Code 2.1.272.

Integrity. Each bundle carries a SHA256SUMS list, 544 lines for the comparison and 1046 for the full-plugin run, and every file each bundle holds matches its list. The three SKILL.md files in both bundles match the upstream commit named above byte for byte. The files published beside this report are copied from the bundles unchanged and match the same lists. Receipt hashes: code-review-and-quality 4e3e1dddda5592ac, git-workflow-and-versioning 2df794c7f2539a9c, documentation-and-adrs 3f52ddb6b849ee54.

Read from: docs/reports/009/evidence/three-skill-comparison--driftproof--code-review-and-quality--receipts--code-review-and-quality-claude-opus-5-2026-09-15.json; docs/reports/009/evidence/three-skill-comparison--driftproof--git-workflow-and-versioning--receipts--git-workflow-and-versioning-claude-opus-5-2026-09-15.json; docs/reports/009/evidence/three-skill-comparison--driftproof--documentation-and-adrs--receipts--documentation-and-adrs-claude-opus-5-2026-09-15.json; docs/reports/009/evidence/three-skill-comparison--native--code-review-and-quality--results--aggregate-result.json; docs/reports/009/evidence/three-skill-comparison--native--git-workflow-and-versioning--results--aggregate-result.json; docs/reports/009/evidence/three-skill-comparison--native--documentation-and-adrs--results--aggregate-result.json; docs/reports/009/evidence/three-skill-full-plugin--native--results--aggregate-result.json; docs/reports/009/evidence/three-skill-comparison--SHA256SUMS; docs/reports/009/evidence/three-skill-full-plugin--SHA256SUMS.

Published evidence

The files this report makes public. The three Driftproof receipts with their summaries, the native tool’s aggregate results, each run’s protocol, and each bundle’s SHA-256 list.

Validate a receipt with npx driftproof validate <file>. Each published copy is named for its path inside its bundle, with each / written as --; its sha256 is the one its bundle’s SHA256SUMS lists for that path.