Driftproof

Report #005 — what does a skill cost to run?

Value report — three axes (accuracy, cost, latency), held across 3 substrates. The substrate does not move under the skill here; the axes widen.

In 20 of 30 skill × substrate pairs at least one case cleared the effect floor with separated bands: 18 with an improving case (5 of them also with a regressing one), 2 with only a regressing case. 14 cleared the floor on aggregate: 10 carry a price and 4 report a saving instead, having improved quality while reducing cost.

Reports #001–#004 asked whether a skill still helps while something moved underneath it. This one holds that question and adds the half a reader needs before acting: what the skill costs to run, in money and in wall-clock. The same 10 suites (Report #001 v1.2 rubrics) and the same fixed judge (claude-haiku-4-5, n=5) run on 3 substrates; each row shows the accuracy lift with its bands, what the skill does to the call's token cost per 1,000 calls (which is not always upward), and what it does to latency.

How to read the three axes — and what they do NOT mean.
Disclosure.

Per-skill economics

skill claude-sonnet-5 claude-fable-5 gpt-5.6-sol
lift (band)Δ tokens (cost)Δ latencycost / benefit lift (band)Δ tokens (cost)Δ latencycost / benefit lift (band)Δ tokens (cost)Δ latencycost / benefit
code-review-and-quality 0.845 ± 0.038
base 0.810 · Δ +0.035 (driver-only)
+6,922 in · +229 out
derived: $24.20/1k calls
+3.58sn/a (driver-only) 0.889 ± 0.023
base 0.786 · Δ +0.103
+6,968 in · −3 out
derived: $69.52/1k calls
+1.84s$6.78 per 0.01 lift 0.798 ± 0.131
base 0.809 · Δ -0.011 (within noise)
+4,343 in · +93 out
derived: $24.50/1k calls
+2.62sn/a (within noise)
git-workflow-and-versioning 0.862 ± 0.016
base 0.746 · Δ +0.116
+4,956 in · −66 out
derived: $13.88/1k calls
−0.12s$1.19 per 0.01 lift 0.869 ± 0.018
base 0.882 · Δ -0.013 (within noise)
+4,954 in · −62 out
derived: $46.42/1k calls
−0.49sn/a (within noise) 0.819 ± 0.099
base 0.821 · Δ -0.002 (driver-only)
+3,256 in · +12 out
derived: $16.64/1k calls
−0.02sn/a (driver-only)
documentation-and-adrs 0.659 ± 0.338
base 0.707 · Δ -0.048 (driver-only)
+16,327 in · +372 out
derived: $54.56/1k calls
+2.15sn/a (driver-only) 0.868 ± 0.033
base 0.711 · Δ +0.157
+12,743 in · +37 out
derived: $129.29/1k calls
−1.62s$8.26 per 0.01 lift 0.799 ± 0.102
base 0.760 · Δ +0.038 (within noise)
−5,401 in · −72 out
derived: −$29.17/1k calls
+0.16sn/a (within noise)
commit-work 0.859 ± 0.025
base 0.837 · Δ +0.021 (within noise)
−4,054 in · +187 out
derived: −$9.36/1k calls
+3.48sn/a (within noise) 0.857 ± 0.055
base 0.867 · Δ -0.010 (driver-only)
+4,418 in · +234 out
derived: $55.88/1k calls
+2.50sn/a (driver-only) 0.837 ± 0.068
base 0.782 · Δ +0.054
+584 in · +160 out
derived: $7.72/1k calls
+4.28s$1.42 per 0.01 lift
writing-clearly-and-concisely 0.864 ± 0.022
base 0.853 · Δ +0.010 (within noise)
+1,394 in · +109 out
derived: $5.81/1k calls
+0.11sn/a (within noise) 0.864 ± 0.032
base 0.827 · Δ +0.037 (within noise)
+1,394 in · −25 out
derived: $12.68/1k calls
−0.40sn/a (within noise) 0.845 ± 0.050
base 0.853 · Δ -0.008 (within noise)
+845 in · +51 out
derived: $5.75/1k calls
+0.61sn/a (within noise)
crafting-effective-readmes 0.795 ± 0.234
base 0.811 · Δ -0.016 (within noise)
+11,627 in · +245 out
derived: $38.55/1k calls
+0.95sn/a (within noise) 0.883 ± 0.011
base 0.864 · Δ +0.019 (within noise)
+954 in · +46 out
derived: $11.82/1k calls
+0.78sn/a (within noise) 0.857 ± 0.038
base 0.863 · Δ -0.006 (driver-only)
−11,853 in · −165 out
derived: −$64.22/1k calls
−2.30sn/a (driver-only)
naming-analyzer 0.815 ± 0.096
base 0.687 · Δ +0.129
+4,025 in · +55 out
derived: $12.91/1k calls
+0.27s$1.00 per 0.01 lift 0.857 ± 0.037
base 0.805 · Δ +0.052
+4,000 in · +24 out
derived: $41.22/1k calls
−0.11s$7.97 per 0.01 lift 0.815 ± 0.089
base 0.747 · Δ +0.068
+2,393 in · +44 out
derived: $13.28/1k calls
+0.44s$1.94 per 0.01 lift
requesting-code-review 0.637 ± 0.291
base 0.496 · Δ +0.141
+25,429 in · −474 out
derived: $69.18/1k calls
−23.04s$4.92 per 0.01 lift 0.852 ± 0.052
base 0.678 · Δ +0.173
−45,045 in · −980 out
derived: −$499.46/1k calls
−15.29ssaves $499.46/1k calls 0.802 ± 0.052
base 0.721 · Δ +0.081
−14,804 in · −61 out
derived: −$75.86/1k calls
+0.92ssaves $75.86/1k calls
writing-plans 0.822 ± 0.078
base 0.646 · Δ +0.176
+1,078 in · +1,455 out
derived: $25.05/1k calls
−5.17s$1.42 per 0.01 lift 0.831 ± 0.092
base 0.654 · Δ +0.177
−23,161 in · +697 out
derived: −$196.77/1k calls
−4.55ssaves $196.77/1k calls 0.779 ± 0.112
base 0.717 · Δ +0.062
−253 in · +826 out
derived: $23.53/1k calls
+3.39s$3.79 per 0.01 lift
skill-creator 0.867 ± 0.022
base 0.803 · Δ +0.065
−6,530 in · −499 out
derived: −$27.08/1k calls
+2.94ssaves $27.08/1k calls 0.884 ± 0.013
base 0.880 · Δ +0.004 (within noise)
+10,959 in · −21 out
derived: $108.55/1k calls
+0.03sn/a (within noise) 0.857 ± 0.021
base 0.814 · Δ +0.042 (driver-only)
+29,494 in · +14 out
derived: $147.88/1k calls
−2.71sn/a (driver-only)

Per substrate: the with_skill band with its baseline and lift Δ; the token delta the skill adds (input and output), with the derived dollar cost per 1,000 calls beneath it; the median latency delta; and the cost per unit of measured benefit (dollars per 0.01 lift), which renders only where the aggregate lift cleared the floor — cells whose drivers cleared it while the aggregate did not read n/a (driver-only) and are listed under Verdict basis. Every number is re-derived from the receipts under receipts/report-005/; nothing is hand-entered. Deviation from the shared table anatomy, stated as REPORT-STYLE requires: a value report replaces the value-per-token, post-checks and composed-verdict columns with the three axes and the cost/benefit cell — the accuracy verdict lives in Verdict basis, per-case, rather than as a composed label.

Verdict basis — the per-case with/without drivers on each substrate. A case drives the accuracy axis only when its with_skill and baseline bands (mean ± stddev, n=5) do not overlap AND the mean moves ≥ 0.05. Cases marked †point-band rest on a zero-width band (all 5 judge samples identical).
Low-resolution: judge quantization — 13 of 30 cells. 10 of the 14 priced cells are among them. A zero-width point band means all 5 judge samples landed identically: the judge grades on a coarse grid, so the effect is a clean grid step the floor still gates, but one whose finer structure the judge cannot resolve. The verdict stands; it is flagged, not suppressed.

Run record

Run started 2026-08-18 19:01 UTC — receipt-attested, identical on all 30 receipts. Compute: 20.47 h (2.21 h generation over 420 calls, 18.26 h judging over 420 case rows). At the run's concurrency of 2 that is ≈10.2 h elapsed, a derived finish of ≈2026-08-19 05:15 UTC — derived, not attested: a receipt records per-call wall-clock, not a run finish, and concurrency is a launch parameter rather than receipt evidence. 30 receipts (10 skills × 3 substrates, with/without × n=5, judged by claude-haiku-4-5). Surfaces: claude-sonnet-5 → claude-cli, claude-fable-5 → claude-cli, gpt-5.6-sol → openai-cli. Metered spend $0 (vendor subscription CLIs); metered-equivalent $192.79 ($105.22 generation + $87.57 judge measurement, derived from the receipts at their frozen rates; verification level TESTED, the weakest among the 30 receipts summed). The $40 guard did not fire: it checked the $20.89 projection, not accrued spend — see the disclosure above. All receipts complete. [resume: 0 restored, 30 skipped]

Amendments

v1.2 · 2026-09-01. The three cells v1.1 flagged have been re-measured, and their published lifts are named here as what they are: single-draw, judge-spread figures. Each of the three lifts on this page rests on one generation per arm, and its band is the spread of the judge re-scoring that single response. Report 007 re-ran all three cells with the generation sampled adaptively, three draws per arm minimum and more until the across-draw spread settled, and found lower lifts in every one: code-review-and-quality on claude-fable-5 +0.103 to +0.055; git-workflow-and-versioning on claude-sonnet-5 +0.116 to -0.002; writing-plans on claude-fable-5 +0.177 to +0.131. At the cell level none of the three separates from its own band.

These are not corrections, and Report 007 does not present them as such. Two of the three comparisons were refused before any verdict was formed, by the same baseline-reproduction control that refused every cell of Report 006: the no-skill arm did not reproduce the arm this page measured. On the third, writing-plans on claude-fable-5, the control passed and the comparison rendered WITHIN NOISE, which says the two measurements are the same as far as the instrument can tell. All three cells also changed skill.content_hash upstream between the two runs, so none of the pairs is a straight re-measurement of the same text.

v1.1's "no cause is asserted" holds, for the same reason and one more. Sampling noise accounts for the observations without needing a cause; whether anything else also changed is not established either way. Report 007 adds a second, independent instrument change to the same gap: a declared 300 s call timeout on this surface had never executed, because a numeric literal at the call site shadowed it from 2026-07-27 onward. Every run behind this page used 120 s. The question stays open.

What this qualifies on the page above, in that page's own words. Two sentences already published here are the ones this amendment bears on, and they are quoted rather than paraphrased. The first is the cost-per-benefit rule: “The denominator is the cell’s aggregate lift, not the lift of the single case that drove it. The cost is paid on every case in the suite, so the benefit it buys must be averaged over those same cases; pricing a whole-suite cost against one case’s lift would mix populations and read 2.5–6× cheaper than the measurement supports.” Every $ per 0.01 lift figure on this page divides by a cell's aggregate lift, and two of those denominators are the lifts under revision here. The band this page prints beside each cell is the with_skill band, not the lift band, and the two are different quantities: where this page reads 0.889 ± 0.023 the receipt's delta_uncertainty is 0.216857. Nothing above has been swapped for anything else.

The second is this page's own fairness disclosure: “Checked against upstream on 2026-08-19: 8 of 10 are still byte-identical; code-review-and-quality and writing-plans have been revised upstream.” That sentence is why the 005 to 007 comparison is a revision pair on all three cells rather than a re-measurement of the same text, and why v1.2 cannot present Report 007's lower lifts as this page's figures measured again.

Unaffected, and unchanged from v1.1. The cost-driver correlation, the substrate-disagreement result and the three-axis presentation do not rest on single-cell baselines and are not qualified here. No figure on this page has been edited; every number, count and verdict above is as published and re-derives from the same 30 receipts.

v1.1 · 2026-08-29Three cells rest on single-draw baselines now shown unstable. code-review-and-quality on claude-fable-5, git-workflow-and-versioning on claude-sonnet-5, and writing-plans on claude-fable-5 each report a lift measured against a baseline arm generated once. A 20-draw stability probe run for Report #006 shows two of those baselines sit in distributions wide enough that a single draw does not pin them: one spans 0.300–0.808 across ten draws, the other sits between 0.214 and 0.300 in nine draws and reaches 0.856 in one. Generation-level variation exceeds judge-level variation by 3.2× and 7.5× on those cases, and this report sampled the judge five times and the generation once.

What this does not say. It does not say these lifts are wrong, and it does not supply corrected ones. Report #006 re-measured the same three cells and got lower lifts, but those re-measurements are equally single-draw and are not replacements — on commit-message-conventional-type the published 0.282 is the modal value across ten draws while #006’s 0.852 was the roughly 1-in-10 outcome, so the fresh number is the likelier fluke of the two. No cause is asserted. Sampling noise accounts for the observations without needing one; whether anything else also changed is not established either way, and the one-generation design of this report cannot settle it. The question is recorded as open.

Unaffected. The headline findings of this report do not rest on single-cell baselines and are not qualified here: the cost-driver correlation (input-token delta tracks cost at r = +0.92 against skill length at r = +0.33), the substrate-disagreement result, and the three-axis presentation stand as published. Every figure above is left legible and unedited. Methods: Report #006, stability probe.

v1.0.1 · 2026-08-19 — Publication-chrome correction. No measured value changed: every figure, count and verdict on this page is identical to v1.0, and all 30 receipts are unchanged. At first publication the page’s <title> element and its footer both still carried the draft marker. The publication switch removed the noindex robots tag and the draft banner, but was never applied to those two places, so a page that was published in every other respect still announced itself as unpublished in the browser tab. Both are now rendered from the same switch, and the repo gate asserts the property rather than the two known places: no published report page may carry that marker anywhere, checked across every published report, so the defect cannot recur silently on a later one.

v1.0 · 2026-08-19 — First publication.

Receipts

All 30 receipts for this report, each one resolvable. Every number on this page re-derives from these files. They are the evidence, not a description of it — open any one and check it against the table above.

Directory: receipts/report-005/. Validate any of them with npx driftproof validate <file>.

The value axes in this report type were prompted by a question from a former colleague: why not show what a skill costs to run, not just whether it helps.