Report #005 — what does a skill cost to run?
Value report — three axes (accuracy, cost, latency), held across 3 substrates. The substrate does not move under the skill here; the axes widen.
In 20 of 30 skill × substrate pairs at least one case cleared the effect floor with separated bands: 18 with an improving case (5 of them also with a regressing one), 2 with only a regressing case. 14 cleared the floor on aggregate: 10 carry a price and 4 report a saving instead, having improved quality while reducing cost.
Reports #001–#004 asked whether a skill still helps while something moved underneath it. This one holds that question and adds the half a reader needs before acting: what the skill costs to run, in money and in wall-clock. The same 10 suites (Report #001 v1.2 rubrics) and the same fixed judge (claude-haiku-4-5, n=5) run on 3 substrates; each row shows the accuracy lift with its bands, what the skill does to the call's token cost per 1,000 calls (which is not always upward), and what it does to latency.
- Accuracy lift (band).
with_skillmean ± stddev over 5 judge samples, against the no-skill baseline on the same substrate. A lift is only called real when a per-case band separation clears the 0.05 effect floor — the same anti-cry-wolf rule as every earlier report. Not a model ranking: the comparison is always with-skill vs without-skill within one substrate. - Δ tokens (cost). The change in the whole call's token footprint with the skill loaded — input and output. The skill's own text is one part of the input delta and usually the smaller part: across these 30 cells the input delta tracks cost at r = +0.92 while the skill's own length tracks it at only r = +0.33, and one 738-token skill drew 34× its own size in extra input. What moves the number is how the skill changes the path the model takes, multiplied by the substrate's price per token — the same skill at near-identical token deltas costs 3.3× more on
claude-fable-5than onclaude-sonnet-5, which is exactly their input-rate ratio in the frozen snapshot — the rates differ, not the behaviour. The change is not always upward: in 7 of 30 cells the skill made the call cheaper, and in 8 it pulled in less input — not the same set, since a skill can read less and still write more. This is the durable cost fact: it is what the skill actually consumed, and it does not change when a vendor reprices. The dollar figure beneath is a derived view of exactly these tokens at the rates frozen into the receipt. Not a bill: on subscription surfaces actual metered spend is $0 and the dollars are metered-equivalent. Not a total cost of ownership — it is the skill's marginal cost, the number that changes when you adopt it. Output length is not a quality signal in either direction; longer is not better. - Δ latency. Median wall-clock added, observed on subscription CLI surface, indicative. Not a serving-latency benchmark: it includes CLI cold starts and the vendor's own harness, so it describes what a user of that surface experiences, not the model's speed.
- Cost / benefit. What one unit of measured benefit costs — dollars per 0.01 lift, derived from the two columns to its left. The denominator is the cell's aggregate lift, not the lift of the single case that drove it. The cost is paid on every case in the suite, so the benefit it buys must be averaged over those same cases; pricing a whole-suite cost against one case's lift would mix populations and read 2.5–6× cheaper than the measurement supports. Four states, and only one of them is a price: a floor-clearing positive lift is priced; a cell whose drivers cleared the floor while its aggregate did not reads
n/a (driver-only), with those drivers named under Verdict basis; a cell with no separated driver at all readsn/a (within noise); and a skill that measurably hurt prices nothing. Where a skill improved quality and reduced cost, the cell states the saving instead of a price — there is no cost per unit of benefit when the benefit is free. - There is no composite score. The three axes have different units, different error bars, and different owners. Collapsing them into one number would manufacture a figure no reader could trace to evidence, so the report shows the three and leaves the weighing to you.
- Tokens are the measurement; dollars are derived. The cost column leads with the token delta — what the skill actually consumed, which does not change when a vendor reprices. The dollar figure beneath it is a derived view of exactly those tokens at the rates frozen into the receipt, and every one of them re-derives as
(input/1e6 × input_rate) + (output/1e6 × output_rate). - Pricing snapshot. Metered-equivalent, computed from each receipt's own frozen
run.pricing_snapshot— priced as frozen 2026-08-18 (registry prices at run time) — never from the live registry, so these numbers do not change meaning when a vendor changes prices. Cached input tokens are costed at the list input rate (an over-estimate where caching is heavy); cached_tokens is recorded per call so a reader can recompute. - Surfaces.
claude-sonnet-5→claude-cli(subscription; metered spend $0);claude-fable-5→claude-cli(subscription; metered spend $0);gpt-5.6-sol→openai-cli(subscription; metered spend $0). Disclosed per the neutrality policy. - The judge is excluded — and measurement is the expensive half in time. Judging consumed 18.26 of the run's 20.47 compute hours (89%), against 2.21 hours of generation (n=5 judge calls per generation). In dollars it was $87.57 of $192.79 — cheaper than generation in aggregate. It cost more than the work it graded on
claude-sonnet-5(judge $29.37 vs generation $28.93, 1.5% more, in 7 of its 10 receipts) andgpt-5.6-sol(judge $28.97 vs generation $17.12, 69.3% more, in 9 of its 10 receipts), and 16 of 30 receipts overall — so "grading costs more" is substrate-dependent, and onclaude-sonnet-5the margin is thin enough that a small pricing move would flip it. Judge tokens are recorded separately on every receipt asjudge_usageand enter no value figure — the receipt schema pinseconomics.judge_excludedtotrue. That spend is ours, for measuring; it is not a cost of running the skill. - What this run cost us, and what the guard did. The run was projected at $20.89 (the figure produced at run time, frozen here — the token constants behind it are being replaced, and a later re-render must not silently restate what was projected on the day) and cost $192.79 metered-equivalent — 9.2× the projection — against a $40.00 guard that never fired, because it checked the projection rather than accrued spend. Actual metered spend was $0 (subscription surfaces), so nothing was billed. The projection's token constants did not model the fixed CLI harness preamble; the guard is being rebuilt to check accrual before the next run. We publish the measured figure, not the estimate: it is the same argument this report makes about skills, applied to us.
- Skill versions are pinned, not current. Every skill was measured at the commit pinned in
suites/manifest.json. Checked against upstream on 2026-08-19: 8 of 10 are still byte-identical;code-review-and-qualityandwriting-planshave been revised upstream. No upstream revision responds to a Driftproof finding — no commit message or diff references this project. The influence has run the other way: a maintainer's fairness audit produced our in-text grounding policy, amended 2 of 10 suites, and changed a published verdict (Report #001 v1.2). - Ratios are floor-gated. The cost-per-benefit cell prices one unit of measured benefit — dollars per 0.01 lift. Where a lift did not clear the effect floor, it reads
n/a (within noise)rather than a number: dividing noise by a real cost produces a precise-looking figure with nothing under it.
Per-skill economics
| skill | claude-sonnet-5 | claude-fable-5 | gpt-5.6-sol | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| lift (band) | Δ tokens (cost) | Δ latency | cost / benefit | lift (band) | Δ tokens (cost) | Δ latency | cost / benefit | lift (band) | Δ tokens (cost) | Δ latency | cost / benefit | |
code-review-and-quality |
0.845 ± 0.038 base 0.810 · Δ +0.035 (driver-only) | +6,922 in · +229 out derived: $24.20/1k calls | +3.58s | n/a (driver-only) | 0.889 ± 0.023 base 0.786 · Δ +0.103 | +6,968 in · −3 out derived: $69.52/1k calls | +1.84s | $6.78 per 0.01 lift | 0.798 ± 0.131 base 0.809 · Δ -0.011 (within noise) | +4,343 in · +93 out derived: $24.50/1k calls | +2.62s | n/a (within noise) |
git-workflow-and-versioning |
0.862 ± 0.016 base 0.746 · Δ +0.116 | +4,956 in · −66 out derived: $13.88/1k calls | −0.12s | $1.19 per 0.01 lift | 0.869 ± 0.018 base 0.882 · Δ -0.013 (within noise) | +4,954 in · −62 out derived: $46.42/1k calls | −0.49s | n/a (within noise) | 0.819 ± 0.099 base 0.821 · Δ -0.002 (driver-only) | +3,256 in · +12 out derived: $16.64/1k calls | −0.02s | n/a (driver-only) |
documentation-and-adrs |
0.659 ± 0.338 base 0.707 · Δ -0.048 (driver-only) | +16,327 in · +372 out derived: $54.56/1k calls | +2.15s | n/a (driver-only) | 0.868 ± 0.033 base 0.711 · Δ +0.157 | +12,743 in · +37 out derived: $129.29/1k calls | −1.62s | $8.26 per 0.01 lift | 0.799 ± 0.102 base 0.760 · Δ +0.038 (within noise) | −5,401 in · −72 out derived: −$29.17/1k calls | +0.16s | n/a (within noise) |
commit-work |
0.859 ± 0.025 base 0.837 · Δ +0.021 (within noise) | −4,054 in · +187 out derived: −$9.36/1k calls | +3.48s | n/a (within noise) | 0.857 ± 0.055 base 0.867 · Δ -0.010 (driver-only) | +4,418 in · +234 out derived: $55.88/1k calls | +2.50s | n/a (driver-only) | 0.837 ± 0.068 base 0.782 · Δ +0.054 | +584 in · +160 out derived: $7.72/1k calls | +4.28s | $1.42 per 0.01 lift |
writing-clearly-and-concisely |
0.864 ± 0.022 base 0.853 · Δ +0.010 (within noise) | +1,394 in · +109 out derived: $5.81/1k calls | +0.11s | n/a (within noise) | 0.864 ± 0.032 base 0.827 · Δ +0.037 (within noise) | +1,394 in · −25 out derived: $12.68/1k calls | −0.40s | n/a (within noise) | 0.845 ± 0.050 base 0.853 · Δ -0.008 (within noise) | +845 in · +51 out derived: $5.75/1k calls | +0.61s | n/a (within noise) |
crafting-effective-readmes |
0.795 ± 0.234 base 0.811 · Δ -0.016 (within noise) | +11,627 in · +245 out derived: $38.55/1k calls | +0.95s | n/a (within noise) | 0.883 ± 0.011 base 0.864 · Δ +0.019 (within noise) | +954 in · +46 out derived: $11.82/1k calls | +0.78s | n/a (within noise) | 0.857 ± 0.038 base 0.863 · Δ -0.006 (driver-only) | −11,853 in · −165 out derived: −$64.22/1k calls | −2.30s | n/a (driver-only) |
naming-analyzer |
0.815 ± 0.096 base 0.687 · Δ +0.129 | +4,025 in · +55 out derived: $12.91/1k calls | +0.27s | $1.00 per 0.01 lift | 0.857 ± 0.037 base 0.805 · Δ +0.052 | +4,000 in · +24 out derived: $41.22/1k calls | −0.11s | $7.97 per 0.01 lift | 0.815 ± 0.089 base 0.747 · Δ +0.068 | +2,393 in · +44 out derived: $13.28/1k calls | +0.44s | $1.94 per 0.01 lift |
requesting-code-review |
0.637 ± 0.291 base 0.496 · Δ +0.141 | +25,429 in · −474 out derived: $69.18/1k calls | −23.04s | $4.92 per 0.01 lift | 0.852 ± 0.052 base 0.678 · Δ +0.173 | −45,045 in · −980 out derived: −$499.46/1k calls | −15.29s | saves $499.46/1k calls | 0.802 ± 0.052 base 0.721 · Δ +0.081 | −14,804 in · −61 out derived: −$75.86/1k calls | +0.92s | saves $75.86/1k calls |
writing-plans |
0.822 ± 0.078 base 0.646 · Δ +0.176 | +1,078 in · +1,455 out derived: $25.05/1k calls | −5.17s | $1.42 per 0.01 lift | 0.831 ± 0.092 base 0.654 · Δ +0.177 | −23,161 in · +697 out derived: −$196.77/1k calls | −4.55s | saves $196.77/1k calls | 0.779 ± 0.112 base 0.717 · Δ +0.062 | −253 in · +826 out derived: $23.53/1k calls | +3.39s | $3.79 per 0.01 lift |
skill-creator |
0.867 ± 0.022 base 0.803 · Δ +0.065 | −6,530 in · −499 out derived: −$27.08/1k calls | +2.94s | saves $27.08/1k calls | 0.884 ± 0.013 base 0.880 · Δ +0.004 (within noise) | +10,959 in · −21 out derived: $108.55/1k calls | +0.03s | n/a (within noise) | 0.857 ± 0.021 base 0.814 · Δ +0.042 (driver-only) | +29,494 in · +14 out derived: $147.88/1k calls | −2.71s | n/a (driver-only) |
Per substrate: the with_skill band with its baseline and lift Δ; the token delta the skill adds (input and output), with the derived dollar cost per 1,000 calls beneath it; the median latency delta; and the cost per unit of measured benefit (dollars per 0.01 lift), which renders only where the aggregate lift cleared the floor — cells whose drivers cleared it while the aggregate did not read n/a (driver-only) and are listed under Verdict basis. Every number is re-derived from the receipts under receipts/report-005/; nothing is hand-entered. Deviation from the shared table anatomy, stated as REPORT-STYLE requires: a value report replaces the value-per-token, post-checks and composed-verdict columns with the three axes and the cost/benefit cell — the accuracy verdict lives in Verdict basis, per-case, rather than as a composed label.
Verdict basis — the per-case with/without drivers on each substrate. A case drives the accuracy axis only when its with_skill and baseline bands (mean ± stddev, n=5) do not overlap AND the mean moves ≥ 0.05. Cases marked †point-band rest on a zero-width band (all 5 judge samples identical).
code-review-and-quality— claude-sonnet-5: 🔼severity-labeled-findings(Δ+0.196) · claude-fable-5: 🔼severity-labeled-findings(Δ+0.636) †point-bandgit-workflow-and-versioning— claude-sonnet-5: 🔼commit-message-conventional-type(Δ+0.554), 🔼changelog-curated-by-impact(Δ+0.238), 🔼semver-hidden-breaking-change(Δ+0.060) · gpt-5.6-sol: 🔻split-into-atomic-commits(Δ-0.264) †point-band, 🔼changelog-curated-by-impact(Δ+0.216), 🔼semver-hidden-breaking-change(Δ+0.064)documentation-and-adrs— claude-sonnet-5: 🔼document-public-api-function(Δ+0.200) †point-band, 🔼adr-for-costly-to-reverse-decision(Δ+0.120), 🔻match-existing-adr-convention(Δ-0.758) · claude-fable-5: 🔼comment-intent-not-implementation(Δ+0.268), 🔼document-public-api-function(Δ+0.200) †point-band, 🔼surface-conflicting-adr-conventions(Δ+0.602)commit-work— claude-fable-5: 🔻two-sentence-describability-test(Δ-0.122) · gpt-5.6-sol: 🔼conventional-commit-single-change(Δ+0.384), 🔻two-sentence-describability-test(Δ-0.138)crafting-effective-readmes— gpt-5.6-sol: 🔻categorize-task-before-writing(Δ-0.078) †point-bandnaming-analyzer— claude-sonnet-5: 🔼boolean-prefixes-js(Δ+0.344), 🔻misleading-name-mutation-js(Δ-0.054), 🔼abbreviations-wellknown-js(Δ+0.362), 🔼language-casing-python(Δ+0.300) †point-band · claude-fable-5: 🔼language-casing-python(Δ+0.270) †point-band · gpt-5.6-sol: 🔼boolean-prefixes-js(Δ+0.174), 🔼language-casing-python(Δ+0.266) †point-bandrequesting-code-review— claude-sonnet-5: 🔼request-package-basic(Δ+0.566) †point-band, 🔼crafted-context-not-session-history(Δ+0.582), 🔼mandatory-vs-optional-triggers(Δ+0.546), 🔻resist-simple-self-review(Δ-0.644) · claude-fable-5: 🔼triage-review-findings(Δ+0.168), 🔼crafted-context-not-session-history(Δ+0.572) †point-band, 🔼resist-simple-self-review(Δ+0.108), 🔼full-handoff-before-merge(Δ+0.410) †point-band · gpt-5.6-sol: 🔼triage-review-findings(Δ+0.134), 🔼resist-simple-self-review(Δ+0.176), 🔼full-handoff-before-merge(Δ+0.274) †point-bandwriting-plans— claude-sonnet-5: 🔼repair-placeholder-steps(Δ+0.056), 🔼interfaces-exact-signatures(Δ+0.350) †point-band, 🔼task-right-sizing-testable-deliverable(Δ+0.354), 🔼full-small-plan-header-and-tasks(Δ+0.442) · claude-fable-5: 🔼bite-sized-tdd-steps(Δ+0.280), 🔼repair-placeholder-steps(Δ+0.072), 🔼interfaces-exact-signatures(Δ+0.280), 🔼full-small-plan-header-and-tasks(Δ+0.422) · gpt-5.6-sol: 🔼bite-sized-tdd-steps(Δ+0.272) †point-bandskill-creator— claude-sonnet-5: 🔼write-triggering-description(Δ+0.152) · gpt-5.6-sol: 🔼write-triggering-description(Δ+0.192), 🔼draft-full-skillmd(Δ+0.154)
Low-resolution: judge quantization — 13 of 30 cells. 10 of the 14 priced cells are among them. A zero-width point band means all 5 judge samples landed identically: the judge grades on a coarse grid, so the effect is a clean grid step the floor still gates, but one whose finer structure the judge cannot resolve. The verdict stands; it is flagged, not suppressed.
code-review-and-qualityonclaude-fable-5(priced) —severity-labeled-findings(Δ+0.636)git-workflow-and-versioningongpt-5.6-sol—split-into-atomic-commits(Δ-0.264)documentation-and-adrsonclaude-sonnet-5—document-public-api-function(Δ+0.200)documentation-and-adrsonclaude-fable-5(priced) —document-public-api-function(Δ+0.200)crafting-effective-readmesongpt-5.6-sol—categorize-task-before-writing(Δ-0.078)naming-analyzeronclaude-sonnet-5(priced) —language-casing-python(Δ+0.300)naming-analyzeronclaude-fable-5(priced) —language-casing-python(Δ+0.270)naming-analyzerongpt-5.6-sol(priced) —language-casing-python(Δ+0.266)requesting-code-reviewonclaude-sonnet-5(priced) —request-package-basic(Δ+0.566)requesting-code-reviewonclaude-fable-5(priced) —crafted-context-not-session-history(Δ+0.572),full-handoff-before-merge(Δ+0.410)requesting-code-reviewongpt-5.6-sol(priced) —full-handoff-before-merge(Δ+0.274)writing-plansonclaude-sonnet-5(priced) —interfaces-exact-signatures(Δ+0.350)writing-plansongpt-5.6-sol(priced) —bite-sized-tdd-steps(Δ+0.272)
Run record
Run started 2026-08-18 19:01 UTC — receipt-attested, identical on all 30 receipts. Compute: 20.47 h (2.21 h generation over 420 calls, 18.26 h judging over 420 case rows). At the run's concurrency of 2 that is ≈10.2 h elapsed, a derived finish of ≈2026-08-19 05:15 UTC — derived, not attested: a receipt records per-call wall-clock, not a run finish, and concurrency is a launch parameter rather than receipt evidence. 30 receipts (10 skills × 3 substrates, with/without × n=5, judged by claude-haiku-4-5). Surfaces: claude-sonnet-5 → claude-cli, claude-fable-5 → claude-cli, gpt-5.6-sol → openai-cli. Metered spend $0 (vendor subscription CLIs); metered-equivalent $192.79 ($105.22 generation + $87.57 judge measurement, derived from the receipts at their frozen rates; verification level TESTED, the weakest among the 30 receipts summed). The $40 guard did not fire: it checked the $20.89 projection, not accrued spend — see the disclosure above. All receipts complete. [resume: 0 restored, 30 skipped]
Amendments
v1.2 · 2026-09-01. The three cells v1.1 flagged have been re-measured, and their published lifts are named here as what they are: single-draw, judge-spread figures. Each of the three lifts on this page rests on one generation per arm, and its band is the spread of the judge re-scoring that single response. Report 007 re-ran all three cells with the generation sampled adaptively, three draws per arm minimum and more until the across-draw spread settled, and found lower lifts in every one: code-review-and-quality on claude-fable-5 +0.103 to +0.055; git-workflow-and-versioning on claude-sonnet-5 +0.116 to -0.002; writing-plans on claude-fable-5 +0.177 to +0.131. At the cell level none of the three separates from its own band.
These are not corrections, and Report 007 does not present them as such. Two of the three comparisons were refused before any verdict was formed, by the same baseline-reproduction control that refused every cell of Report 006: the no-skill arm did not reproduce the arm this page measured. On the third, writing-plans on claude-fable-5, the control passed and the comparison rendered WITHIN NOISE, which says the two measurements are the same as far as the instrument can tell. All three cells also changed skill.content_hash upstream between the two runs, so none of the pairs is a straight re-measurement of the same text.
v1.1's "no cause is asserted" holds, for the same reason and one more. Sampling noise accounts for the observations without needing a cause; whether anything else also changed is not established either way. Report 007 adds a second, independent instrument change to the same gap: a declared 300 s call timeout on this surface had never executed, because a numeric literal at the call site shadowed it from 2026-07-27 onward. Every run behind this page used 120 s. The question stays open.
What this qualifies on the page above, in that page's own words. Two sentences already published here are the ones this amendment bears on, and they are quoted rather than paraphrased. The first is the cost-per-benefit rule: “The denominator is the cell’s aggregate lift, not the lift of the single case that drove it. The cost is paid on every case in the suite, so the benefit it buys must be averaged over those same cases; pricing a whole-suite cost against one case’s lift would mix populations and read 2.5–6× cheaper than the measurement supports.” Every $ per 0.01 lift figure on this page divides by a cell's aggregate lift, and two of those denominators are the lifts under revision here. The band this page prints beside each cell is the with_skill band, not the lift band, and the two are different quantities: where this page reads 0.889 ± 0.023 the receipt's delta_uncertainty is 0.216857. Nothing above has been swapped for anything else.
The second is this page's own fairness disclosure: “Checked against upstream on 2026-08-19: 8 of 10 are still byte-identical; code-review-and-quality and writing-plans have been revised upstream.” That sentence is why the 005 to 007 comparison is a revision pair on all three cells rather than a re-measurement of the same text, and why v1.2 cannot present Report 007's lower lifts as this page's figures measured again.
Unaffected, and unchanged from v1.1. The cost-driver correlation, the substrate-disagreement result and the three-axis presentation do not rest on single-cell baselines and are not qualified here. No figure on this page has been edited; every number, count and verdict above is as published and re-derives from the same 30 receipts.
v1.1 · 2026-08-29 — Three cells rest on single-draw baselines now shown unstable. code-review-and-quality on claude-fable-5, git-workflow-and-versioning on claude-sonnet-5, and writing-plans on claude-fable-5 each report a lift measured against a baseline arm generated once. A 20-draw stability probe run for Report #006 shows two of those baselines sit in distributions wide enough that a single draw does not pin them: one spans 0.300–0.808 across ten draws, the other sits between 0.214 and 0.300 in nine draws and reaches 0.856 in one. Generation-level variation exceeds judge-level variation by 3.2× and 7.5× on those cases, and this report sampled the judge five times and the generation once.
What this does not say. It does not say these lifts are wrong, and it does not supply corrected ones. Report #006 re-measured the same three cells and got lower lifts, but those re-measurements are equally single-draw and are not replacements — on commit-message-conventional-type the published 0.282 is the modal value across ten draws while #006’s 0.852 was the roughly 1-in-10 outcome, so the fresh number is the likelier fluke of the two. No cause is asserted. Sampling noise accounts for the observations without needing one; whether anything else also changed is not established either way, and the one-generation design of this report cannot settle it. The question is recorded as open.
Unaffected. The headline findings of this report do not rest on single-cell baselines and are not qualified here: the cost-driver correlation (input-token delta tracks cost at r = +0.92 against skill length at r = +0.33), the substrate-disagreement result, and the three-axis presentation stand as published. Every figure above is left legible and unedited. Methods: Report #006, stability probe.
v1.0.1 · 2026-08-19 — Publication-chrome correction. No measured value changed: every figure, count and verdict on this page is identical to v1.0, and all 30 receipts are unchanged. At first publication the page’s <title> element and its footer both still carried the draft marker. The publication switch removed the noindex robots tag and the draft banner, but was never applied to those two places, so a page that was published in every other respect still announced itself as unpublished in the browser tab. Both are now rendered from the same switch, and the repo gate asserts the property rather than the two known places: no published report page may carry that marker anywhere, checked across every published report, so the defect cannot recur silently on a later one.
v1.0 · 2026-08-19 — First publication.
Receipts
All 30 receipts for this report, each one resolvable. Every number on this page re-derives from these files. They are the evidence, not a description of it — open any one and check it against the table above.
code-review-and-quality— claude-fable-5 · claude-sonnet-5 · gpt-5.6-solcommit-work— claude-fable-5 · claude-sonnet-5 · gpt-5.6-solcrafting-effective-readmes— claude-fable-5 · claude-sonnet-5 · gpt-5.6-soldocumentation-and-adrs— claude-fable-5 · claude-sonnet-5 · gpt-5.6-solgit-workflow-and-versioning— claude-fable-5 · claude-sonnet-5 · gpt-5.6-solnaming-analyzer— claude-fable-5 · claude-sonnet-5 · gpt-5.6-solrequesting-code-review— claude-fable-5 · claude-sonnet-5 · gpt-5.6-solskill-creator— claude-fable-5 · claude-sonnet-5 · gpt-5.6-solwriting-clearly-and-concisely— claude-fable-5 · claude-sonnet-5 · gpt-5.6-solwriting-plans— claude-fable-5 · claude-sonnet-5 · gpt-5.6-sol
Directory: receipts/report-005/. Validate any of them with npx driftproof validate <file>.
The value axes in this report type were prompted by a question from a former colleague: why not show what a skill costs to run, not just whether it helps.