The Driftproof receipt for code-review-and-quality on gpt-5.6-sol, run on 2026-08-18, and what it can say.
code-review-and-quality on gpt-5.6-sol
Not measured
This receipt carries no verdict: it does not record that a model answered the run.
The two arms
| arm | mean | spread across cases | cases |
|---|---|---|---|
| with the skill | 0.798 | ± 0.131 | 7 |
| without the skill | 0.809 | ± 0.110 | 7 |
No lift is stated for a state that did not measure one. A spread is the sample standard deviation of the per-case means: a descriptive spread with no coverage probability. The verdict above reads each case's two bands, not these aggregates.
What was verified
receipt_hash verified when this page was built from receipts/report-005/code-review-and-quality__gpt-5.6-sol.json; run.transcripts records retained-local.
Receipt hash: 43bd1219e2bd2045bad5b35b1bc797326db10aac7c7f041cba64f2a7fef6b9ca
Verify this receipt yourself
curl -fsSLO https://raw.githubusercontent.com/driftproofhq/driftproof/main/receipts/report-005/code-review-and-quality__gpt-5.6-sol.json
npx driftproof validate code-review-and-quality__gpt-5.6-sol.json