# Driftproof > A public instrument and public record measuring whether agent skills still > deliver their claimed lift as models, providers and releases move underneath > them. Every published number carries a dated, hash-verified receipt. Driftproof runs a skill's own eval suite twice, with the skill and without it, scores each answer several times so every result is a range rather than one fragile number, and claims a verdict only when the two ranges do not overlap and the move clears the effect floor. When it cannot stand behind a number it publishes a refusal instead of a guess. ## Reports - [Report 008: the skill stabilises the floor, not the ceiling](https://driftproofhq.com/reports/008/): Release drift report, 2026-09-02. Both cells came back WITHIN NOISE. Models: claude-fable-5-1. - [Report 007: the baseline is the noisy arm](https://driftproofhq.com/reports/007/): Instrument re-measurement report, 2026-08-31. Three skills re-measured under generation sampling, with a corrected instrument. Models: claude-fable-5, claude-sonnet-5. - [Report 006: the skill text moves, the substrate holds still](https://driftproofhq.com/reports/006/): Revision drift report, 2026-08-28. Three cells, one per skill upstream had revised since Report #005 pinned it, and none of the three returned a verdict. Models: claude-fable-5, claude-sonnet-5. - [Report 005: what does a skill cost to run?](https://driftproofhq.com/reports/005/): Value report, 2026-08-18. In 20 of 30 skill × substrate pairs at least one case cleared the effect floor with separated bands: 18 with an improving case (5 of them also with a regressing one), 2 with only a regressing case. Models: claude-fable-5, claude-sonnet-5, gpt-5.6-sol. - [Report 004: does encoded expertise still lift output on the frontier tier?](https://driftproofhq.com/reports/004/): Capability-gap report, 2026-08-16. 3 durable · 5 tier-dependent · 0 regresses · 2 no effect Models: claude-fable-5, claude-opus-5. - [Report 003: do skill verdicts hold across a model release?](https://driftproofhq.com/reports/003/): Release drift report, 2026-08-12. 2 regressed · 4 improved · 0 mixed · 4 within noise Models: claude-opus-4-8, claude-opus-5. - [Report 002: does a skill's benefit hold across substrates?](https://driftproofhq.com/reports/002/): Substrate durability report, 2026-07-31. Cross-provider skill durability, not a model ranking. Models: claude-sonnet-5, gpt-5.6-sol. - [Report 001: do skill verdicts hold across a model release?](https://driftproofhq.com/reports/001/): Release drift report, 2026-07-30. 9 of 10 skills moved beyond noise on this model pair. Models: claude-sonnet-4-6, claude-sonnet-5. ## Method and definitions - [Methodology](https://driftproofhq.com/methodology/): how a run is scored, what a band is, and what the effect floor does. - [Neutrality](https://driftproofhq.com/neutrality/): what Driftproof will not claim, and the limitations it discloses. - [Glossary](https://driftproofhq.com/glossary/): drift report, receipt, band, effect floor, substrate, surface, refusal, cell, arm, and the verification lattice. - [Report types](https://driftproofhq.com/report-types/): the six kinds of report and what moves underneath the skill in each. - [Receipt specification](https://github.com/driftproofhq/driftproof/blob/main/spec/RECEIPT.md): the receipt format, its schema versions, and the UNVERIFIED / DECLARED / TESTED / FORMAL levels. ## Optional - [Interop](https://driftproofhq.com/interop/): importing results from other tools as DECLARED, and exporting a summary. - [Authoring](https://driftproofhq.com/authoring/): writing a skill and an eval suite worth measuring. - [Judge policy](https://driftproofhq.com/judge-policy/): which model judges, and why it is held fixed. - [Atom feed](https://driftproofhq.com/feed.xml)