Driftproof

A skill that helped last month may not help today.

Driftproof re-runs an agent skill's own eval suite across model versions, samples the judge so every score carries a confidence band, and only claims a regression when the bands actually separate. The phenomenon is the story: verdicts age because the substrate moves.

Read Report #001 → View on GitHub

What it does

For each skill, Driftproof runs its eval suite with the skill and without it, on two model versions. Each response is judged several times, so a case carries a mean ± stddev band rather than one fragile number. Diffing two runs yields a drift report: a case only counts as regressed or improved when its two confidence bands do not overlap. Everything is recorded in a signed, dated receipt whose every hash is reproducible.

Why band-based

An LLM judge is noisy. A tool that cries wolf on that noise is worse than useless. Driftproof's core rule is that a verdict is claimed only on non-overlapping bands — so the honest answer is allowed to be "nothing moved beyond the noise." That restraint is the point.

Latest report

Report #001 — do skill verdicts hold across a Sonnet model release? Ten public, text-representable agent skills, run on a current-vs-previous Sonnet pair with sampled, band-based verdicts.

Roadmap