# What each evaluation tool measures

Canonical: <https://driftproofhq.com/compare/>

What each tool measures, cell by cell, read from that tool's own docs or code and linked to where it says so. No tool is placed above another.

| Aspect | [claude plugin eval (Claude Code)](https://code.claude.com/docs/en/plugin-evals) | [alibaba/skill-up](https://github.com/alibaba/skill-up) | [NVIDIA SkillEvaluator](https://github.com/NVIDIA/SkillEvaluator) | [agent-skills-eval](https://github.com/darkrishabh/agent-skills-eval) | [Driftproof](https://github.com/driftproofhq/driftproof) |
| --- | --- | --- | --- | --- | --- |
| What it measures | How reliably a plugin (including a skill shipped as a plugin) leads Claude to the right outcome on eval cases, and what it contributes against a no-plugin baseline ([source, read 30 September 2026](https://code.claude.com/docs/en/plugin-evals)) | A Skill's behaviour across declarative cases (response, tool calls and workspace changes), optionally compared with and without the Skill; it also evaluates agents and workspaces without a Skill ([source, read 30 September 2026](https://github.com/alibaba/skill-up/blob/7f1ff9b8e2d7c654728de526867f2f7e7b78ea51/README.md#L54); [source, read 30 September 2026](https://github.com/alibaba/skill-up/blob/7f1ff9b8e2d7c654728de526867f2f7e7b78ea51/README.md#L52)) | Three tiers: static validation (safety, well-formedness), deduplication (repeated or overlapping guidance), and live evaluation of whether the skill helps a real agent, reported on five dimensions (security, correctness, discoverability, effectiveness, efficiency) ([source, read 30 September 2026](https://github.com/NVIDIA/SkillEvaluator/blob/d0157411fe6c6eb949f832983bf671dfd1b9c701/README.md#L14); [source, read 30 September 2026](https://github.com/NVIDIA/SkillEvaluator/blob/d0157411fe6c6eb949f832983bf671dfd1b9c701/docs/tier3-live-evaluation.mdx#L7)) | Whether a SKILL.md makes a model better at the task on the skill's evals/evals.json prompts, judged against declared assertions ([source, read 30 September 2026](https://github.com/darkrishabh/agent-skills-eval/blob/f10ae2b3c93055cadd82c02a3b7fe4cfbd9a4645/README.md#L11)) | Not shown |
| How it scores | Pass/fail graders (regex, tool_used, tool_order, file_exists, llm judge, baseline judge); a run scores the (optionally weighted) fraction of graders passed; a case scores the mean across runs and passes at --threshold (default [1.0](https://code.claude.com/docs/en/plugin-evals)). An llm grader passes on at least two of three judge votes ([source, read 30 September 2026](https://code.claude.com/docs/en/plugin-evals)) | Per-case assertions graded by rule_based, script or agent_judge (LLM rubric, pass_threshold default [0.7](https://github.com/alibaba/skill-up/blob/7f1ff9b8e2d7c654728de526867f2f7e7b78ea51/docs/guide/writing-evals.md#L810)); a case is PASS when all assertions pass; results roll up to pass_rate ([source, read 30 September 2026](https://github.com/alibaba/skill-up/blob/7f1ff9b8e2d7c654728de526867f2f7e7b78ea51/README.md#L69); [source, read 30 September 2026](https://github.com/alibaba/skill-up/blob/7f1ff9b8e2d7c654728de526867f2f7e7b78ea51/docs/guide/cli-reference.md#L386); [source, read 30 September 2026](https://github.com/alibaba/skill-up/blob/7f1ff9b8e2d7c654728de526867f2f7e7b78ea51/docs/guide/writing-evals.md#L810)) | Five dimension scores from [0.0](https://github.com/NVIDIA/SkillEvaluator/blob/d0157411fe6c6eb949f832983bf671dfd1b9c701/docs/tier3-live-evaluation.mdx#L460) to [1.0](https://github.com/NVIDIA/SkillEvaluator/blob/d0157411fe6c6eb949f832983bf671dfd1b9c701/docs/tier3-live-evaluation.mdx#L460) with fixed bands (PASS at [0.50](https://github.com/NVIDIA/SkillEvaluator/blob/d0157411fe6c6eb949f832983bf671dfd1b9c701/docs/tier3-live-evaluation.mdx#L460) or above, NEUTRAL [0.40](https://github.com/NVIDIA/SkillEvaluator/blob/d0157411fe6c6eb949f832983bf671dfd1b9c701/docs/tier3-live-evaluation.mdx#L460) to below [0.50](https://github.com/NVIDIA/SkillEvaluator/blob/d0157411fe6c6eb949f832983bf671dfd1b9c701/docs/tier3-live-evaluation.mdx#L460), FAIL below [0.40](https://github.com/NVIDIA/SkillEvaluator/blob/d0157411fe6c6eb949f832983bf671dfd1b9c701/docs/tier3-live-evaluation.mdx#L460)); an agent passes only when every dimension passes. Skill Lift (with-skill minus baseline) has its own bands (PASS at [+0.05](https://github.com/NVIDIA/SkillEvaluator/blob/d0157411fe6c6eb949f832983bf671dfd1b9c701/docs/tier3-live-evaluation.mdx#L460) or above, FAIL at [-0.10](https://github.com/NVIDIA/SkillEvaluator/blob/d0157411fe6c6eb949f832983bf671dfd1b9c701/docs/tier3-live-evaluation.mdx#L460) or below) and is diagnostic, not the verdict ([source, read 30 September 2026](https://github.com/NVIDIA/SkillEvaluator/blob/d0157411fe6c6eb949f832983bf671dfd1b9c701/docs/tier3-live-evaluation.mdx#L460); [source, read 30 September 2026](https://github.com/NVIDIA/SkillEvaluator/blob/d0157411fe6c6eb949f832983bf671dfd1b9c701/docs/tier3-live-evaluation.mdx#L470)) | A judge model grades each output pass/fail against the eval's assertions (expected_output is promoted to an assertion when none are given); deterministic tool_assertions are also available; results roll up to pass_rate ([source, read 30 September 2026](https://github.com/darkrishabh/agent-skills-eval/blob/f10ae2b3c93055cadd82c02a3b7fe4cfbd9a4645/README.md#L65); [source, read 30 September 2026](https://github.com/darkrishabh/agent-skills-eval/blob/f10ae2b3c93055cadd82c02a3b7fe4cfbd9a4645/README.md#L118)) | An LLM judge scores each response against rubrics anchored at [0.80](https://github.com/driftproofhq/driftproof/blob/v0.13.0/README.md#L83); the lift is the with-skill mean minus the baseline mean; the verdict is PASSED, NO_EFFECT, REGRESSED or UNDERPOWERED against a [0.05](https://github.com/driftproofhq/driftproof/blob/v0.13.0/README.md#L103) effect floor, and a case is called improved or regressed only when the bands do not overlap and the move clears the floor ([source, read 1 October 2026](https://github.com/driftproofhq/driftproof/blob/v0.13.0/README.md#L382); [source, read 1 October 2026](https://github.com/driftproofhq/driftproof/blob/v0.13.0/README.md#L83); [source, read 1 October 2026](https://github.com/driftproofhq/driftproof/blob/v0.13.0/README.md#L104); [source, read 30 September 2026](https://driftproofhq.com/methodology/); [source, read 1 October 2026](https://github.com/driftproofhq/driftproof/blob/v0.13.0/README.md#L103)) |
| With and without the skill | Yes, by default: a with-plugin arm and a no-plugin arm, reported as WITH, W/OUT and delta; --ablation none turns it off ([source, read 30 September 2026](https://code.claude.com/docs/en/plugin-evals)) | Yes, opt-in: benchmark mode (benchmark.enabled: true or --baseline) runs each case with_skill and without_skill and reports the delta; disabled by default ([source, read 30 September 2026](https://github.com/alibaba/skill-up/blob/7f1ff9b8e2d7c654728de526867f2f7e7b78ea51/docs/guide/writing-evals.md#L865); [source, read 30 September 2026](https://github.com/alibaba/skill-up/blob/7f1ff9b8e2d7c654728de526867f2f7e7b78ea51/docs/guide/writing-evals.md#L879)) | Yes, by default in Tier [3](https://github.com/NVIDIA/SkillEvaluator/blob/d0157411fe6c6eb949f832983bf671dfd1b9c701/docs/tier3-live-evaluation.mdx#L17): a with-skill arm and a baseline arm in a sandbox; --skip-baseline removes the baseline and the lift ([source, read 30 September 2026](https://github.com/NVIDIA/SkillEvaluator/blob/d0157411fe6c6eb949f832983bf671dfd1b9c701/docs/tier3-live-evaluation.mdx#L17)) | Yes, opt-in: --baseline (config baseline: true) runs each eval with_skill and without_skill; the default is false ([source, read 30 September 2026](https://github.com/darkrishabh/agent-skills-eval/blob/f10ae2b3c93055cadd82c02a3b7fe4cfbd9a4645/README.md#L118); [source, read 30 September 2026](https://github.com/darkrishabh/agent-skills-eval/blob/f10ae2b3c93055cadd82c02a3b7fe4cfbd9a4645/README.md#L173)) | Yes, always: each case is run with the SKILL.md supplied and without it (bare baseline) ([source, read 30 September 2026](https://driftproofhq.com/methodology/)) |
| Repeats and spread | Not shown | Repeats: yes, --iteration N runs N samples for stability or flakiness sampling. Spread: benchmark.json carries mean and stddev of pass_rate, time and tokens; in the code the pass_rate stddev is taken over the per-case pass/fail values of a run, and the Anthropic-compatible benchmark stamps runs_per_configuration: [1](https://github.com/alibaba/skill-up/blob/7f1ff9b8e2d7c654728de526867f2f7e7b78ea51/internal/report/benchmark_anthropic.go#L169) ([source, read 30 September 2026](https://github.com/alibaba/skill-up/blob/7f1ff9b8e2d7c654728de526867f2f7e7b78ea51/docs/guide/cli-reference.md#L31); [source, read 30 September 2026](https://github.com/alibaba/skill-up/blob/7f1ff9b8e2d7c654728de526867f2f7e7b78ea51/internal/report/benchmark.go#L4); [source, read 30 September 2026](https://github.com/alibaba/skill-up/blob/7f1ff9b8e2d7c654728de526867f2f7e7b78ea51/internal/report/benchmark.go#L104); [source, read 30 September 2026](https://github.com/alibaba/skill-up/blob/7f1ff9b8e2d7c654728de526867f2f7e7b78ea51/internal/report/benchmark_anthropic.go#L169)) | Repeats: optional, --n-attempts (default [1](https://github.com/NVIDIA/SkillEvaluator/blob/d0157411fe6c6eb949f832983bf671dfd1b9c701/docs/tier3-live-evaluation.mdx#L417)) runs each case k times per arm and adds pass@k. Spread: each arm records a case-level [95](https://github.com/NVIDIA/SkillEvaluator/blob/d0157411fe6c6eb949f832983bf671dfd1b9c701/docs/reports.mdx#L295)% Wilson score interval for its pass rate, and paired cases get an exact McNemar diagnostic ([source, read 30 September 2026](https://github.com/NVIDIA/SkillEvaluator/blob/d0157411fe6c6eb949f832983bf671dfd1b9c701/docs/tier3-live-evaluation.mdx#L417); [source, read 30 September 2026](https://github.com/NVIDIA/SkillEvaluator/blob/d0157411fe6c6eb949f832983bf671dfd1b9c701/docs/reports.mdx#L295)) | Not shown | Yes: since receipt spec v0.5 each arm is generated at least [3](https://github.com/driftproofhq/driftproof/blob/v0.13.0/README.md#L108) and at most [10](https://github.com/driftproofhq/driftproof/blob/v0.13.0/README.md#L108) times and each draw is judged several times; every score carries a band (mean plus or minus one sample standard deviation), and cases that could not be separated at the effect floor are reported as "Not enough draws to conclude at this effect floor" ([source, read 1 October 2026](https://github.com/driftproofhq/driftproof/blob/v0.13.0/README.md#L108); [source, read 1 October 2026](https://github.com/driftproofhq/driftproof/blob/v0.13.0/spec/RECEIPT.md#L6); [source, read 1 October 2026](https://github.com/driftproofhq/driftproof/blob/v0.13.0/spec/RECEIPT.md#L150)) |
| Across model releases | Not shown | Not shown | Not shown | Not shown | Yes: receipts are bound to one model version, and driftproof diff compares two receipts across a model release into a drift report ([source, read 1 October 2026](https://github.com/driftproofhq/driftproof/blob/v0.13.0/README.md#L161); [source, read 1 October 2026](https://github.com/driftproofhq/driftproof/blob/v0.13.0/spec/RECEIPT.md#L8)) |
| What it writes | A terminal summary table; results/<timestamp>/aggregate-result.json (schemaVersion [1](https://code.claude.com/docs/en/plugin-evals)) and a self-contained report.html; --json output; optional publication as a private claude.ai artifact; exit codes for CI ([source, read 30 September 2026](https://code.claude.com/docs/en/plugin-evals)) | grading.json, benchmark.json and benchmark.md (Anthropic-compatible), result.json, JUnit XML and HTML reports, in <skill-name>-workspace/iteration-N/ ([source, read 30 September 2026](https://github.com/alibaba/skill-up/blob/7f1ff9b8e2d7c654728de526867f2f7e7b78ea51/README.md#L70)) | Reports in cli, json, html, markdown and sarif formats (html and json by default); Tier [3](https://github.com/NVIDIA/SkillEvaluator/blob/d0157411fe6c6eb949f832983bf671dfd1b9c701/docs/reports.mdx#L11) also writes a timestamped results tree (scores, trials, lift data) and BENCHMARK.md cards ([source, read 30 September 2026](https://github.com/NVIDIA/SkillEvaluator/blob/d0157411fe6c6eb949f832983bf671dfd1b9c701/docs/reports.mdx#L23); [source, read 30 September 2026](https://github.com/NVIDIA/SkillEvaluator/blob/d0157411fe6c6eb949f832983bf671dfd1b9c701/docs/reports.mdx#L11)) | A workspace per run (iteration-N/) with meta.json, benchmark.json, per-eval with_skill and without_skill outputs, timing and grading.json, a static HTML report, and optional JSONL event logs ([source, read 30 September 2026](https://github.com/darkrishabh/agent-skills-eval/blob/f10ae2b3c93055cadd82c02a3b7fe4cfbd9a4645/README.md#L44); [source, read 30 September 2026](https://github.com/darkrishabh/agent-skills-eval/blob/f10ae2b3c93055cadd82c02a3b7fe4cfbd9a4645/README.md#L377)) | A receipt (JSON, published schema, self-hash) plus a human summary per run; Markdown drift reports; a shields.io badge JSON; GitHub Action outputs (verdict, delta and others) ([source, read 1 October 2026](https://github.com/driftproofhq/driftproof/blob/v0.13.0/README.md#L259); [source, read 1 October 2026](https://github.com/driftproofhq/driftproof/blob/v0.13.0/README.md#L94)) |
| Licence | Not shown | Apache-2.0 ([source, read 30 September 2026](https://github.com/alibaba/skill-up/blob/7f1ff9b8e2d7c654728de526867f2f7e7b78ea51/README.md#L399)) | Apache-2.0 ([source, read 30 September 2026](https://github.com/NVIDIA/SkillEvaluator/blob/d0157411fe6c6eb949f832983bf671dfd1b9c701/README.md#L238)) | MIT ([source, read 30 September 2026](https://github.com/darkrishabh/agent-skills-eval/blob/f10ae2b3c93055cadd82c02a3b7fe4cfbd9a4645/LICENSE#L1)) | Apache-2.0 ([source, read 1 October 2026](https://github.com/driftproofhq/driftproof/blob/v0.13.0/README.md#L681)) |

Not shown: this cell could not be confirmed from that tool's own docs or code as they were read for this page, so it is left out. It does not mean the tool lacks it.

## More answers

- [What is Driftproof?](https://driftproofhq.com/what-is-driftproof/)
- [How to tell whether a skill improved output](https://driftproofhq.com/agent-skill-evaluation/)
- [Does a skill still help after a model upgrade?](https://driftproofhq.com/agent-skill-regression-testing/)
- [The paper](https://driftproofhq.com/paper/)
