Why this benchmark is useful
It turns code ability into a patch or pass-rate signal you can inspect. That makes it useful for comparing agent harnesses, not just base models.
Benchmarks / Coding
When a coding agent (Cursor Agent, Factory Droid) patches a small repo fixture, does the diff stay in scope, pass tests, obey project rules, and remain reviewable before merge? We grade seven frozen tasks on a diff-review rubric — the failure mode coursework gets burned by when agents touch extra files or ship green tests with unsafe patches.
It turns code ability into a patch or pass-rate signal you can inspect. That makes it useful for comparing agent harnesses, not just base models.
Read Agent Diff Review (VerdictPal Lab) as a active signal with low contamination risk. Compare models only when the source uses the same harness, prompting setup, sampling policy, and score unit.
No results published yet.
The protocol is public before the run ships.
Protocol v0.1 is frozen; results are pending the first VerdictPal Lab desk run — the atlas card is the protocol, not a padded leaderboard.
No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.
Frozen protocol v0.1 (2026-07-06). Seven tasks (T01–T07) on fixture repo v0.1. Pre-run: verify Privacy Mode ON (Cursor) or documented no-training data-use (Factory); screenshot settings. Per task: agent produces a diff; two reviewers grade independently on test-pass, scope, safety, rules, reviewability; disagreements resolved on the record. Post-run: log files touched, `npm test` / `npm run lint` output, and quota burn. Full runbook: docs/lab/agent-diff-review-v0.1.md.
The Pack · Editorial newsletter
One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.