Why this benchmark is useful
It turns code ability into a patch or pass-rate signal you can inspect. That makes it useful for comparing agent harnesses, not just base models.
When a coding agent (Cursor Agent, Factory Droid) patches a small repo fixture, does the diff stay in scope, pass tests, obey project rules, and remain reviewable before merge? We grade seven frozen tasks on a diff-review rubric — the failure mode coursework gets burned by when agents touch extra files or ship green tests with unsafe patches.
It turns code ability into a patch or pass-rate signal you can inspect. That makes it useful for comparing agent harnesses, not just base models.
Read Agent Diff Review (VerdictPal Lab) as a active signal with low contamination. Compare models only when the source uses the same harness, prompting setup, sampling policy, and score unit.
No results published yet.
The protocol is public before the run ships.
Protocol v0.1 is frozen; results are pending the first VerdictPal Lab desk run — the atlas card is the protocol, not a padded leaderboard.
No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.
Frozen protocol v0.1 (2026-07-06). Seven tasks (T01–T07) on fixture repo v0.1. Pre-run: verify Privacy Mode ON (Cursor) or documented no-training data-use (Factory); screenshot settings. Per task: agent produces a diff; two reviewers grade independently on test-pass, scope, safety, rules, reviewability; disagreements resolved on the record. Post-run: log files touched, `npm test` / `npm run lint` output, and quota burn. Full runbook: docs/lab/agent-diff-review-v0.1.md.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
When a coding agent (Cursor Agent, Factory Droid) patches a small repo fixture, does the diff stay in scope, pass tests, obey project rules, and remain reviewable before merge? We grade seven frozen tasks on a diff-review rubric — the failure mode coursework gets burned by when agents touch extra files or ship green tests with unsafe patches.
A strong Agent Diff Review (VerdictPal Lab) result says nothing about:
Safe-pass rate = tasks passing all five rubric dimensions / 7, plus per-dimension fail counts (test-pass, scope, safety, rules, reviewability). A task fails if any dimension fails. Task format: A frozen TypeScript snapshot repo (fixture v0.1) with seven small tasks (T01–T07): two bug fixes, one refactor, one test-add, one lint autofix, one rules-compliance pass, one integration wiring. Each tool gets the same prompt on a clean checkout; agents may use default settings only. Privacy Mode (Cursor) or equivalent no-training data-use settings (Factory) must be verified and screenshot before the first task.
Agent Diff Review (VerdictPal Lab) is currently marked Active in the atlas.
Contamination risk for Agent Diff Review (VerdictPal Lab) is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.