Benchmarks / Coding

Agent Diff Review (VerdictPal Lab)

When a coding agent (Cursor Agent, Factory Droid) patches a small repo fixture, does the diff stay in scope, pass tests, obey project rules, and remain reviewable before merge? We grade seven frozen tasks on a diff-review rubric — the failure mode coursework gets burned by when agents touch extra files or ship green tests with unsafe patches.

CodingDraftActiveLow contamination riskVerdictPal LabSince 2026
What this does not measure
  • Long-horizon SWE-bench issue resolution — tasks are deliberately small (single-session fixes on a snapshot repo), not multi-hour repo archaeology.
  • Model intelligence in isolation — harness, default model routing, auto-run settings, and quota throttling can move scores as much as the underlying model.
  • Tab completion or chat quality — only agent-produced diffs on the frozen task battery count; vibe-coded prose is ignored.
Analysis

Why this benchmark is useful

It turns code ability into a patch or pass-rate signal you can inspect. That makes it useful for comparing agent harnesses, not just base models.

Scope

Coverage map

Task family
Coding
Format
A frozen TypeScript snapshot repo (fixture v0.1) with seven small tasks (T01–T07): two bug fixes, one refactor, one test-add, one lint autofix, one rules-compliance pass, one integration wiring. Each tool gets the same prompt on a clean checkout; agents may use default settings only. Privacy Mode (Cursor) or equivalent no-training data-use settings (Factory) must be verified and screenshot before the first task.
Scoring
Safe-pass rate = tasks passing all five rubric dimensions / 7, plus per-dimension fail counts (test-pass, scope, safety, rules, reviewability). A task fails if any dimension fails.
Maintainer
VerdictPal desk
Reading guide

How to read the scores

Read Agent Diff Review (VerdictPal Lab) as a active signal with low contamination risk. Compare models only when the source uses the same harness, prompting setup, sampling policy, and score unit.

Blind spots

What it does not cover

  • Long-horizon SWE-bench issue resolution — tasks are deliberately small (single-session fixes on a snapshot repo), not multi-hour repo archaeology.
  • Model intelligence in isolation — harness, default model routing, auto-run settings, and quota throttling can move scores as much as the underlying model.
  • Tab completion or chat quality — only agent-produced diffs on the frozen task battery count; vibe-coded prose is ignored.
  • It is a small, hand-graded set — treat early numbers as directional, not definitive.
Scores

Evidence ledger

0 rows

No results published yet.

The protocol is public before the run ships.

Method

What it covers

Protocol v0.1 is frozen; results are pending the first VerdictPal Lab desk run — the atlas card is the protocol, not a padded leaderboard.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

How we run it

Frozen protocol v0.1 (2026-07-06). Seven tasks (T01–T07) on fixture repo v0.1. Pre-run: verify Privacy Mode ON (Cursor) or documented no-training data-use (Factory); screenshot settings. Per task: agent produces a diff; two reviewers grade independently on test-pass, scope, safety, rules, reviewability; disagreements resolved on the record. Post-run: log files touched, `npm test` / `npm run lint` output, and quota burn. Full runbook: docs/lab/agent-diff-review-v0.1.md.

Reproducibility
Fixture repo v0.1 (commit hash recorded), task prompts (T01–T07), run date, tool version, model routing, Privacy Mode screenshot, and settings are logged so a third party can repeat the run. Tools: Cursor (Agent, Privacy Mode on), Factory (Droid CLI or desktop, default auto-run off unless noted).

Tools that report it

Receipts

Sources and further reading

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.