Benchmarks / Coding

Agent Diff Review (VerdictPal Lab)

When a coding agent (Cursor Agent, Factory Droid) patches a small repo fixture, does the diff stay in scope, pass tests, obey project rules, and remain reviewable before merge? We grade seven frozen tasks on a diff-review rubric — the failure mode coursework gets burned by when agents touch extra files or ship green tests with unsafe patches.

CodingDraftActiveLow contaminationVerdictPal LabSince 2026
What this does not measure
  • Long-horizon SWE-bench issue resolution — tasks are deliberately small (single-session fixes on a snapshot repo), not multi-hour repo archaeology.
  • Model intelligence in isolation — harness, default model routing, auto-run settings, and quota throttling can move scores as much as the underlying model.
  • Tab completion or chat quality — only agent-produced diffs on the frozen task battery count; vibe-coded prose is ignored.
Analysis

Why this benchmark is useful

It turns code ability into a patch or pass-rate signal you can inspect. That makes it useful for comparing agent harnesses, not just base models.

Scope

Coverage map

Task family
Coding
Format
A frozen TypeScript snapshot repo (fixture v0.1) with seven small tasks (T01–T07): two bug fixes, one refactor, one test-add, one lint autofix, one rules-compliance pass, one integration wiring. Each tool gets the same prompt on a clean checkout; agents may use default settings only. Privacy Mode (Cursor) or equivalent no-training data-use settings (Factory) must be verified and screenshot before the first task.
Scoring
Safe-pass rate = tasks passing all five rubric dimensions / 7, plus per-dimension fail counts (test-pass, scope, safety, rules, reviewability). A task fails if any dimension fails.
Maintainer
VerdictPal desk
Reading guide

How to read the scores

Read Agent Diff Review (VerdictPal Lab) as a active signal with low contamination. Compare models only when the source uses the same harness, prompting setup, sampling policy, and score unit.

Blind spots

What it does not cover

  • Long-horizon SWE-bench issue resolution — tasks are deliberately small (single-session fixes on a snapshot repo), not multi-hour repo archaeology.
  • Model intelligence in isolation — harness, default model routing, auto-run settings, and quota throttling can move scores as much as the underlying model.
  • Tab completion or chat quality — only agent-produced diffs on the frozen task battery count; vibe-coded prose is ignored.
  • It is a small, hand-graded set — treat early numbers as directional, not definitive.
Scores

Evidence ledger

0 rows

No results published yet.

The protocol is public before the run ships.

Method

What it covers

Protocol v0.1 is frozen; results are pending the first VerdictPal Lab desk run — the atlas card is the protocol, not a padded leaderboard.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

How we run it

Frozen protocol v0.1 (2026-07-06). Seven tasks (T01–T07) on fixture repo v0.1. Pre-run: verify Privacy Mode ON (Cursor) or documented no-training data-use (Factory); screenshot settings. Per task: agent produces a diff; two reviewers grade independently on test-pass, scope, safety, rules, reviewability; disagreements resolved on the record. Post-run: log files touched, `npm test` / `npm run lint` output, and quota burn. Full runbook: docs/lab/agent-diff-review-v0.1.md.

Reproducibility
Fixture repo v0.1 (commit hash recorded), task prompts (T01–T07), run date, tool version, model routing, Privacy Mode screenshot, and settings are logged so a third party can repeat the run. Tools: Cursor (Agent, Privacy Mode on), Factory (Droid CLI or desktop, default auto-run off unless noted).

Tools that report it

Receipts

Sources and further reading

Questions about this benchmark

Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.

What does Agent Diff Review (VerdictPal Lab) measure?

When a coding agent (Cursor Agent, Factory Droid) patches a small repo fixture, does the diff stay in scope, pass tests, obey project rules, and remain reviewable before merge? We grade seven frozen tasks on a diff-review rubric — the failure mode coursework gets burned by when agents touch extra files or ship green tests with unsafe patches.

What does a high Agent Diff Review (VerdictPal Lab) score not prove?

A strong Agent Diff Review (VerdictPal Lab) result says nothing about:

  • Long-horizon SWE-bench issue resolution — tasks are deliberately small (single-session fixes on a snapshot repo), not multi-hour repo archaeology.
  • Model intelligence in isolation — harness, default model routing, auto-run settings, and quota throttling can move scores as much as the underlying model.
  • Tab completion or chat quality — only agent-produced diffs on the frozen task battery count; vibe-coded prose is ignored.
  • It is a small, hand-graded set — treat early numbers as directional, not definitive.

How is Agent Diff Review (VerdictPal Lab) scored?

Safe-pass rate = tasks passing all five rubric dimensions / 7, plus per-dimension fail counts (test-pass, scope, safety, rules, reviewability). A task fails if any dimension fails. Task format: A frozen TypeScript snapshot repo (fixture v0.1) with seven small tasks (T01–T07): two bug fixes, one refactor, one test-add, one lint autofix, one rules-compliance pass, one integration wiring. Each tool gets the same prompt on a clean checkout; agents may use default settings only. Privacy Mode (Cursor) or equivalent no-training data-use settings (Factory) must be verified and screenshot before the first task.

Is Agent Diff Review (VerdictPal Lab) saturated?

Agent Diff Review (VerdictPal Lab) is currently marked Active in the atlas.

Can Agent Diff Review (VerdictPal Lab) results be contaminated by training data?

Contamination risk for Agent Diff Review (VerdictPal Lab) is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.