VerdictPal · editorial desk · 2026VerdictPal
The Lab

Benchmarks we run ourselves.

Vendor scores are marketing-adjacent. The Lab is where the desk runs its own tests — on the failure modes that actually burn students and researchers — and publishes the method, prompts, and raw examples so you can repeat the run.

Honest by default

A Lab benchmark ships a score only after human grading and sign-off. Until then you see the protocol and run progress — not a padded leaderboard.

Citation Fidelity v0.1: 156/156 captures complete. Human grading and sign-off still required before scores ship — desk runbook

Agent Diff Review (VerdictPal Lab)

DraftProtocol v0.17 questions
Full protocol →

When a coding agent (Cursor Agent, Factory Droid) patches a small repo fixture, does the diff stay in scope, pass tests, obey project rules, and remain reviewable before merge? We grade seven frozen tasks on a diff-review rubric — the failure mode coursework gets burned by when agents touch extra files or ship green tests with unsafe patches.

Results pending

Protocol is frozen and published. Score rows ship only after the desk completes an end-to-end run — no padded leaderboard until then.

What it won't tell you

  • Long-horizon SWE-bench issue resolution — tasks are deliberately small (single-session fixes on a snapshot repo), not multi-hour repo archaeology.
  • Model intelligence in isolation — harness, default model routing, auto-run settings, and quota throttling can move scores as much as the underlying model.
  • Tab completion or chat quality — only agent-produced diffs on the frozen task battery count; vibe-coded prose is ignored.
  • It is a small, hand-graded set — treat early numbers as directional, not definitive.
Method
Frozen protocol v0.1 (2026-07-06). Seven tasks (T01–T07) on fixture repo v0.1. Pre-run: verify Privacy Mode ON (Cursor) or documented no-training data-use (Factory); screenshot settings. Per task: agent produces a diff; two reviewers grade independently on test-pass, scope, safety, rules, reviewability; disagreements resolved on the record. Post-run: log files touched, `npm test` / `npm run lint` output, and quota burn. Full runbook: docs/lab/agent-diff-review-v0.1.md.

Agent Diff Review v0.1 runbook · Desk runbook (repo) · methodology: /methodology

AI Builder Design Test (VerdictPal Lab)

DraftProtocol v0.1
Full protocol →

Give four AI website builders the same multi-turn design prompts and watch how far apart the results land — and what it costs to get there. We run identical prompts through Lovable, v0, Replit, and Bolt and grade the output variation, design quality, and spend.

Results pending

Protocol is frozen and published. Score rows ship only after the desk completes an end-to-end run — no padded leaderboard until then.

What it won't tell you

  • Model quality in isolation — each builder wraps a model, a scaffold, and tooling. The output is the product, not the model.
  • General coding ability — this is a visual design task, not a correctness or refactor benchmark.
  • Reproducibility across accounts and tiers — each run used the tester's own plan and defaults; vendors may route different models or scaffolds on other tiers.
  • Statistical significance — one tester, one pass per builder. Treat the ranking as directional, not definitive.
Method
Frozen four-turn prompt bank (T1–T4), identical across builders, default settings, no manual edits. Run 01 outcomes — Turn 1 ranking: Lovable (best) → v0 → Replit → Bolt. Spend after four turns: Replit about $13, Bolt about 3M tokens, Lovable about 14.2 credits, v0 about $14. The four outputs diverged sharply on turns 2–4 despite identical prompts; screenshots are filed in the runbook and, when supplied, on the atlas card. Formal scored ledger rows stay pending a repeatable rubric and a second reviewer. Full runbook: docs/lab/ai-builder-design-test-v0.1.md.

AI Builder Design Test v0.1 runbook · Desk runbook (repo) · methodology: /methodology

Citation Fidelity (VerdictPal Lab)

DraftProtocol v0.112 questions
Full protocol →

When an answer-engine cites a source, does the source actually say what the answer claims? We grade real research questions on whether every cited claim is supported by the linked page — the failure mode students get burned by most.

Captures complete — grading pending

All capture slots are filled. Score rows appear only after independent human grading and desk-lead sign-off. No padded leaderboard until then.

What it won't tell you

  • Answer quality or fluency — a beautifully written answer with a fabricated citation fails here on purpose.
  • Coverage — we test whether what is cited is true, not whether the model found every relevant source.
  • It is a small, hand-graded set — treat early numbers as directional, not definitive.
Method
Frozen protocol v0.1 (2026-07-06). Twelve questions (Q01–Q12): three policy brief, three product comparison, three academic-adjacent, three local/current. Each engine gets the same prompt and default settings; two reviewers grade every cited claim independently (supported / unsupported / broken-link); disagreements are resolved on the record. Full runbook: docs/lab/citation-fidelity-v0.1.md.

Citation Fidelity v0.1 runbook · Desk runbook (repo) · methodology: /methodology

← All benchmarks