Benchmarks / Agentic

AI Builder Design Test (VerdictPal Lab)

Vibecoder Design Test

Give four AI website builders the same multi-turn design prompts and watch how far apart the results land — and what it costs to get there. We run identical prompts through Lovable, v0, Replit, and Bolt and grade the output variation, design quality, and spend.

AgenticDraftActiveLow contamination riskVerdictPal LabSince 2026
What this does not measure
  • Model quality in isolation — each builder wraps a model, a scaffold, and tooling. The output is the product, not the model.
  • General coding ability — this is a visual design task, not a correctness or refactor benchmark.
  • Reproducibility across accounts and tiers — each run used the tester's own plan and defaults; vendors may route different models or scaffolds on other tiers.
Analysis

Why this benchmark is useful

A score on a coding benchmark tells you what a model can do in a harness. It does not tell you what ships when a human types a vague design brief into a product. This test holds the prompt constant and lets the product vary — the variable readers actually face when they pick a builder.

Scope

Coverage map

Task family
Agentic
Format
Four identical prompts, sent in sequence to each builder with default settings and no manual code edits. Each turn escalates: design masterpiece in the style of a reference plant site → add mystery, paper gradients, motion → refine to competition-winning quality with flawless function and legible type → fully animated brand experience, clutter removed. The only variable is which builder receives the prompt. Full prompt text lives in the runbook.
Scoring
Editorial. After each turn the tester ranks the four outputs 1–4 on visual design, and records the builder's reported spend (credits, tokens, or USD, depending on the plan). No normalized 0–100 score is published from a single run; the record carries the directional ranking and the raw spend.
Maintainer
VerdictPal desk
Reading guide

How to read the scores

Read the Turn 1 ranking as a first-impression design verdict, and the final spend as a cost-of-arrival signal — not as a quality score. The four outputs diverged sharply on later turns despite identical prompts; that divergence, not the rank order, is the finding.

Blind spots

What it does not cover

  • Model quality in isolation — each builder wraps a model, a scaffold, and tooling. The output is the product, not the model.
  • General coding ability — this is a visual design task, not a correctness or refactor benchmark.
  • Reproducibility across accounts and tiers — each run used the tester's own plan and defaults; vendors may route different models or scaffolds on other tiers.
  • Statistical significance — one tester, one pass per builder. Treat the ranking as directional, not definitive.
Scores

Evidence ledger

0 rows

No results published yet.

The protocol is public before the run ships.

Method

What it covers

Run 01 (2026-07-12) is a single-judge editorial pass with four builders. Spend is recorded in each vendor's native unit — Lovable in credits, Bolt in tokens, Replit and v0 in USD — and is not normalized, because plans and tiers differ. One Turn 1 plan detail was unclear at capture time and is flagged in the runbook. Rankings are directional, not definitive; a repeatable rubric and a second reviewer are the gate to formal scored rows.

Score ceiling

Where it breaks down

There is no normalized 0–100 score. The ranking is a single judge's editorial order after Turn 1; spend is recorded in each vendor's native unit (credits, tokens, or USD) and is not converted, because plans and tiers are not comparable one-for-one.

How we run it

Frozen four-turn prompt bank (T1–T4), identical across builders, default settings, no manual edits. Run 01 outcomes — Turn 1 ranking: Lovable (best) → v0 → Replit → Bolt. Spend after four turns: Replit about $13, Bolt about 3M tokens, Lovable about 14.2 credits, v0 about $14. The four outputs diverged sharply on turns 2–4 despite identical prompts; screenshots are filed in the runbook and, when supplied, on the atlas card. Formal scored ledger rows stay pending a repeatable rubric and a second reviewer. Full runbook: docs/lab/ai-builder-design-test-v0.1.md.

Reproducibility
Repeat with the same four prompts (T1–T4 verbatim from the runbook), default settings, no manual edits; record builder version, account tier, timestamp, per-turn spend, and one screenshot per turn. The prompt bank is frozen so a third party can re-run and compare divergence and cost against Run 01.

Tools that report it

Receipts

Sources and further reading

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.