How we run it
Frozen four-turn prompt bank (T1–T4), identical across builders, default settings, no manual edits. Run 01 outcomes — Turn 1 ranking: Lovable (best) → v0 → Replit → Bolt. Spend after four turns: Replit about $13, Bolt about 3M tokens, Lovable about 14.2 credits, v0 about $14. The four outputs diverged sharply on turns 2–4 despite identical prompts; screenshots are filed in the runbook and, when supplied, on the atlas card. Formal scored ledger rows stay pending a repeatable rubric and a second reviewer. Full runbook: docs/lab/ai-builder-design-test-v0.1.md.
Reproducibility
Repeat with the same four prompts (T1–T4 verbatim from the runbook), default settings, no manual edits; record builder version, account tier, timestamp, per-turn spend, and one screenshot per turn. The prompt bank is frozen so a third party can re-run and compare divergence and cost against Run 01.