Why this benchmark is useful
When you care about sustained business judgment over hundreds of simulated days — inventory, suppliers, pricing — not single-shot tool calls.
Long-horizon business coherence: a model runs a simulated vending-machine company for 365 days — inventory, supplier email, pricing, refunds — scored on final bank balance (USD).
When you care about sustained business judgment over hundreds of simulated days — inventory, suppliers, pricing — not single-shot tool calls.
Mean final bank balance in USD across runs (higher is better). Start capital and daily fees are fixed in the harness; this is ledger simulation, not robotics or a human ceiling claim.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Claude Opus 5Anthropic | 2026-07-24 | 11181.87 | Andon Labs Vending-Bench 22026-08-04 | Source-checked |
| 2 | Claude Opus 4.7Anthropic | 2026-05-01 | 10936.76 | Andon Labs Vending-Bench 22026-08-04 | Source-checked |
| 3 | GPT-5.6 SolOpenAI | 2026-07-09 | 9619.37 | Andon Labs Vending-Bench 22026-08-04 | Source-checked |
| 4 | GLM 5.2Zhipu | 2026-05-15 | 8313.78 | Andon Labs Vending-Bench 22026-08-04 | Source-checked |
| 5 | Claude Opus 4.6Anthropic | 2026-03-09 | 8017.59 | Andon Labs Vending-Bench 22026-08-04 | Source-checked |
| 6 | GPT 5.5OpenAI | 2026-04-23 | 7523.84 | Andon Labs Vending-Bench 22026-08-04 | Source-checked |
| 7 | GPT-5.6 TerraOpenAI | 2026-07-09 | 7343.21 | Andon Labs Vending-Bench 22026-08-04 | Source-checked |
| 8 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 7204.14 | Andon Labs Vending-Bench 22026-08-04 | Source-checked |
| 9 | Muse Spark 1.1Meta | 2026-07-09 | 6520.47 | Andon Labs Vending-Bench 22026-08-04 | Source-checked |
| 10 | Claude Sonnet 5Anthropic | 2026-06-30 | 6377.7 | Andon Labs Vending-Bench 22026-08-04 | Source-checked |
This benchmark sits in the agentic family. It uses Start $500; $2/day location fee; 365 simulated days; mean final balance across 5 runs. The score should travel with its task format, scoring method, source date, and benchmark version.
No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
Long-horizon business coherence: a model runs a simulated vending-machine company for 365 days — inventory, supplier email, pricing, refunds — scored on final bank balance (USD).
A strong Vending-Bench 2 result says nothing about:
Final bank balance in USD (higher is better). Task format: Start $500; $2/day location fee; 365 simulated days; mean final balance across 5 runs.
Vending-Bench 2 is currently marked Active in the atlas.
Contamination risk for Vending-Bench 2 is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.