Benchmarks / Agentic

Vending-Bench 2

Long-horizon business coherence: a model runs a simulated vending-machine company for 365 days — inventory, supplier email, pricing, refunds — scored on final bank balance (USD).

What this does not measure
  • Single-shot tool calls — runs exceed 20M output tokens; the test is sustained judgment.
  • Physical robotics — everything is email and ledger simulation.
  • Human ceiling — Andon estimates ~$63k/year for a strong human operator; models reach a fraction.
Analysis

Why this benchmark is useful

When you care about sustained business judgment over hundreds of simulated days — inventory, suppliers, pricing — not single-shot tool calls.

Scope

Coverage map

Task family
Agentic
Format
Start $500; $2/day location fee; 365 simulated days; mean final balance across 5 runs.
Scoring
Final bank balance in USD (higher is better).
Maintainer
Andon Labs
Reading guide

How to read the scores

Mean final bank balance in USD across runs (higher is better). Start capital and daily fees are fixed in the harness; this is ledger simulation, not robotics or a human ceiling claim.

Blind spots

What it does not cover

  • Single-shot tool calls — runs exceed 20M output tokens; the test is sustained judgment.
  • Physical robotics — everything is email and ledger simulation.
  • Human ceiling — Andon estimates ~$63k/year for a strong human operator; models reach a fraction.
Scores

Evidence ledger

10 rows
1010 rows
11181.87best score
10source-checked
1sources
2026-08-04source date
Vending-Bench10

official-page · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Claude Opus 511181.87 · Andon Labs Vending-Bench 2
  2. Claude Opus 4.710936.76 · Andon Labs Vending-Bench 2
  3. GPT-5.6 Sol9619.37 · Andon Labs Vending-Bench 2
  4. GLM 5.28313.78 · Andon Labs Vending-Bench 2
  5. Claude Opus 4.68017.59 · Andon Labs Vending-Bench 2

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Claude Opus 5Anthropic2026-07-2411181.87
Andon Labs Vending-Bench 22026-08-04
Source-checked
2Claude Opus 4.7Anthropic2026-05-0110936.76
Andon Labs Vending-Bench 22026-08-04
Source-checked
3GPT-5.6 SolOpenAI2026-07-099619.37
Andon Labs Vending-Bench 22026-08-04
Source-checked
4GLM 5.2Zhipu2026-05-158313.78
Andon Labs Vending-Bench 22026-08-04
Source-checked
5Claude Opus 4.6Anthropic2026-03-098017.59
Andon Labs Vending-Bench 22026-08-04
Source-checked
6GPT 5.5OpenAI2026-04-237523.84
Andon Labs Vending-Bench 22026-08-04
Source-checked
7GPT-5.6 TerraOpenAI2026-07-097343.21
Andon Labs Vending-Bench 22026-08-04
Source-checked
8Claude Sonnet 4.6Anthropic2026-02-177204.14
Andon Labs Vending-Bench 22026-08-04
Source-checked
9Muse Spark 1.1Meta2026-07-096520.47
Andon Labs Vending-Bench 22026-08-04
Source-checked
10Claude Sonnet 5Anthropic2026-06-306377.7
Andon Labs Vending-Bench 22026-08-04
Source-checked
Method

What it covers

This benchmark sits in the agentic family. It uses Start $500; $2/day location fee; 365 simulated days; mean final balance across 5 runs. The score should travel with its task format, scoring method, source date, and benchmark version.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

Receipts

Sources and further reading

Questions about this benchmark

Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.

What does Vending-Bench 2 measure?

Long-horizon business coherence: a model runs a simulated vending-machine company for 365 days — inventory, supplier email, pricing, refunds — scored on final bank balance (USD).

What does a high Vending-Bench 2 score not prove?

A strong Vending-Bench 2 result says nothing about:

  • Single-shot tool calls — runs exceed 20M output tokens; the test is sustained judgment.
  • Physical robotics — everything is email and ledger simulation.
  • Human ceiling — Andon estimates ~$63k/year for a strong human operator; models reach a fraction.

How is Vending-Bench 2 scored?

Final bank balance in USD (higher is better). Task format: Start $500; $2/day location fee; 365 simulated days; mean final balance across 5 runs.

Is Vending-Bench 2 saturated?

Vending-Bench 2 is currently marked Active in the atlas.

Can Vending-Bench 2 results be contaminated by training data?

Contamination risk for Vending-Bench 2 is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.