Benchmarks / Agentic

Vending-Bench 2

Long-horizon business coherence: a model runs a simulated vending-machine company for 365 days — inventory, supplier email, pricing, refunds — scored on final bank balance (USD).

AgenticSolidActiveLow contamination riskSince 2026
What this does not measure
  • Single-shot tool calls — runs exceed 20M output tokens; the test is sustained judgment.
  • Physical robotics — everything is email and ledger simulation.
  • Human ceiling — Andon estimates ~$63k/year for a strong human operator; models reach a fraction.
Analysis

Why this benchmark is useful

Editorial brief pendingWe publish the methodology and ledger first; benchmark-specific analysis ships after desk review.

Scope

Coverage map

Task family
Agentic
Format
Start $500; $2/day location fee; 365 simulated days; mean final balance across 5 runs.
Scoring
Final bank balance in USD (higher is better).
Maintainer
Andon Labs
Reading guide

How to read the scores

Reading guide pending. Use task format, scoring method, and source dates in the ledger until the desk brief ships.

Blind spots

What it does not cover

  • Single-shot tool calls — runs exceed 20M output tokens; the test is sustained judgment.
  • Physical robotics — everything is email and ledger simulation.
  • Human ceiling — Andon estimates ~$63k/year for a strong human operator; models reach a fraction.
Scores

Evidence ledger

5 rows
55 rows
10936.76best score
5source-checked
1sources
2026-06-01source date
Vending-Bench5

official-page · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Claude Opus 4.710936.76 · Andon Labs Vending-Bench 2
  2. GLM 5.28313.78 · Andon Labs Vending-Bench 2
  3. Claude Opus 4.68017.59 · Andon Labs Vending-Bench 2
  4. GPT 5.57523.84 · Andon Labs Vending-Bench 2
  5. Claude Fable 5 (high)5680.26 · Andon Labs Vending-Bench 2

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Claude Opus 4.7Anthropic2026-05-0110936.76
Andon Labs Vending-Bench 22026-06-01
Source-checked
2GLM 5.2Zhipu2026-05-158313.78
Andon Labs Vending-Bench 22026-06-01
Source-checked
3Claude Opus 4.6Anthropic2026-03-098017.59
Andon Labs Vending-Bench 22026-06-01
Source-checked
4GPT 5.5OpenAI2026-04-237523.84
Andon Labs Vending-Bench 22026-06-01
Source-checked
10Claude Fable 5 (high)Anthropic2026-06-095680.26
Andon Labs Vending-Bench 22026-06-01
Source-checked
Method

What it covers

Data-quality note pending. Every ledger row still carries source URL, source date, and ingest timestamp.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

Receipts

Sources and further reading

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.