Benchmarks / Reasoning

LiveBench

A contamination-limited benchmark that refreshes questions over time across reasoning, coding, math, data analysis, language, and instruction-following tasks.

ReasoningSolidActiveLow contamination riskSince 2024
What this does not measure
  • Static benchmark mastery — the point is freshness, so old snapshots should not be treated like permanent model truth.
  • Domain-specific reliability — a broad score can hide weak behavior in the domain your reader actually cares about.
  • Cost, latency, or privacy — LiveBench is capability evidence, not product-readiness evidence.
Analysis

Why this benchmark is useful

Many classic tests are memorized, saturated, or stale. LiveBench earns a slot because it explicitly refreshes the question pool on a schedule.

Scope

Coverage map

Task family
Reasoning
Format
Freshly refreshed tasks across multiple categories; recent releases are published on the LiveBench site.
Scoring
Category and global averages by LiveBench release. Compare only inside the same release.
Maintainer
LiveBench team
Reading guide

How to read the scores

Always read the LiveBench release date with the score. A model's position on an older snapshot may not survive the next refresh.

Blind spots

What it does not cover

  • Static benchmark mastery — the point is freshness, so old snapshots should not be treated like permanent model truth.
  • Domain-specific reliability — a broad score can hide weak behavior in the domain your reader actually cares about.
  • Cost, latency, or privacy — LiveBench is capability evidence, not product-readiness evidence.
Scores

Evidence ledger

21 rows
2121 rows
71.4%best score
0source-checked
1sources
2026-04-15source date
LiveBench21

official-page · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Claude Opus 4.671.4% · LiveBench leaderboard
  2. Claude Opus 4.870.8% · LiveBench leaderboard
  3. GPT 5.569.8% · LiveBench leaderboard
  4. o368.5% · LiveBench leaderboard
  5. Claude Sonnet 4.667.2% · LiveBench leaderboard

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Claude Opus 4.6AnthropicLatest LiveBench release. Compare within same release only.2026-02-0571.4%
LiveBench leaderboard2026-04-15
Needs audit
2Claude Opus 4.8Anthropic2026-05-2870.8%
LiveBench leaderboard2026-04-15
Needs audit
3GPT 5.5OpenAI2026-04-2369.8%
LiveBench leaderboard2026-04-15
Needs audit
4o3OpenAIStrong in math and reasoning categories; softer in instruction-following.2025-04-1668.5%
LiveBench leaderboard2026-04-15
Needs audit
5Claude Sonnet 4.6Anthropic2026-02-1767.2%
LiveBench leaderboard2026-04-15
Needs audit
6GPT 5.4OpenAI2026-03-0566.4%
LiveBench leaderboard2026-04-15
Needs audit
7Gemini 3.1 ProGoogle2026-02-1965.8%
LiveBench leaderboard2026-04-15
Needs audit
8Gemini 3.5 FlashGoogle2026-05-1964.2%
LiveBench leaderboard2026-04-15
Needs audit
9DeepSeek R1DeepSeek2025-01-2064.1%
LiveBench leaderboard2026-04-15
Needs audit
10Gemini 3 Flash PreviewGoogle2025-12-1763.5%
LiveBench leaderboard2026-04-15
Needs audit
11Mistral Large 3Mistral2025-12-0262.8%
LiveBench leaderboard2026-04-15
Needs audit
12Kimi K2.6Moonshot2026-04-2062.4%
LiveBench leaderboard2026-04-15
Needs audit
13Llama 4 MaverickMeta2026-04-0562.1%
LiveBench leaderboard2026-04-15
Needs audit
14DeepSeek V4DeepSeek2026-04-2461.9%
LiveBench leaderboard2026-04-15
Needs audit
15Qwen 3.7 MaxAlibaba2026-05-2061.6%
LiveBench leaderboard2026-04-15
Needs audit
16Kimi K2.5Moonshot2026-01-2761.2%
LiveBench leaderboard2026-04-15
Needs audit
17Grok 4.3xAI2026-04-3060.8%
LiveBench leaderboard2026-04-15
Needs audit
19MiniMax M3MiniMax2026-05-3159.6%
LiveBench leaderboard2026-04-15
Needs audit
20Llama 4 ScoutMeta2026-04-0558.9%
LiveBench leaderboard2026-04-15
Needs audit
21Claude Haiku 4.5Anthropic2025-10-0157.8%
LiveBench leaderboard2026-04-15
Needs audit
22Grok 4.20 ReasoningxAI2026-03-0556.4%
LiveBench leaderboard2026-04-15
Needs audit
Method

What it covers

Primary source is livebench.ai. VerdictPal should record the LiveBench release label with every score row.

Score ceiling

Where it breaks down

The broad average is a map, not a verdict. Drill into task families before using it to choose a tool.

Tools that report it

Receipts

Sources and further reading

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.