Why this benchmark is useful
Many classic tests are memorized, saturated, or stale. LiveBench earns a slot because it explicitly refreshes the question pool on a schedule.
Benchmarks / Reasoning
A contamination-limited benchmark that refreshes questions over time across reasoning, coding, math, data analysis, language, and instruction-following tasks.
Many classic tests are memorized, saturated, or stale. LiveBench earns a slot because it explicitly refreshes the question pool on a schedule.
Always read the LiveBench release date with the score. A model's position on an older snapshot may not survive the next refresh.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Claude Opus 4.6AnthropicLatest LiveBench release. Compare within same release only. | 2026-02-05 | 71.4% | LiveBench leaderboard2026-04-15 | Needs audit |
| 2 | Claude Opus 4.8Anthropic | 2026-05-28 | 70.8% | LiveBench leaderboard2026-04-15 | Needs audit |
| 3 | GPT 5.5OpenAI | 2026-04-23 | 69.8% | LiveBench leaderboard2026-04-15 | Needs audit |
| 4 | o3OpenAIStrong in math and reasoning categories; softer in instruction-following. | 2025-04-16 | 68.5% | LiveBench leaderboard2026-04-15 | Needs audit |
| 5 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 67.2% | LiveBench leaderboard2026-04-15 | Needs audit |
| 6 | GPT 5.4OpenAI | 2026-03-05 | 66.4% | LiveBench leaderboard2026-04-15 | Needs audit |
| 7 | Gemini 3.1 ProGoogle | 2026-02-19 | 65.8% | LiveBench leaderboard2026-04-15 | Needs audit |
| 8 | Gemini 3.5 FlashGoogle | 2026-05-19 | 64.2% | LiveBench leaderboard2026-04-15 | Needs audit |
| 9 | DeepSeek R1DeepSeek | 2025-01-20 | 64.1% | LiveBench leaderboard2026-04-15 | Needs audit |
| 10 | Gemini 3 Flash PreviewGoogle | 2025-12-17 | 63.5% | LiveBench leaderboard2026-04-15 | Needs audit |
| 11 | Mistral Large 3Mistral | 2025-12-02 | 62.8% | LiveBench leaderboard2026-04-15 | Needs audit |
| 12 | Kimi K2.6Moonshot | 2026-04-20 | 62.4% | LiveBench leaderboard2026-04-15 | Needs audit |
| 13 | Llama 4 MaverickMeta | 2026-04-05 | 62.1% | LiveBench leaderboard2026-04-15 | Needs audit |
| 14 | DeepSeek V4DeepSeek | 2026-04-24 | 61.9% | LiveBench leaderboard2026-04-15 | Needs audit |
| 15 | Qwen 3.7 MaxAlibaba | 2026-05-20 | 61.6% | LiveBench leaderboard2026-04-15 | Needs audit |
| 16 | Kimi K2.5Moonshot | 2026-01-27 | 61.2% | LiveBench leaderboard2026-04-15 | Needs audit |
| 17 | Grok 4.3xAI | 2026-04-30 | 60.8% | LiveBench leaderboard2026-04-15 | Needs audit |
| 19 | MiniMax M3MiniMax | 2026-05-31 | 59.6% | LiveBench leaderboard2026-04-15 | Needs audit |
| 20 | Llama 4 ScoutMeta | 2026-04-05 | 58.9% | LiveBench leaderboard2026-04-15 | Needs audit |
| 21 | Claude Haiku 4.5Anthropic | 2025-10-01 | 57.8% | LiveBench leaderboard2026-04-15 | Needs audit |
| 22 | Grok 4.20 ReasoningxAI | 2026-03-05 | 56.4% | LiveBench leaderboard2026-04-15 | Needs audit |
Primary source is livebench.ai. VerdictPal should record the LiveBench release label with every score row.
The broad average is a map, not a verdict. Drill into task families before using it to choose a tool.
The Pack · Editorial newsletter
One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.