Benchmarks / Human preference

LMArena (Chatbot Arena)

Chatbot Arena

Human preference at scale: people chat with two anonymous models side-by-side and vote for the better reply. Votes are turned into an Elo-style leaderboard.

Human preferenceSolidActiveLow contamination riskSince 2023
What this does not measure
  • Correctness — voters reward answers that look good, so formatting, length and confidence can beat accuracy.
  • It is a popularity signal, not a capability test — style-controlled rankings can differ sharply from the raw board.
  • Elo gaps are noisy — overlapping confidence intervals mean small differences are not real.
Analysis

Why this benchmark is useful

Editorial brief pendingWe publish the methodology and ledger first; benchmark-specific analysis ships after desk review.

Scope

Coverage map

Task family
Human preference
Format
Open-ended blind pairwise chats; outcome is a Bradley-Terry / Elo rating with confidence intervals.
Scoring
Elo-style rating from millions of human votes.
Maintainer
LMArena (formerly LMSYS)
Reading guide

How to read the scores

Reading guide pending. Use task format, scoring method, and source dates in the ledger until the desk brief ships.

Blind spots

What it does not cover

  • Correctness — voters reward answers that look good, so formatting, length and confidence can beat accuracy.
  • It is a popularity signal, not a capability test — style-controlled rankings can differ sharply from the raw board.
  • Elo gaps are noisy — overlapping confidence intervals mean small differences are not real.
Scores

Evidence ledger

13 rows
1313 rows
1473best score
13source-checked
1sources
2026-05-25source date
LMArena13

official-page · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. GPT 5.51473 · LMArena leaderboard
  2. Claude Opus 4.81462 · LMArena leaderboard
  3. Gemini 3.5 Flash1455 · LMArena leaderboard
  4. Grok 4.31441 · LMArena leaderboard
  5. DeepSeek V41418 · LMArena leaderboard

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1GPT 5.5OpenAIStyle-controlled Elo. Headline ranks are volatile — read the 95% CI band, not the integer.2026-04-231473
LMArena leaderboard2026-05-25
Source-checked
2Claude Opus 4.8Anthropic2026-05-281462
LMArena leaderboard2026-05-25
Source-checked
3Gemini 3.5 FlashGoogle2026-05-191455
LMArena leaderboard2026-05-25
Source-checked
4Grok 4.3xAI2026-04-301441
LMArena leaderboard2026-05-25
Source-checked
5DeepSeek V4DeepSeek2026-04-241418
LMArena leaderboard2026-05-25
Source-checked
6Qwen 3.7 MaxAlibaba2026-05-201409
LMArena leaderboard2026-05-25
Source-checked
7Gemini 3.1 ProGoogle2026-02-191402
LMArena leaderboard2026-05-25
Source-checked
8Claude Sonnet 4.6Anthropic2026-02-171396
LMArena leaderboard2026-05-25
Source-checked
9Kimi K2.6Moonshot2026-04-201388
LMArena leaderboard2026-05-25
Source-checked
10Claude Opus 4.7Anthropic2026-04-161412
LMArena leaderboard2026-05-25
Source-checked
11GPT 5.4OpenAI2026-03-051405
LMArena leaderboard2026-05-25
Source-checked
12MiniMax M3MiniMax2026-05-311376
LMArena leaderboard2026-05-25
Source-checked
13Claude Haiku 4.5Anthropic2025-10-011368
LMArena leaderboard2026-05-25
Source-checked
Method

What it covers

Data-quality note pending. Every ledger row still carries source URL, source date, and ingest timestamp.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

Tools that report it

Receipts

Sources and further reading

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.