Benchmarks / Human preference

LMArena (Chatbot Arena)

Chatbot Arena

Human preference at scale: people chat with two anonymous models side-by-side and vote for the better reply. Votes are turned into an Elo-style leaderboard.

What this does not measure
  • Correctness — voters reward answers that look good, so formatting, length and confidence can beat accuracy.
  • It is a popularity signal, not a capability test — style-controlled rankings can differ sharply from the raw board.
  • Elo gaps are noisy — overlapping confidence intervals mean small differences are not real.
Analysis

Why this benchmark is useful

When you need a human preference signal at scale — how replies feel in blind side-by-side chat — not a correctness exam.

Scope

Coverage map

Task family
Human preference
Format
Open-ended blind pairwise chats; outcome is a Bradley-Terry / Elo rating with confidence intervals.
Scoring
Elo-style rating from millions of human votes.
Maintainer
LMArena (formerly LMSYS)
Reading guide

How to read the scores

Read Elo-style ratings with confidence intervals. Overlapping intervals mean the gap is noise; style and length can beat accuracy, so treat it as preference, not capability proof.

Blind spots

What it does not cover

  • Correctness — voters reward answers that look good, so formatting, length and confidence can beat accuracy.
  • It is a popularity signal, not a capability test — style-controlled rankings can differ sharply from the raw board.
  • Elo gaps are noisy — overlapping confidence intervals mean small differences are not real.
Scores

Evidence ledger

13 rows
1313 rows
1473best score
13source-checked
1sources
2026-05-25source date
LMArena13

official-page · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. GPT 5.51473 · LMArena leaderboard
  2. Claude Opus 4.81462 · LMArena leaderboard
  3. Gemini 3.5 Flash1455 · LMArena leaderboard
  4. Grok 4.31441 · LMArena leaderboard
  5. DeepSeek V41418 · LMArena leaderboard

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1GPT 5.5OpenAIStyle-controlled Elo. Headline ranks are volatile — read the 95% CI band, not the integer.2026-04-231473
LMArena leaderboard2026-05-25
Source-checked
2Claude Opus 4.8Anthropic2026-05-281462
LMArena leaderboard2026-05-25
Source-checked
3Gemini 3.5 FlashGoogle2026-05-191455
LMArena leaderboard2026-05-25
Source-checked
4Grok 4.3xAI2026-04-301441
LMArena leaderboard2026-05-25
Source-checked
5DeepSeek V4DeepSeek2026-04-241418
LMArena leaderboard2026-05-25
Source-checked
6Qwen 3.7 MaxAlibaba2026-05-201409
LMArena leaderboard2026-05-25
Source-checked
7Gemini 3.1 ProGoogle2026-02-191402
LMArena leaderboard2026-05-25
Source-checked
8Claude Sonnet 4.6Anthropic2026-02-171396
LMArena leaderboard2026-05-25
Source-checked
9Kimi K2.6Moonshot2026-04-201388
LMArena leaderboard2026-05-25
Source-checked
10Claude Opus 4.7Anthropic2026-04-161412
LMArena leaderboard2026-05-25
Source-checked
11GPT 5.4OpenAI2026-03-051405
LMArena leaderboard2026-05-25
Source-checked
12MiniMax M3MiniMax2026-05-311376
LMArena leaderboard2026-05-25
Source-checked
13Claude Haiku 4.5Anthropic2025-10-011368
LMArena leaderboard2026-05-25
Source-checked
Method

What it covers

This benchmark sits in the human preference family. It uses Open-ended blind pairwise chats; outcome is a Bradley-Terry / Elo rating with confidence intervals. The score should travel with its task format, scoring method, source date, and benchmark version.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

Tools that report it

Receipts

Sources and further reading

Questions about this benchmark

Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.

What does LMArena (Chatbot Arena) measure?

Human preference at scale: people chat with two anonymous models side-by-side and vote for the better reply. Votes are turned into an Elo-style leaderboard.

What does a high LMArena (Chatbot Arena) score not prove?

A strong LMArena (Chatbot Arena) result says nothing about:

  • Correctness — voters reward answers that look good, so formatting, length and confidence can beat accuracy.
  • It is a popularity signal, not a capability test — style-controlled rankings can differ sharply from the raw board.
  • Elo gaps are noisy — overlapping confidence intervals mean small differences are not real.

How is LMArena (Chatbot Arena) scored?

Elo-style rating from millions of human votes. Task format: Open-ended blind pairwise chats; outcome is a Bradley-Terry / Elo rating with confidence intervals.

Is LMArena (Chatbot Arena) saturated?

LMArena (Chatbot Arena) is currently marked Active in the atlas.

Can LMArena (Chatbot Arena) results be contaminated by training data?

Contamination risk for LMArena (Chatbot Arena) is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.