Why this benchmark is useful
When you need a human preference signal at scale — how replies feel in blind side-by-side chat — not a correctness exam.
Chatbot Arena
Human preference at scale: people chat with two anonymous models side-by-side and vote for the better reply. Votes are turned into an Elo-style leaderboard.
When you need a human preference signal at scale — how replies feel in blind side-by-side chat — not a correctness exam.
Read Elo-style ratings with confidence intervals. Overlapping intervals mean the gap is noise; style and length can beat accuracy, so treat it as preference, not capability proof.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | GPT 5.5OpenAIStyle-controlled Elo. Headline ranks are volatile — read the 95% CI band, not the integer. | 2026-04-23 | 1473 | LMArena leaderboard2026-05-25 | Source-checked |
| 2 | Claude Opus 4.8Anthropic | 2026-05-28 | 1462 | LMArena leaderboard2026-05-25 | Source-checked |
| 3 | Gemini 3.5 FlashGoogle | 2026-05-19 | 1455 | LMArena leaderboard2026-05-25 | Source-checked |
| 4 | Grok 4.3xAI | 2026-04-30 | 1441 | LMArena leaderboard2026-05-25 | Source-checked |
| 5 | DeepSeek V4DeepSeek | 2026-04-24 | 1418 | LMArena leaderboard2026-05-25 | Source-checked |
| 6 | Qwen 3.7 MaxAlibaba | 2026-05-20 | 1409 | LMArena leaderboard2026-05-25 | Source-checked |
| 7 | Gemini 3.1 ProGoogle | 2026-02-19 | 1402 | LMArena leaderboard2026-05-25 | Source-checked |
| 8 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 1396 | LMArena leaderboard2026-05-25 | Source-checked |
| 9 | Kimi K2.6Moonshot | 2026-04-20 | 1388 | LMArena leaderboard2026-05-25 | Source-checked |
| 10 | Claude Opus 4.7Anthropic | 2026-04-16 | 1412 | LMArena leaderboard2026-05-25 | Source-checked |
| 11 | GPT 5.4OpenAI | 2026-03-05 | 1405 | LMArena leaderboard2026-05-25 | Source-checked |
| 12 | MiniMax M3MiniMax | 2026-05-31 | 1376 | LMArena leaderboard2026-05-25 | Source-checked |
| 13 | Claude Haiku 4.5Anthropic | 2025-10-01 | 1368 | LMArena leaderboard2026-05-25 | Source-checked |
This benchmark sits in the human preference family. It uses Open-ended blind pairwise chats; outcome is a Bradley-Terry / Elo rating with confidence intervals. The score should travel with its task format, scoring method, source date, and benchmark version.
No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
Human preference at scale: people chat with two anonymous models side-by-side and vote for the better reply. Votes are turned into an Elo-style leaderboard.
A strong LMArena (Chatbot Arena) result says nothing about:
Elo-style rating from millions of human votes. Task format: Open-ended blind pairwise chats; outcome is a Bradley-Terry / Elo rating with confidence intervals.
LMArena (Chatbot Arena) is currently marked Active in the atlas.
Contamination risk for LMArena (Chatbot Arena) is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.