Why this benchmark is useful
When you want contest-style coding with fresher problems that resist training-data leakage better than classic static suites.
Contamination-resistant coding problems drawn from recent contest releases — AA reruns models on the public harness.
When you want contest-style coding with fresher problems that resist training-data leakage better than classic static suites.
pass@1 from AA public harness runs. Timed contest slices, not multi-file IDE agents — do not treat as production codebase skill.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Gemini 3 ProGoogle | 2025-11-18 | 91.7% | 2026-09-01 | Source-checked |
| 2 | Gemini 3 Flash Preview (Reasoning)Google | 2025-12-17 | 90.8% | 2026-09-01 | Source-checked |
| 3 | GLM-4.7 (Reasoning)Zhipu | 2025-12-22 | 89.4% | 2026-09-01 | Source-checked |
| 4 | GPT-5.2 (medium)OpenAI | 2025-12-11 | 89.4% | 2026-09-01 | Source-checked |
| 5 | GPT-5.2OpenAI | 2025-12-11 | 88.9% | 2026-09-01 | Source-checked |
| 6 | Claude Opus 4.5 (Reasoning)Anthropic | 2025-11-24 | 87.1% | 2026-09-01 | Source-checked |
| 7 | GPT-5.1OpenAI | 2025-11-13 | 86.8% | 2026-09-01 | Source-checked |
| 8 | MiMo-V2-Flash (Reasoning)Xiaomi | 2025-12-16 | 86.8% | 2026-09-01 | Source-checked |
| 9 | DeepSeek V3.2 (Reasoning)DeepSeek | 2025-12-01 | 86.2% | 2026-09-01 | Source-checked |
| 10 | Gemini 3 Pro Preview (low)Google | 2025-11-18 | 85.7% | 2026-09-01 | Source-checked |
| 11 | Kimi K2 ThinkingMoonshot | 2025-11-06 | 85.3% | 2026-09-01 | Source-checked |
| 12 | GPT-5.1 Codex (high)OpenAI | 2025-11-13 | 84.9% | 2026-09-01 | Source-checked |
| 13 | GPT-5 (high)OpenAI | 2025-08-07 | 84.6% | 2026-09-01 | Source-checked |
| 14 | GPT-5 Codex (high)OpenAI | 2025-09-23 | 84% | 2026-09-01 | Source-checked |
| 15 | Grok 4xAI | 2025-07-10 | 81.9% | 2026-09-01 | Source-checked |
| 16 | MiniMax-M2.1MiniMax | 2025-12-23 | 81% | 2026-09-01 | Source-checked |
| 17 | OpenAI o3OpenAI | 2025-04-16 | 80.8% | 2026-09-01 | Source-checked |
| 18 | Gemini 3 Flash PreviewGoogle | 2025-12-17 | 79.7% | 2026-09-01 | Source-checked |
| 19 | DeepSeek R1DeepSeek | 2025-01-20 | 77% | 2026-09-01 | Source-checked |
| 20 | GPT-5 (low)OpenAI | 2025-08-07 | 76.3% | 2026-09-01 | Source-checked |
| 21 | Claude Opus 4.5Anthropic | 2025-11-24 | 73.8% | 2026-09-01 | Source-checked |
| 22 | Claude 4.5 Sonnet (Reasoning)Anthropic | 2025-09-29 | 71.4% | 2026-09-01 | Source-checked |
| 23 | GPT-5 (medium)OpenAI | 2025-08-07 | 70.3% | 2026-09-01 | Source-checked |
| 24 | OpenAI o1OpenAI | 2024-09-12 | 67.9% | 2026-09-01 | Source-checked |
| 25 | Claude 4.1 Opus (Reasoning)Anthropic | 2025-08-05 | 65.4% | 2026-09-01 | Source-checked |
| 26 | Claude 4 Opus (Reasoning)Anthropic | 2025-05-22 | 63.6% | 2026-09-01 | Source-checked |
| 27 | Mistral Large 3Mistral | 2025-12-02 | 46.5% | 2026-09-01 | Source-checked |
| 28 | Llama 4 MaverickMeta | 2026-04-05 | 39.7% | 2026-09-01 | Source-checked |
| 29 | Llama 4 ScoutMeta | 2026-04-05 | 29.9% | 2026-09-01 | Source-checked |
| 30 | Command ACohere | 2026-03-15 | 28.7% | 2026-09-01 | Source-checked |
| 31 | Claude 3 OpusAnthropic | 2024-03-04 | 27.9% | 2026-09-01 | Source-checked |
This benchmark sits in the coding family. It uses Timed coding tasks with public grader. The score should travel with its task format, scoring method, source date, and benchmark version.
No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
Contamination-resistant coding problems drawn from recent contest releases — AA reruns models on the public harness.
A strong LiveCodeBench result says nothing about:
pass@1 percentage from AA runs. Task format: Timed coding tasks with public grader.
LiveCodeBench is currently marked Active in the atlas.
Contamination risk for LiveCodeBench is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.