Benchmarks / Coding

LiveCodeBench

Contamination-resistant coding problems drawn from recent contest releases — AA reruns models on the public harness.

What this does not measure
  • Multi-file refactors or IDE agent loops.
  • Private codebase context — problems are public contest slices only.
Analysis

Why this benchmark is useful

When you want contest-style coding with fresher problems that resist training-data leakage better than classic static suites.

Scope

Coverage map

Task family
Coding
Format
Timed coding tasks with public grader.
Scoring
pass@1 percentage from AA runs.
Maintainer
LiveCodeBench authors
Reading guide

How to read the scores

pass@1 from AA public harness runs. Timed contest slices, not multi-file IDE agents — do not treat as production codebase skill.

Blind spots

What it does not cover

  • Multi-file refactors or IDE agent loops.
  • Private codebase context — problems are public contest slices only.
Scores

Evidence ledger

31 rows
3131 rows
91.7%best score
31source-checked
1sources
2026-09-01source date
31

api · api

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Gemini 3 Pro91.7% ·
  2. Gemini 3 Flash Preview (Reasoning)90.8% ·
  3. GLM-4.7 (Reasoning)89.4% ·
  4. GPT-5.2 (medium)89.4% ·
  5. GPT-5.288.9% ·

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Gemini 3 ProGoogle2025-11-1891.7%
2026-09-01
Source-checked
2Gemini 3 Flash Preview (Reasoning)Google2025-12-1790.8%
2026-09-01
Source-checked
3GLM-4.7 (Reasoning)Zhipu2025-12-2289.4%
2026-09-01
Source-checked
4GPT-5.2 (medium)OpenAI2025-12-1189.4%
2026-09-01
Source-checked
5GPT-5.2OpenAI2025-12-1188.9%
2026-09-01
Source-checked
6Claude Opus 4.5 (Reasoning)Anthropic2025-11-2487.1%
2026-09-01
Source-checked
7GPT-5.1OpenAI2025-11-1386.8%
2026-09-01
Source-checked
8MiMo-V2-Flash (Reasoning)Xiaomi2025-12-1686.8%
2026-09-01
Source-checked
9DeepSeek V3.2 (Reasoning)DeepSeek2025-12-0186.2%
2026-09-01
Source-checked
10Gemini 3 Pro Preview (low)Google2025-11-1885.7%
2026-09-01
Source-checked
11Kimi K2 ThinkingMoonshot2025-11-0685.3%
2026-09-01
Source-checked
12GPT-5.1 Codex (high)OpenAI2025-11-1384.9%
2026-09-01
Source-checked
13GPT-5 (high)OpenAI2025-08-0784.6%
2026-09-01
Source-checked
14GPT-5 Codex (high)OpenAI2025-09-2384%
2026-09-01
Source-checked
15Grok 4xAI2025-07-1081.9%
2026-09-01
Source-checked
16MiniMax-M2.1MiniMax2025-12-2381%
2026-09-01
Source-checked
17OpenAI o3OpenAI2025-04-1680.8%
2026-09-01
Source-checked
18Gemini 3 Flash PreviewGoogle2025-12-1779.7%
2026-09-01
Source-checked
19DeepSeek R1DeepSeek2025-01-2077%
2026-09-01
Source-checked
20GPT-5 (low)OpenAI2025-08-0776.3%
2026-09-01
Source-checked
21Claude Opus 4.5Anthropic2025-11-2473.8%
2026-09-01
Source-checked
22Claude 4.5 Sonnet (Reasoning)Anthropic2025-09-2971.4%
2026-09-01
Source-checked
23GPT-5 (medium)OpenAI2025-08-0770.3%
2026-09-01
Source-checked
24OpenAI o1OpenAI2024-09-1267.9%
2026-09-01
Source-checked
25Claude 4.1 Opus (Reasoning)Anthropic2025-08-0565.4%
2026-09-01
Source-checked
26Claude 4 Opus (Reasoning)Anthropic2025-05-2263.6%
2026-09-01
Source-checked
27Mistral Large 3Mistral2025-12-0246.5%
2026-09-01
Source-checked
28Llama 4 MaverickMeta2026-04-0539.7%
2026-09-01
Source-checked
29Llama 4 ScoutMeta2026-04-0529.9%
2026-09-01
Source-checked
30Command ACohere2026-03-1528.7%
2026-09-01
Source-checked
31Claude 3 OpusAnthropic2024-03-0427.9%
2026-09-01
Source-checked
Method

What it covers

This benchmark sits in the coding family. It uses Timed coding tasks with public grader. The score should travel with its task format, scoring method, source date, and benchmark version.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

Receipts

Sources and further reading

Questions about this benchmark

Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.

What does LiveCodeBench measure?

Contamination-resistant coding problems drawn from recent contest releases — AA reruns models on the public harness.

What does a high LiveCodeBench score not prove?

A strong LiveCodeBench result says nothing about:

  • Multi-file refactors or IDE agent loops.
  • Private codebase context — problems are public contest slices only.

How is LiveCodeBench scored?

pass@1 percentage from AA runs. Task format: Timed coding tasks with public grader.

Is LiveCodeBench saturated?

LiveCodeBench is currently marked Active in the atlas.

Can LiveCodeBench results be contaminated by training data?

Contamination risk for LiveCodeBench is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.