Benchmarks / Coding

LiveCodeBench

Contamination-resistant coding problems drawn from recent contest releases — AA reruns models on the public harness.

CodingSolidActiveLow contamination riskSince 2024
What this does not measure
  • Multi-file refactors or IDE agent loops.
  • Private codebase context — problems are public contest slices only.
Analysis

Why this benchmark is useful

Editorial brief pendingWe publish the methodology and ledger first; benchmark-specific analysis ships after desk review.

Scope

Coverage map

Task family
Coding
Format
Timed coding tasks with public grader.
Scoring
pass@1 percentage from AA runs.
Maintainer
LiveCodeBench authors
Reading guide

How to read the scores

Reading guide pending. Use task format, scoring method, and source dates in the ledger until the desk brief ships.

Blind spots

What it does not cover

  • Multi-file refactors or IDE agent loops.
  • Private codebase context — problems are public contest slices only.
Scores

Evidence ledger

13 rows
1313 rows
91.7%best score
13source-checked
1sources
2026-07-21source date
13

api · api

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Gemini 3 Pro91.7% ·
  2. GPT-5.288.9% ·
  3. GPT-5.186.8% ·
  4. OpenAI o380.8% ·
  5. Gemini 3 Flash Preview79.7% ·

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Gemini 3 ProGoogle2025-11-1891.7%
2026-07-21
Source-checked
2GPT-5.2OpenAI2025-12-1188.9%
2026-07-21
Source-checked
3GPT-5.1OpenAI2025-11-1386.8%
2026-07-21
Source-checked
4OpenAI o3OpenAI2025-04-1680.8%
2026-07-21
Source-checked
5Gemini 3 Flash PreviewGoogle2025-12-1779.7%
2026-07-21
Source-checked
6DeepSeek R1DeepSeek2025-01-2077%
2026-07-21
Source-checked
7Claude Opus 4.5Anthropic2025-11-2473.8%
2026-07-21
Source-checked
8OpenAI o1OpenAI2024-09-1267.9%
2026-07-21
Source-checked
9Mistral Large 3Mistral2025-12-0246.5%
2026-07-21
Source-checked
10Llama 4 MaverickMeta2026-04-0539.7%
2026-07-21
Source-checked
11Llama 4 ScoutMeta2026-04-0529.9%
2026-07-21
Source-checked
12Command ACohere2026-03-1528.7%
2026-07-21
Source-checked
13Claude 3 OpusAnthropic2024-03-0427.9%
2026-07-21
Source-checked
Method

What it covers

Data-quality note pending. Every ledger row still carries source URL, source date, and ingest timestamp.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

Receipts

Sources and further reading

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.