Benchmarks / Coding

HumanEval

Whether a model can write a short Python function from a docstring that passes a handful of unit tests — the original code-generation benchmark.

CodingSolidDefunctHigh contamination riskSince 2021
What this does not measure
  • Real software work — 164 self-contained toy functions are nothing like a production codebase.
  • It is heavily contaminated: the problems have circulated online for years and almost certainly sit in training data.
  • Debugging, multi-file changes, or reading existing code — none of it is tested here.
Analysis

Why this benchmark is useful

It turns code ability into a patch or pass-rate signal you can inspect. That makes it useful for comparing agent harnesses, not just base models.

Scope

Coverage map

Task family
Coding
Format
164 hand-written programming problems; metric is pass@1 (first attempt passes all tests).
Scoring
pass@k — fraction solved within k samples.
Maintainer
OpenAI (Chen et al.)
Reading guide

How to read the scores

Read HumanEval as a defunct signal with high contamination risk. Compare models only when the source uses the same harness, prompting setup, sampling policy, and score unit.

Blind spots

What it does not cover

  • Real software work — 164 self-contained toy functions are nothing like a production codebase.
  • It is heavily contaminated: the problems have circulated online for years and almost certainly sit in training data.
  • Debugging, multi-file changes, or reading existing code — none of it is tested here.
Scores

Evidence ledger

15 rows
1515 rows
96.3% pass@1best score
4source-checked
3sources
2026-05-28to 2023-03-15
CodeSOTA6

manual-snapshot · manual

Papers / model cards3

manual-snapshot · manual

Vendor claims6

manual-snapshot · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Claude Opus 4.895.8% pass@1 · Anthropic Claude 4 family
  2. Qwen 3.7 Max93.1% pass@1 · CodeSOTA HumanEval leaderboard
  3. Gemini 3.5 Flash94.6% pass@1 · CodeSOTA HumanEval leaderboard
  4. DeepSeek V4 Pro Max93.4% pass@1 · DeepSeek V4 release
  5. GPT 5.596.2% pass@1 · CodeSOTA HumanEval leaderboard

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Claude Opus 4.6Anthropic2026-02-0596.3% pass@1
Anthropic Claude 3.7/4.6 model card2026-01-01
Source-checked
2GPT 5.5OpenAI2026-04-2396.2% pass@1
CodeSOTA HumanEval leaderboard2026-04-23
Needs audit
3Claude Opus 4.8Anthropic2026-05-2895.8% pass@1
Anthropic Claude 4 family2026-05-28
Needs audit
4GPT-5OpenAICross-reference with official releases before treating as ground truth.2025-12-1195.1% pass@1
CodeSOTA HumanEval leaderboard2025-12-01
Needs audit
5o3OpenAI2025-04-1694.8% pass@1
CodeSOTA HumanEval leaderboard2025-04-01
Needs audit
6Gemini 3.5 FlashGoogle2026-05-1994.6% pass@1
CodeSOTA HumanEval leaderboard2026-05-19
Needs audit
7Claude Sonnet 4.6Anthropic2026-02-1794.1% pass@1
Anthropic Claude 3.7/4.6 model card2026-01-01
Source-checked
8Kimi K2.6Moonshot2026-04-2093.8% pass@1
CodeSOTA HumanEval leaderboard2026-04-20
Needs audit
9DeepSeek V4 Pro MaxDeepSeek2026-04-2493.4% pass@1
DeepSeek V4 release2026-04-24
Needs audit
10Qwen 3.7 MaxAlibaba2026-05-2093.1% pass@1
CodeSOTA HumanEval leaderboard2026-05-20
Needs audit
11DeepSeek Coder V2DeepSeekOpen-weight model; treat as third-party result.2024-05-1592.7% pass@1
DeepSeek Coder V2 repository2024-05-15
Needs audit
12Mistral Large 3Mistral2025-12-0292.4% pass@1
Mistral Large 3 model card2025-12-02
Needs audit
13Claude 3.5 Sonnet (2024-06)Anthropic2024-06-2092% pass@1
Anthropic Claude 3.5 Sonnet2024-06-20
Source-checked
14Llama 4 MaverickMeta2026-04-0591.8% pass@1
Meta Llama 4 model card2026-04-05
Needs audit
15GPT-4 (2023)OpenAI2023-03-1467% pass@1
GPT-4 Technical Report2023-03-15
Source-checked
Method

What it covers

Archived — multiple frontier models exceed 95% pass@1. Use SWE-bench Verified or Terminal-Bench for coding discrimination.

Score ceiling

Where it breaks down

Frontier models exceed 90% pass@1, so the benchmark no longer separates them — treat high scores as table stakes.

Tools that report it

Receipts

Sources and further reading

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.