Benchmarks / Coding

HumanEval

Whether a model can write a short Python function from a docstring that passes a handful of unit tests — the original code-generation benchmark.

What this does not measure
  • Real software work — 164 self-contained toy functions are nothing like a production codebase.
  • It is heavily contaminated: the problems have circulated online for years and almost certainly sit in training data.
  • Debugging, multi-file changes, or reading existing code — none of it is tested here.
Analysis

Why this benchmark is useful

Historical code-generation baseline — short Python functions from docstrings. Solved and heavily contaminated; keep for lineage, not frontier coding rank.

Scope

Coverage map

Task family
Coding
Format
164 hand-written programming problems; metric is pass@1 (first attempt passes all tests).
Scoring
pass@k — fraction solved within k samples.
Maintainer
OpenAI (Chen et al.)
Reading guide

How to read the scores

pass@k on 164 toy functions. Scores above ~90–95% pass@1 no longer separate frontier models — prefer SWE-bench Verified, Terminal-Bench, or DeepSWE for discrimination.

Blind spots

What it does not cover

  • Real software work — 164 self-contained toy functions are nothing like a production codebase.
  • It is heavily contaminated: the problems have circulated online for years and almost certainly sit in training data.
  • Debugging, multi-file changes, or reading existing code — none of it is tested here.
Scores

Evidence ledger

15 rows
1515 rows
96.3% pass@1best score
4source-checked
3sources
2026-05-28to 2023-03-15
CodeSOTA6

manual-snapshot · manual

Papers / model cards3

manual-snapshot · manual

Vendor claims6

manual-snapshot · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Claude Opus 4.895.8% pass@1 · Anthropic Claude 4 family
  2. Qwen 3.7 Max93.1% pass@1 · CodeSOTA HumanEval leaderboard
  3. Gemini 3.5 Flash94.6% pass@1 · CodeSOTA HumanEval leaderboard
  4. DeepSeek V4 Pro Max93.4% pass@1 · DeepSeek V4 release
  5. GPT 5.596.2% pass@1 · CodeSOTA HumanEval leaderboard

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Claude Opus 4.6Anthropic2026-02-0596.3% pass@1
Anthropic Claude 3.7/4.6 model card2026-01-01
Source-checked
2GPT 5.5OpenAI2026-04-2396.2% pass@1
CodeSOTA HumanEval leaderboard2026-04-23
Needs audit
3Claude Opus 4.8Anthropic2026-05-2895.8% pass@1
Anthropic Claude 4 family2026-05-28
Needs audit
4GPT-5OpenAICross-reference with official releases before treating as ground truth.2025-12-1195.1% pass@1
CodeSOTA HumanEval leaderboard2025-12-01
Needs audit
5o3OpenAI2025-04-1694.8% pass@1
CodeSOTA HumanEval leaderboard2025-04-01
Needs audit
6Gemini 3.5 FlashGoogle2026-05-1994.6% pass@1
CodeSOTA HumanEval leaderboard2026-05-19
Needs audit
7Claude Sonnet 4.6Anthropic2026-02-1794.1% pass@1
Anthropic Claude 3.7/4.6 model card2026-01-01
Source-checked
8Kimi K2.6Moonshot2026-04-2093.8% pass@1
CodeSOTA HumanEval leaderboard2026-04-20
Needs audit
9DeepSeek V4 Pro MaxDeepSeek2026-04-2493.4% pass@1
DeepSeek V4 release2026-04-24
Needs audit
10Qwen 3.7 MaxAlibaba2026-05-2093.1% pass@1
CodeSOTA HumanEval leaderboard2026-05-20
Needs audit
11DeepSeek Coder V2DeepSeekOpen-weight model; treat as third-party result.2024-05-1592.7% pass@1
DeepSeek Coder V2 repository2024-05-15
Needs audit
12Mistral Large 3Mistral2025-12-0292.4% pass@1
Mistral Large 3 model card2025-12-02
Needs audit
13Claude 3.5 Sonnet (2024-06)Anthropic2024-06-2092% pass@1
Anthropic Claude 3.5 Sonnet2024-06-20
Source-checked
14Llama 4 MaverickMeta2026-04-0591.8% pass@1
Meta Llama 4 model card2026-04-05
Needs audit
15GPT-4 (2023)OpenAI2023-03-1467% pass@1
GPT-4 Technical Report2023-03-15
Source-checked
Method

What it covers

Archived — multiple frontier models exceed 95% pass@1. Use SWE-bench Verified or Terminal-Bench for coding discrimination.

Score ceiling

Where it breaks down

Frontier models exceed 90% pass@1, so the benchmark no longer separates them — treat high scores as table stakes.

Tools that report it

Receipts

Sources and further reading

Questions about this benchmark

Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.

What does HumanEval measure?

Whether a model can write a short Python function from a docstring that passes a handful of unit tests — the original code-generation benchmark.

What does a high HumanEval score not prove?

A strong HumanEval result says nothing about:

  • Real software work — 164 self-contained toy functions are nothing like a production codebase.
  • It is heavily contaminated: the problems have circulated online for years and almost certainly sit in training data.
  • Debugging, multi-file changes, or reading existing code — none of it is tested here.

How is HumanEval scored?

pass@k — fraction solved within k samples. Task format: 164 hand-written programming problems; metric is pass@1 (first attempt passes all tests).

Is HumanEval saturated?

HumanEval is currently marked Defunct in the atlas. Ceiling context: Frontier models exceed 90% pass@1, so the benchmark no longer separates them — treat high scores as table stakes.

Can HumanEval results be contaminated by training data?

Contamination risk for HumanEval is graded High contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.