Why this benchmark is useful
Historical code-generation baseline — short Python functions from docstrings. Solved and heavily contaminated; keep for lineage, not frontier coding rank.
Whether a model can write a short Python function from a docstring that passes a handful of unit tests — the original code-generation benchmark.
Historical code-generation baseline — short Python functions from docstrings. Solved and heavily contaminated; keep for lineage, not frontier coding rank.
pass@k on 164 toy functions. Scores above ~90–95% pass@1 no longer separate frontier models — prefer SWE-bench Verified, Terminal-Bench, or DeepSWE for discrimination.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Claude Opus 4.6Anthropic | 2026-02-05 | 96.3% pass@1 | Anthropic Claude 3.7/4.6 model card2026-01-01 | Source-checked |
| 2 | GPT 5.5OpenAI | 2026-04-23 | 96.2% pass@1 | CodeSOTA HumanEval leaderboard2026-04-23 | Needs audit |
| 3 | Claude Opus 4.8Anthropic | 2026-05-28 | 95.8% pass@1 | Anthropic Claude 4 family2026-05-28 | Needs audit |
| 4 | GPT-5OpenAICross-reference with official releases before treating as ground truth. | 2025-12-11 | 95.1% pass@1 | CodeSOTA HumanEval leaderboard2025-12-01 | Needs audit |
| 5 | o3OpenAI | 2025-04-16 | 94.8% pass@1 | CodeSOTA HumanEval leaderboard2025-04-01 | Needs audit |
| 6 | Gemini 3.5 FlashGoogle | 2026-05-19 | 94.6% pass@1 | CodeSOTA HumanEval leaderboard2026-05-19 | Needs audit |
| 7 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 94.1% pass@1 | Anthropic Claude 3.7/4.6 model card2026-01-01 | Source-checked |
| 8 | Kimi K2.6Moonshot | 2026-04-20 | 93.8% pass@1 | CodeSOTA HumanEval leaderboard2026-04-20 | Needs audit |
| 9 | DeepSeek V4 Pro MaxDeepSeek | 2026-04-24 | 93.4% pass@1 | DeepSeek V4 release2026-04-24 | Needs audit |
| 10 | Qwen 3.7 MaxAlibaba | 2026-05-20 | 93.1% pass@1 | CodeSOTA HumanEval leaderboard2026-05-20 | Needs audit |
| 11 | DeepSeek Coder V2DeepSeekOpen-weight model; treat as third-party result. | 2024-05-15 | 92.7% pass@1 | DeepSeek Coder V2 repository2024-05-15 | Needs audit |
| 12 | Mistral Large 3Mistral | 2025-12-02 | 92.4% pass@1 | Mistral Large 3 model card2025-12-02 | Needs audit |
| 13 | Claude 3.5 Sonnet (2024-06)Anthropic | 2024-06-20 | 92% pass@1 | Anthropic Claude 3.5 Sonnet2024-06-20 | Source-checked |
| 14 | Llama 4 MaverickMeta | 2026-04-05 | 91.8% pass@1 | Meta Llama 4 model card2026-04-05 | Needs audit |
| 15 | GPT-4 (2023)OpenAI | 2023-03-14 | 67% pass@1 | GPT-4 Technical Report2023-03-15 | Source-checked |
Archived — multiple frontier models exceed 95% pass@1. Use SWE-bench Verified or Terminal-Bench for coding discrimination.
Frontier models exceed 90% pass@1, so the benchmark no longer separates them — treat high scores as table stakes.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
Whether a model can write a short Python function from a docstring that passes a handful of unit tests — the original code-generation benchmark.
A strong HumanEval result says nothing about:
pass@k — fraction solved within k samples. Task format: 164 hand-written programming problems; metric is pass@1 (first attempt passes all tests).
HumanEval is currently marked Defunct in the atlas. Ceiling context: Frontier models exceed 90% pass@1, so the benchmark no longer separates them — treat high scores as table stakes.
Contamination risk for HumanEval is graded High contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.