Why this benchmark is useful
It turns code ability into a patch or pass-rate signal you can inspect. That makes it useful for comparing agent harnesses, not just base models.
Benchmarks / Coding
Whether a model can write a short Python function from a docstring that passes a handful of unit tests — the original code-generation benchmark.
It turns code ability into a patch or pass-rate signal you can inspect. That makes it useful for comparing agent harnesses, not just base models.
Read HumanEval as a defunct signal with high contamination risk. Compare models only when the source uses the same harness, prompting setup, sampling policy, and score unit.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Claude Opus 4.6Anthropic | 2026-02-05 | 96.3% pass@1 | Anthropic Claude 3.7/4.6 model card2026-01-01 | Source-checked |
| 2 | GPT 5.5OpenAI | 2026-04-23 | 96.2% pass@1 | CodeSOTA HumanEval leaderboard2026-04-23 | Needs audit |
| 3 | Claude Opus 4.8Anthropic | 2026-05-28 | 95.8% pass@1 | Anthropic Claude 4 family2026-05-28 | Needs audit |
| 4 | GPT-5OpenAICross-reference with official releases before treating as ground truth. | 2025-12-11 | 95.1% pass@1 | CodeSOTA HumanEval leaderboard2025-12-01 | Needs audit |
| 5 | o3OpenAI | 2025-04-16 | 94.8% pass@1 | CodeSOTA HumanEval leaderboard2025-04-01 | Needs audit |
| 6 | Gemini 3.5 FlashGoogle | 2026-05-19 | 94.6% pass@1 | CodeSOTA HumanEval leaderboard2026-05-19 | Needs audit |
| 7 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 94.1% pass@1 | Anthropic Claude 3.7/4.6 model card2026-01-01 | Source-checked |
| 8 | Kimi K2.6Moonshot | 2026-04-20 | 93.8% pass@1 | CodeSOTA HumanEval leaderboard2026-04-20 | Needs audit |
| 9 | DeepSeek V4 Pro MaxDeepSeek | 2026-04-24 | 93.4% pass@1 | DeepSeek V4 release2026-04-24 | Needs audit |
| 10 | Qwen 3.7 MaxAlibaba | 2026-05-20 | 93.1% pass@1 | CodeSOTA HumanEval leaderboard2026-05-20 | Needs audit |
| 11 | DeepSeek Coder V2DeepSeekOpen-weight model; treat as third-party result. | 2024-05-15 | 92.7% pass@1 | DeepSeek Coder V2 repository2024-05-15 | Needs audit |
| 12 | Mistral Large 3Mistral | 2025-12-02 | 92.4% pass@1 | Mistral Large 3 model card2025-12-02 | Needs audit |
| 13 | Claude 3.5 Sonnet (2024-06)Anthropic | 2024-06-20 | 92% pass@1 | Anthropic Claude 3.5 Sonnet2024-06-20 | Source-checked |
| 14 | Llama 4 MaverickMeta | 2026-04-05 | 91.8% pass@1 | Meta Llama 4 model card2026-04-05 | Needs audit |
| 15 | GPT-4 (2023)OpenAI | 2023-03-14 | 67% pass@1 | GPT-4 Technical Report2023-03-15 | Source-checked |
Archived — multiple frontier models exceed 95% pass@1. Use SWE-bench Verified or Terminal-Bench for coding discrimination.
Frontier models exceed 90% pass@1, so the benchmark no longer separates them — treat high scores as table stakes.
The Pack · Editorial newsletter
One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.