Why this benchmark is useful
It stresses abstraction and multi-step inference. Use it to find models that can hold a problem together after the prompt stops looking like a memorized exam.
Benchmarks / Reasoning
HLE
Cross-domain expert-level questions — math, sciences, humanities, professional law and medicine — designed to be the hardest public exam a frontier model can take.
It stresses abstraction and multi-step inference. Use it to find models that can hold a problem together after the prompt stops looking like a memorized exam.
Read Humanity's Last Exam as a active signal with low contamination risk. Compare models only when the source uses the same harness, prompting setup, sampling policy, and score unit.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Claude Fable 5Anthropic | 2026-06-09 | 53.3% | 2026-07-21 | Source-checked |
| 2 | GPT-5.6 SolOpenAI | 2026-07-09 | 47.2% | 2026-07-21 | Source-checked |
| 3 | Claude Opus 4.8Anthropic | 2026-05-28 | 45.7% | 2026-07-21 | Source-checked |
| 4 | Muse Spark 1.1Meta | 2026-07-09 | 45.1% | 2026-07-21 | Source-checked |
| 5 | Gemini 3.1 Pro PreviewGoogle | 2026-02-19 | 44.7% | 2026-07-21 | Source-checked |
| 6 | GPT 5.5OpenAI | 2026-04-23 | 44.3% | 2026-07-21 | Source-checked |
| 7 | Kimi K3Moonshot | 2026-07-16 | 44.3% | 2026-07-21 | Source-checked |
| 8 | GPT-5.5 (high)OpenAI | 2026-04-23 | 43% | 2026-07-21 | Source-checked |
| 9 | GPT-5.6 TerraOpenAI | 2026-07-09 | 41.8% | 2026-07-21 | Source-checked |
| 10 | GPT 5.4OpenAI | 2026-03-05 | 41.6% | 2026-07-21 | Source-checked |
| 11 | Gemini 3.5 FlashGoogle | 2026-05-19 | 41% | 2026-07-21 | Source-checked |
| 12 | GPT-5.5 (medium)OpenAI | 2026-04-23 | 40.6% | 2026-07-21 | Source-checked |
| 13 | Grok 4.5xAI | 2026-07-08 | 40.3% | 2026-07-21 | Source-checked |
| 14 | GLM-5.2Zhipu | 2026-06-13 | 40.1% | 2026-07-21 | Source-checked |
| 15 | Gemini 3.5 Flash (medium)Google | 2026-05-19 | 39.9% | 2026-07-21 | Source-checked |
| 16 | GPT-5.3 Codex (xhigh)OpenAI | 2026-02-05 | 39.9% | 2026-07-21 | Source-checked |
| 17 | Muse SparkOther | 2026-01-01 | 39.9% | 2026-07-21 | Source-checked |
| 18 | Claude Opus 4.7Anthropic | 2026-04-16 | 39.6% | 2026-07-21 | Source-checked |
| 19 | Claude Sonnet 5Anthropic | 2026-06-30 | 39.6% | 2026-07-21 | Source-checked |
| 20 | Qwen3.7 MaxAlibaba | 2026-05-19 | 38.1% | 2026-07-21 | Source-checked |
| 21 | Gemini 3 ProGoogle | 2025-11-18 | 37.2% | 2026-07-21 | Source-checked |
| 22 | GPT-5.6 LunaOpenAI | 2026-07-09 | 37.2% | 2026-07-21 | Source-checked |
| 23 | MiniMax M3MiniMax | 2026-05-31 | 37.1% | 2026-07-21 | Source-checked |
| 24 | Claude Opus 4.6 (Adaptive Reasoning, Max Effort)Anthropic | 2026-02-05 | 36.7% | 2026-07-21 | Source-checked |
| 25 | DeepSeek V4 Pro (Reasoning, Max Effort)DeepSeek | 2026-04-24 | 35.9% | 2026-07-21 | Source-checked |
| 26 | Kimi K2.6Moonshot | 2026-04-20 | 35.9% | 2026-07-21 | Source-checked |
| 27 | GPT-5.2OpenAI | 2025-12-11 | 35.4% | 2026-07-21 | Source-checked |
| 28 | Grok 4.3xAI | 2026-04-30 | 35% | 2026-07-21 | Source-checked |
| 29 | MiMo-V2.5-ProXiaomi | 2026-04-22 | 33.8% | 2026-07-21 | Source-checked |
| 30 | Qwen 3.7 PlusAlibaba | 2026-06-02 | 33.4% | 2026-07-21 | Source-checked |
| 31 | Kimi K2.7 CodeMoonshot | 2026-06-12 | 32.8% | 2026-07-21 | Source-checked |
| 32 | Grok 4.20 ReasoningxAI | 2026-03-05 | 32.2% | 2026-07-21 | Source-checked |
| 33 | LongCat-2.0Meituan | 2026-06-29 | 32.1% | 2026-07-21 | Source-checked |
| 34 | Claude Opus 4.7 (Non-reasoning, High Effort)Anthropic | 2026-04-16 | 31.2% | 2026-07-21 | Source-checked |
| 35 | GPT-5.5 (low)OpenAI | 2026-04-23 | 31% | 2026-07-21 | Source-checked |
| 36 | Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)Anthropic | 2026-02-17 | 30% | 2026-07-21 | Source-checked |
| 37 | Kimi K2.5Moonshot | 2026-01-27 | 29.4% | 2026-07-21 | Source-checked |
| 38 | MiniMax M2.7MiniMax | 2026-04-15 | 28.1% | 2026-07-21 | Source-checked |
| 39 | GLM-5.1Zhipu | 2026-04-07 | 28% | 2026-07-21 | Source-checked |
| 40 | GPT-5.4 MiniOpenAI | 2026-03-17 | 26.6% | 2026-07-21 | Source-checked |
| 41 | Nemotron 3 Ultra 550B A55B (Reasoning)NVIDIA | 2026-06-04 | 26.6% | 2026-07-21 | Source-checked |
| 42 | GPT-5.1OpenAI | 2025-11-13 | 26.5% | 2026-07-21 | Source-checked |
| 43 | GPT-5.4 NanoOpenAI | 2026-03-17 | 26.5% | 2026-07-21 | Source-checked |
| 44 | Gemma 4 31B ITGoogle | 2026-03-01 | 22.7% | 2026-07-21 | Source-checked |
| 45 | OpenAI o3OpenAI | 2025-04-16 | 20% | 2026-07-21 | Source-checked |
| 46 | Step 3.7 FlashStepFun | 2026-06-01 | 19.9% | 2026-07-21 | Source-checked |
| 47 | MiniMax M2.5MiniMax | 2026-04-01 | 19.1% | 2026-07-21 | Source-checked |
| 48 | Claude Opus 4.6Anthropic | 2026-02-05 | 18.6% | 2026-07-21 | Source-checked |
| 49 | Gemini 3.1 Flash LiteGoogle | 2026-05-07 | 16.2% | 2026-07-21 | Source-checked |
| 50 | HyperNova 60B 2605Multiverse | 2026-05-06 | 15.1% | 2026-07-21 | Source-checked |
| 51 | DeepSeek R1DeepSeek | 2025-01-20 | 14.9% | 2026-07-21 | Source-checked |
| 52 | Gemma 4 12B (Reasoning)Google | 2026-06-07 | 14.8% | 2026-07-21 | Source-checked |
| 53 | Gemini 3 Flash PreviewGoogle | 2025-12-17 | 14.1% | 2026-07-21 | Source-checked |
| 54 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 13.2% | 2026-07-21 | Source-checked |
| 55 | Claude Opus 4.5Anthropic | 2025-11-24 | 12.9% | 2026-07-21 | Source-checked |
| 56 | North Mini CodeCohere | 2026-06-09 | 10% | 2026-07-21 | Source-checked |
| 57 | OpenAI o1OpenAI | 2024-09-12 | 7.7% | 2026-07-21 | Source-checked |
| 58 | LFM2.5-8B-A1BLiquid | 2026-06-08 | 6.9% | 2026-07-21 | Source-checked |
| 59 | MiniCPM5-1B (Reasoning)OpenBMB | 2026-06-04 | 6.5% | 2026-07-21 | Source-checked |
| 60 | Gemma 4 12B (Non-reasoning)Google | 2026-06-10 | 6.2% | 2026-07-21 | Source-checked |
| 61 | Llama 4 MaverickMeta | 2026-04-05 | 4.8% | 2026-07-21 | Source-checked |
| 62 | Command ACohere | 2026-03-15 | 4.6% | 2026-07-21 | Source-checked |
| 63 | Llama 4 ScoutMeta | 2026-04-05 | 4.3% | 2026-07-21 | Source-checked |
| 64 | Mistral Large 3Mistral | 2025-12-02 | 4.1% | 2026-07-21 | Source-checked |
| 65 | Claude 3 OpusAnthropic | 2024-03-04 | 3.1% | 2026-07-21 | Source-checked |
This benchmark sits in the reasoning family. It uses Multiple-choice and short-answer questions vetted by domain experts; submissions are reviewed for difficulty before scoring. The score should travel with its task format, scoring method, source date, and benchmark version.
Scores on VerdictPal come from Artificial Analysis' independent HLE run unless noted. The official project site hosts the dataset and paper — not a single maintained public leaderboard URL.
The Pack · Editorial newsletter
One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.