Benchmarks / Reasoning

Humanity's Last Exam

HLE

Cross-domain expert-level questions — math, sciences, humanities, professional law and medicine — designed to be the hardest public exam a frontier model can take.

ReasoningFlagshipActiveLow contamination riskSince 2025
What this does not measure
  • Real research productivity — exam-style questions do not capture the messy search-and-synthesize work a researcher actually does.
  • Tool-augmented use — HLE is closed-book by default; models that leverage calculators, code, or web search get a separate tier.
  • Calibration — HLE reports accuracy, not how well the model knows what it doesn't know.
Analysis

Why this benchmark is useful

It stresses abstraction and multi-step inference. Use it to find models that can hold a problem together after the prompt stops looking like a memorized exam.

Scope

Coverage map

Task family
Reasoning
Format
Multiple-choice and short-answer questions vetted by domain experts; submissions are reviewed for difficulty before scoring.
Scoring
Percent of questions answered correctly on the public + private split; HLE reports a single accuracy figure per model.
Maintainer
Center for AI Safety / Scale AI
Reading guide

How to read the scores

Read Humanity's Last Exam as a active signal with low contamination risk. Compare models only when the source uses the same harness, prompting setup, sampling policy, and score unit.

Blind spots

What it does not cover

  • Real research productivity — exam-style questions do not capture the messy search-and-synthesize work a researcher actually does.
  • Tool-augmented use — HLE is closed-book by default; models that leverage calculators, code, or web search get a separate tier.
  • Calibration — HLE reports accuracy, not how well the model knows what it doesn't know.
Scores

Evidence ledger

65 rows
6565 rows
53.3%best score
65source-checked
1sources
2026-07-21source date
65

api · api

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Claude Fable 553.3% ·
  2. GPT-5.6 Sol47.2% ·
  3. Claude Opus 4.845.7% ·
  4. Muse Spark 1.145.1% ·
  5. Gemini 3.1 Pro Preview44.7% ·

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Claude Fable 5Anthropic2026-06-0953.3%
2026-07-21
Source-checked
2GPT-5.6 SolOpenAI2026-07-0947.2%
2026-07-21
Source-checked
3Claude Opus 4.8Anthropic2026-05-2845.7%
2026-07-21
Source-checked
4Muse Spark 1.1Meta2026-07-0945.1%
2026-07-21
Source-checked
5Gemini 3.1 Pro PreviewGoogle2026-02-1944.7%
2026-07-21
Source-checked
6GPT 5.5OpenAI2026-04-2344.3%
2026-07-21
Source-checked
7Kimi K3Moonshot2026-07-1644.3%
2026-07-21
Source-checked
8GPT-5.5 (high)OpenAI2026-04-2343%
2026-07-21
Source-checked
9GPT-5.6 TerraOpenAI2026-07-0941.8%
2026-07-21
Source-checked
10GPT 5.4OpenAI2026-03-0541.6%
2026-07-21
Source-checked
11Gemini 3.5 FlashGoogle2026-05-1941%
2026-07-21
Source-checked
12GPT-5.5 (medium)OpenAI2026-04-2340.6%
2026-07-21
Source-checked
13Grok 4.5xAI2026-07-0840.3%
2026-07-21
Source-checked
14GLM-5.2Zhipu2026-06-1340.1%
2026-07-21
Source-checked
15Gemini 3.5 Flash (medium)Google2026-05-1939.9%
2026-07-21
Source-checked
16GPT-5.3 Codex (xhigh)OpenAI2026-02-0539.9%
2026-07-21
Source-checked
17Muse SparkOther2026-01-0139.9%
2026-07-21
Source-checked
18Claude Opus 4.7Anthropic2026-04-1639.6%
2026-07-21
Source-checked
19Claude Sonnet 5Anthropic2026-06-3039.6%
2026-07-21
Source-checked
20Qwen3.7 MaxAlibaba2026-05-1938.1%
2026-07-21
Source-checked
21Gemini 3 ProGoogle2025-11-1837.2%
2026-07-21
Source-checked
22GPT-5.6 LunaOpenAI2026-07-0937.2%
2026-07-21
Source-checked
23MiniMax M3MiniMax2026-05-3137.1%
2026-07-21
Source-checked
24Claude Opus 4.6 (Adaptive Reasoning, Max Effort)Anthropic2026-02-0536.7%
2026-07-21
Source-checked
25DeepSeek V4 Pro (Reasoning, Max Effort)DeepSeek2026-04-2435.9%
2026-07-21
Source-checked
26Kimi K2.6Moonshot2026-04-2035.9%
2026-07-21
Source-checked
27GPT-5.2OpenAI2025-12-1135.4%
2026-07-21
Source-checked
28Grok 4.3xAI2026-04-3035%
2026-07-21
Source-checked
29MiMo-V2.5-ProXiaomi2026-04-2233.8%
2026-07-21
Source-checked
30Qwen 3.7 PlusAlibaba2026-06-0233.4%
2026-07-21
Source-checked
31Kimi K2.7 CodeMoonshot2026-06-1232.8%
2026-07-21
Source-checked
32Grok 4.20 ReasoningxAI2026-03-0532.2%
2026-07-21
Source-checked
33LongCat-2.0Meituan2026-06-2932.1%
2026-07-21
Source-checked
34Claude Opus 4.7 (Non-reasoning, High Effort)Anthropic2026-04-1631.2%
2026-07-21
Source-checked
35GPT-5.5 (low)OpenAI2026-04-2331%
2026-07-21
Source-checked
36Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)Anthropic2026-02-1730%
2026-07-21
Source-checked
37Kimi K2.5Moonshot2026-01-2729.4%
2026-07-21
Source-checked
38MiniMax M2.7MiniMax2026-04-1528.1%
2026-07-21
Source-checked
39GLM-5.1Zhipu2026-04-0728%
2026-07-21
Source-checked
40GPT-5.4 MiniOpenAI2026-03-1726.6%
2026-07-21
Source-checked
41Nemotron 3 Ultra 550B A55B (Reasoning)NVIDIA2026-06-0426.6%
2026-07-21
Source-checked
42GPT-5.1OpenAI2025-11-1326.5%
2026-07-21
Source-checked
43GPT-5.4 NanoOpenAI2026-03-1726.5%
2026-07-21
Source-checked
44Gemma 4 31B ITGoogle2026-03-0122.7%
2026-07-21
Source-checked
45OpenAI o3OpenAI2025-04-1620%
2026-07-21
Source-checked
46Step 3.7 FlashStepFun2026-06-0119.9%
2026-07-21
Source-checked
47MiniMax M2.5MiniMax2026-04-0119.1%
2026-07-21
Source-checked
48Claude Opus 4.6Anthropic2026-02-0518.6%
2026-07-21
Source-checked
49Gemini 3.1 Flash LiteGoogle2026-05-0716.2%
2026-07-21
Source-checked
50HyperNova 60B 2605Multiverse2026-05-0615.1%
2026-07-21
Source-checked
51DeepSeek R1DeepSeek2025-01-2014.9%
2026-07-21
Source-checked
52Gemma 4 12B (Reasoning)Google2026-06-0714.8%
2026-07-21
Source-checked
53Gemini 3 Flash PreviewGoogle2025-12-1714.1%
2026-07-21
Source-checked
54Claude Sonnet 4.6Anthropic2026-02-1713.2%
2026-07-21
Source-checked
55Claude Opus 4.5Anthropic2025-11-2412.9%
2026-07-21
Source-checked
56North Mini CodeCohere2026-06-0910%
2026-07-21
Source-checked
57OpenAI o1OpenAI2024-09-127.7%
2026-07-21
Source-checked
58LFM2.5-8B-A1BLiquid2026-06-086.9%
2026-07-21
Source-checked
59MiniCPM5-1B (Reasoning)OpenBMB2026-06-046.5%
2026-07-21
Source-checked
60Gemma 4 12B (Non-reasoning)Google2026-06-106.2%
2026-07-21
Source-checked
61Llama 4 MaverickMeta2026-04-054.8%
2026-07-21
Source-checked
62Command ACohere2026-03-154.6%
2026-07-21
Source-checked
63Llama 4 ScoutMeta2026-04-054.3%
2026-07-21
Source-checked
64Mistral Large 3Mistral2025-12-024.1%
2026-07-21
Source-checked
65Claude 3 OpusAnthropic2024-03-043.1%
2026-07-21
Source-checked
Method

What it covers

This benchmark sits in the reasoning family. It uses Multiple-choice and short-answer questions vetted by domain experts; submissions are reviewed for difficulty before scoring. The score should travel with its task format, scoring method, source date, and benchmark version.

Score ceiling

Where it breaks down

Scores on VerdictPal come from Artificial Analysis' independent HLE run unless noted. The official project site hosts the dataset and paper — not a single maintained public leaderboard URL.

Receipts

Sources and further reading

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.