Benchmarks / Reasoning

GPQA Diamond

Graduate-Level Google-Proof Q&A

Hard graduate-level science questions (biology, physics, chemistry) written so that non-experts cannot solve them even with web access — the 'Diamond' subset is the highest-quality, expert-validated slice.

ReasoningSolidActiveLow contamination riskSince 2023
What this does not measure
  • General usefulness — it is a narrow slice of frontier science, not everyday tasks.
  • Whether a model knows when it is out of its depth; it still answers every question.
  • Small size (198 Diamond questions) means a few lucky guesses move the score noticeably.
Analysis

Why this benchmark is useful

Editorial brief pendingWe publish the methodology and ledger first; benchmark-specific analysis ships after desk review.

Scope

Coverage map

Task family
Reasoning
Format
198 expert-validated multiple-choice questions in the Diamond set; designed to be 'Google-proof'.
Scoring
Accuracy (% correct).
Maintainer
Rein et al.
Reading guide

How to read the scores

Reading guide pending. Use task format, scoring method, and source dates in the ledger until the desk brief ships.

Blind spots

What it does not cover

  • General usefulness — it is a narrow slice of frontier science, not everyday tasks.
  • Whether a model knows when it is out of its depth; it still answers every question.
  • Small size (198 Diamond questions) means a few lucky guesses move the score noticeably.
Scores

Evidence ledger

69 rows
6969 rows
94.6%best score
62source-checked
2sources
2026-07-21to 2026-05-15
57

api · api

Papers / model cards12

manual-snapshot · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Gemini 3.1 Pro Preview94.1% ·
  2. GPT-5.6 Sol94.1% ·
  3. Kimi K393.5% ·
  4. GPT-5.5 (high)93.2% ·
  5. Grok 4.593.1% ·

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Claude Mythos PreviewAnthropic2026-06-0194.6%
GPQA Diamond paper leaderboard2026-06-02
Source-checked
2Gemini 3.1 Pro PreviewGoogle2026-02-1994.1%
2026-07-21
Source-checked
3GPT-5.6 SolOpenAI2026-07-0994.1%
2026-07-21
Source-checked
4Kimi K3Moonshot2026-07-1693.5%
2026-07-21
Source-checked
5GPT-5.5 (high)OpenAI2026-04-2393.2%
2026-07-21
Source-checked
6Grok 4.5xAI2026-07-0893.1%
2026-07-21
Source-checked
7MiniMax M3MiniMax2026-05-3192.9%
2026-07-21
Source-checked
8Claude Fable 5Anthropic2026-06-0992.6%
2026-07-21
Source-checked
9GPT-5.5 (medium)OpenAI2026-04-2392.6%
2026-07-21
Source-checked
10GPT-5.6 TerraOpenAI2026-07-0992.5%
2026-07-21
Source-checked
11Qwen3.7 MaxAlibaba2026-05-1992.3%
2026-07-21
Source-checked
12Gemini 3.5 Flash (medium)Google2026-05-1992.1%
2026-07-21
Source-checked
13GPT 5.4OpenAI2026-03-0592%
2026-07-21
Source-checked
14GPT-5.3 Codex (xhigh)OpenAI2026-02-0591.5%
2026-07-21
Source-checked
15Claude Opus 4.7Anthropic2026-04-1691.4%
2026-07-21
Source-checked
16Claude Sonnet 5Anthropic2026-06-3091.1%
2026-07-21
Source-checked
17GPT-5.6 LunaOpenAI2026-07-0991.1%
2026-07-21
Source-checked
18Grok 4.20 ReasoningxAI2026-03-0591.1%
2026-07-21
Source-checked
19GPT-5.5 (low)OpenAI2026-04-2391%
2026-07-21
Source-checked
20Gemini 3 ProGoogle2025-11-1890.8%
2026-07-21
Source-checked
1Kimi K2.6Moonshot2026-04-2090.5%
GPQA Diamond paper leaderboard2026-05-15
Source-checked
22GPT-5.2OpenAI2025-12-1190.3%
2026-07-21
Source-checked
23Qwen 3.7 PlusAlibaba2026-06-0290%
2026-07-21
Source-checked
24Muse Spark 1.1Meta2026-07-0989.8%
2026-07-21
Source-checked
25Claude Opus 4.6 (Adaptive Reasoning, Max Effort)Anthropic2026-02-0589.6%
2026-07-21
Source-checked
26Kimi K2.7 CodeMoonshot2026-06-1289.6%
2026-07-21
Source-checked
27GLM-5.2Zhipu2026-06-1389.5%
2026-07-21
Source-checked
28DeepSeek V4 Pro (Reasoning, Max Effort)DeepSeek2026-04-2488.8%
2026-07-21
Source-checked
29Claude Opus 4.7 (Non-reasoning, High Effort)Anthropic2026-04-1688.5%
2026-07-21
Source-checked
30Muse SparkOther2026-01-0188.4%
2026-07-21
Source-checked
31Kimi K2.5Moonshot2026-01-2787.9%
2026-07-21
Source-checked
32GPT-5.4 MiniOpenAI2026-03-1787.5%
2026-07-21
Source-checked
33MiniMax M2.7MiniMax2026-04-1587.4%
2026-07-21
Source-checked
34GPT-5.1OpenAI2025-11-1387.3%
2026-07-21
Source-checked
35GLM-5.1Zhipu2026-04-0786.8%
2026-07-21
Source-checked
36Nemotron 3 Ultra 550B A55B (Reasoning)NVIDIA2026-06-0486.7%
2026-07-21
Source-checked
37MiMo-V2.5-ProXiaomi2026-04-2286.6%
2026-07-21
Source-checked
38MiniMax M2.5MiniMax2026-04-0184.8%
2026-07-21
Source-checked
39Claude Opus 4.6Anthropic2026-02-0584%
2026-07-21
Source-checked
40OpenAI o3OpenAI2025-04-1682.7%
2026-07-21
Source-checked
41GPT-5.4 NanoOpenAI2026-03-1781.7%
2026-07-21
Source-checked
42Claude Opus 4.5Anthropic2025-11-2481%
2026-07-21
Source-checked
43Step 3.7 FlashStepFun2026-06-0180.9%
2026-07-21
Source-checked
1Gemini 3.1 ProGoogleCoT prompting. PhD-level science MCQ.2026-02-1978.4%
Artificial Analysis GPQA Diamond evaluation2026-06-09
Needs audit
2Grok 4.3xAI2026-04-3088%
GPQA Diamond paper leaderboard2026-05-15
Source-checked
46Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)Anthropic2026-02-1787.5%
2026-07-21
Source-checked
47Gemma 4 31B ITGoogle2026-03-0185.7%
2026-07-21
Source-checked
48Gemini 3.1 Flash LiteGoogle2026-05-0782.2%
2026-07-21
Source-checked
49DeepSeek R1DeepSeek2025-01-2081.3%
2026-07-21
Source-checked
50LongCat-2.0Meituan2026-06-2978%
2026-07-21
Source-checked
2Claude Sonnet 4.6Anthropic2026-02-1777.2%
Artificial Analysis GPQA Diamond evaluation2026-06-09
Needs audit
3GPT 5.5OpenAI2026-04-2381.2%
GPQA Diamond paper leaderboard2026-05-15
Source-checked
53Gemini 3 Flash PreviewGoogle2025-12-1781.2%
2026-07-21
Source-checked
54North Mini CodeCohere2026-06-0975.7%
2026-07-21
Source-checked
3DeepSeek V4 Pro MaxDeepSeek2026-04-2475.6%
Artificial Analysis GPQA Diamond evaluation2026-06-09
Needs audit
4Claude Opus 4.8Anthropic2026-05-2879.8%
GPQA Diamond paper leaderboard2026-05-15
Source-checked
57Gemma 4 12B (Reasoning)Google2026-06-0775.3%
2026-07-21
Source-checked
58OpenAI o1OpenAI2024-09-1274.7%
2026-07-21
Source-checked
4Gemini 3.5 FlashGoogle2026-05-1974.1%
Artificial Analysis GPQA Diamond evaluation2026-06-09
Needs audit
5Qwen 3.7 MaxAlibaba2026-05-2073.8%
Artificial Analysis GPQA Diamond evaluation2026-06-09
Needs audit
61HyperNova 60B 2605Multiverse2026-05-0673.3%
2026-07-21
Source-checked
6Mistral Large 3Mistral2025-12-0271.2%
Artificial Analysis GPQA Diamond evaluation2026-06-09
Needs audit
7Llama 4 MaverickMeta2026-04-0569.4%
Artificial Analysis GPQA Diamond evaluation2026-06-09
Needs audit
64Gemma 4 12B (Non-reasoning)Google2026-06-1066.1%
2026-07-21
Source-checked
65Llama 4 ScoutMeta2026-04-0558.7%
2026-07-21
Source-checked
66Command ACohere2026-03-1552.7%
2026-07-21
Source-checked
67LFM2.5-8B-A1BLiquid2026-06-0851.3%
2026-07-21
Source-checked
68Claude 3 OpusAnthropic2024-03-0448.9%
2026-07-21
Source-checked
69MiniCPM5-1B (Reasoning)OpenBMB2026-06-0427.8%
2026-07-21
Source-checked
Method

What it covers

Data-quality note pending. Every ledger row still carries source URL, source date, and ingest timestamp.

Score ceiling

Where it breaks down

PhD-level domain experts score ~65-74% in their own field; skilled non-experts with web access ~34%.

Tools that report it

Receipts

Sources and further reading

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.