Why this benchmark is useful
When you want graduate-level science that non-experts struggle to Google — a narrow but still discriminating knowledge/reasoning slice.
Graduate-Level Google-Proof Q&A
Hard graduate-level science questions (biology, physics, chemistry) written so that non-experts cannot solve them even with web access — the 'Diamond' subset is the highest-quality, expert-validated slice.
When you want graduate-level science that non-experts struggle to Google — a narrow but still discriminating knowledge/reasoning slice.
Accuracy on 198 Diamond MCQs. Small set: a few lucky items move the number. Compare to expert (~65–74%) and web-aided non-expert (~34%) ceilings, not everyday usefulness.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Grok 4.6xAI | 2026-08-12 | 94.9% | 2026-09-01 | Source-checked |
| 1 | Claude Mythos PreviewAnthropic | 2026-06-01 | 94.6% | GPQA Diamond paper leaderboard2026-06-02 | Source-checked |
| 3 | Gemini 3.7 FlashGoogle | 2026-08-13 | 94.5% | 2026-09-01 | Source-checked |
| 4 | Gemini 3.1 Pro PreviewGoogle | 2026-02-19 | 94.1% | 2026-09-01 | Source-checked |
| 5 | GPT-5.6 SolOpenAI | 2026-07-09 | 94.1% | 2026-09-01 | Source-checked |
| 6 | Claude Opus 5 (Adaptive Reasoning, High Effort)Anthropic | 2026-07-24 | 93.7% | 2026-09-01 | Source-checked |
| 7 | Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)Anthropic | 2026-07-24 | 93.7% | 2026-09-01 | Source-checked |
| 8 | Kimi K3Moonshot | 2026-07-16 | 93.5% | 2026-09-01 | Source-checked |
| 9 | Qwen3.8 2.4T A95BAlibaba | 2026-08-12 | 93.5% | 2026-09-01 | Source-checked |
| 10 | Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic | 2026-07-24 | 93.2% | 2026-09-01 | Source-checked |
| 11 | GPT-5.5 (high)OpenAI | 2026-04-23 | 93.2% | 2026-09-01 | Source-checked |
| 12 | Grok 4.5xAI | 2026-07-08 | 93.1% | 2026-09-01 | Source-checked |
| 13 | MiniMax M3MiniMax | 2026-05-31 | 92.9% | 2026-09-01 | Source-checked |
| 14 | DeepSeek V4 Pro 0813 (Reasoning, Max Effort)DeepSeek | 2026-08-13 | 92.8% | 2026-09-01 | Source-checked |
| 15 | Gemini 3.6 Flash (high)Google | 2026-07-21 | 92.8% | 2026-09-01 | Source-checked |
| 16 | Qwen3.8 MaxAlibaba | 2026-08-03 | 92.7% | 2026-09-01 | Source-checked |
| 17 | Claude Fable 5Anthropic | 2026-06-09 | 92.6% | 2026-09-01 | Source-checked |
| 18 | GPT-5.5 (medium)OpenAI | 2026-04-23 | 92.6% | 2026-09-01 | Source-checked |
| 19 | GPT-5.6 TerraOpenAI | 2026-07-09 | 92.5% | 2026-09-01 | Source-checked |
| 20 | Qwen3.7 MaxAlibaba | 2026-05-19 | 92.3% | 2026-09-01 | Source-checked |
| 21 | Qwen3.8-Flash-NextAlibaba | 2026-08-26 | 92.3% | 2026-09-01 | Source-checked |
| 22 | Gemini 3.5 Flash (medium)Google | 2026-05-19 | 92.1% | 2026-09-01 | Source-checked |
| 23 | GPT 5.4OpenAI | 2026-03-05 | 92% | 2026-09-01 | Source-checked |
| 24 | Claude Opus 5 (Adaptive Reasoning, Medium Effort)Anthropic | 2026-07-24 | 91.9% | 2026-09-01 | Source-checked |
| 25 | GLM-5.3Zhipu | 2026-08-14 | 91.7% | 2026-09-01 | Source-checked |
| 26 | GPT-5.3 Codex (xhigh)OpenAI | 2026-02-05 | 91.5% | 2026-09-01 | Source-checked |
| 27 | Claude Opus 4.7Anthropic | 2026-04-16 | 91.4% | 2026-09-01 | Source-checked |
| 28 | DeepSeek V4 Flash Vision (Reasoning, Max Effort)DeepSeek | 2026-08-21 | 91.3% | 2026-09-01 | Source-checked |
| 29 | GLM-5.3-FlashZhipu | 2026-08-26 | 91.2% | 2026-09-01 | Source-checked |
| 30 | Claude Sonnet 5Anthropic | 2026-06-30 | 91.1% | 2026-09-01 | Source-checked |
| 31 | GPT-5.6 LunaOpenAI | 2026-07-09 | 91.1% | 2026-09-01 | Source-checked |
| 32 | Grok 4.20 ReasoningxAI | 2026-03-05 | 91.1% | 2026-09-01 | Source-checked |
| 33 | GPT-5.5 (low)OpenAI | 2026-04-23 | 91% | 2026-09-01 | Source-checked |
| 34 | DeepSeek V4 Flash 0731 (Reasoning, Max Effort)DeepSeek | 2026-07-31 | 90.8% | 2026-09-01 | Source-checked |
| 35 | Gemini 3 ProGoogle | 2025-11-18 | 90.8% | 2026-09-01 | Source-checked |
| 1 | Kimi K2.6Moonshot | 2026-04-20 | 90.5% | GPQA Diamond paper leaderboard2026-05-15 | Source-checked |
| 37 | Agnes 2.5 Pro BetaOther | 2026-08-26 | 90.5% | 2026-09-01 | Source-checked |
| 38 | DeepSeek V4 Pro (Reasoning, High Effort)DeepSeek | 2026-04-24 | 90.5% | 2026-09-01 | Source-checked |
| 39 | Qwen3.8 27BAlibaba | 2026-08-14 | 90.5% | 2026-09-01 | Source-checked |
| 40 | Muse Spark 1.2Meta | 2026-08-05 | 90.4% | 2026-09-01 | Source-checked |
| 41 | GPT-5.2OpenAI | 2025-12-11 | 90.3% | 2026-09-01 | Source-checked |
| 42 | Qwen 3.7 PlusAlibaba | 2026-06-02 | 90% | 2026-09-01 | Source-checked |
| 43 | GPT-5.2 Codex (xhigh)OpenAI | 2025-12-11 | 89.9% | 2026-09-01 | Source-checked |
| 44 | Gemini 3 Flash Preview (Reasoning)Google | 2025-12-17 | 89.8% | 2026-09-01 | Source-checked |
| 45 | Muse Spark 1.1Meta | 2026-07-09 | 89.8% | 2026-09-01 | Source-checked |
| 46 | Hy3Other | 2026-07-06 | 89.7% | 2026-09-01 | Source-checked |
| 47 | Claude Opus 4.6 (Adaptive Reasoning, Max Effort)Anthropic | 2026-02-05 | 89.6% | 2026-09-01 | Source-checked |
| 48 | Kimi K2.7 CodeMoonshot | 2026-06-12 | 89.6% | 2026-09-01 | Source-checked |
| 49 | GLM-5.2Zhipu | 2026-06-13 | 89.5% | 2026-09-01 | Source-checked |
| 50 | Grok Build 0.1 0616xAI | 2026-06-16 | 89.5% | 2026-09-01 | Source-checked |
| 51 | Inkling SmallOther | 2026-07-30 | 89.5% | 2026-09-01 | Source-checked |
| 52 | Qwen3.5 397B A17B (Reasoning)Alibaba | 2026-02-16 | 89.3% | 2026-09-01 | Source-checked |
| 53 | Nex-N2-ProOther | 2026-06-02 | 89.2% | 2026-09-01 | Source-checked |
| 54 | Solar Pro 4Other | 2026-08-06 | 89.1% | 2026-09-01 | Source-checked |
| 55 | Grok 4.3 (medium)xAI | 2026-04-30 | 89% | 2026-09-01 | Source-checked |
| 56 | Claude Opus 5 (Adaptive Reasoning, Low Effort)Anthropic | 2026-07-24 | 88.9% | 2026-09-01 | Source-checked |
| 57 | Qwen3.6 Max PreviewAlibaba | 2026-04-20 | 88.8% | 2026-09-01 | Source-checked |
| 58 | Gemini 3 Pro Preview (low)Google | 2025-11-18 | 88.7% | 2026-09-01 | Source-checked |
| 59 | Claude Opus 4.7 (Non-reasoning, High Effort)Anthropic | 2026-04-16 | 88.5% | 2026-09-01 | Source-checked |
| 60 | Grok 4.20 0309 (Reasoning)xAI | 2026-03-10 | 88.5% | 2026-09-01 | Source-checked |
| 61 | Muse SparkOther | 2026-01-01 | 88.4% | 2026-09-01 | Source-checked |
| 62 | Qwen3.6 PlusAlibaba | 2026-04-02 | 88.2% | 2026-09-01 | Source-checked |
| 63 | Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)Anthropic | 2026-02-17 | 87.5% | 2026-09-01 | Source-checked |
| 64 | GPT-5.4 MiniOpenAI | 2026-03-17 | 87.5% | 2026-09-01 | Source-checked |
| 65 | GPT-5.1OpenAI | 2025-11-13 | 87.3% | 2026-09-01 | Source-checked |
| 66 | GPT-5.4 (low)OpenAI | 2026-03-05 | 87.1% | 2026-09-01 | Source-checked |
| 67 | GLM-5.1Zhipu | 2026-04-07 | 86.8% | 2026-09-01 | Source-checked |
| 68 | DeepSeek V4 Flash (Reasoning, High Effort)DeepSeek | 2026-04-24 | 86.7% | 2026-09-01 | Source-checked |
| 69 | Claude Opus 4.5 (Reasoning)Anthropic | 2025-11-24 | 86.6% | 2026-09-01 | Source-checked |
| 70 | Apodex 1.1Other | 2026-08-30 | 86.4% | 2026-09-01 | Source-checked |
| 71 | GPT-5.2 (medium)OpenAI | 2025-12-11 | 86.4% | 2026-09-01 | Source-checked |
| 72 | GPT-5.1 Codex (high)OpenAI | 2025-11-13 | 86% | 2026-09-01 | Source-checked |
| 73 | GLM-4.7 (Reasoning)Zhipu | 2025-12-22 | 85.9% | 2026-09-01 | Source-checked |
| 74 | Gemma 4 31B ITGoogle | 2026-03-01 | 85.7% | 2026-09-01 | Source-checked |
| 75 | GPT-5 (high)OpenAI | 2025-08-07 | 85.4% | 2026-09-01 | Source-checked |
| 76 | GLM-5-TurboZhipu | 2026-03-15 | 84.7% | 2026-09-01 | Source-checked |
| 77 | GPT-5.5 Instant (May 2026)OpenAI | 2026-05-05 | 84.6% | 2026-09-01 | Source-checked |
| 78 | GPT-5 (medium)OpenAI | 2025-08-07 | 84.2% | 2026-09-01 | Source-checked |
| 79 | DeepSeek V3.2 (Reasoning)DeepSeek | 2025-12-01 | 84% | 2026-09-01 | Source-checked |
| 80 | Gemini 3.5 Flash-LiteGoogle | 2026-07-21 | 83.8% | 2026-09-01 | Source-checked |
| 81 | GPT-5 Codex (high)OpenAI | 2025-09-23 | 83.7% | 2026-09-01 | Source-checked |
| 82 | Gemini 3.5 Flash (minimal)Google | 2026-05-19 | 82.8% | 2026-09-01 | Source-checked |
| 83 | GPT-5.5 Instant (June 2026)OpenAI | 2026-06-25 | 82.3% | 2026-09-01 | Source-checked |
| 84 | Gemini 3.1 Flash LiteGoogle | 2026-05-07 | 82.2% | 2026-09-01 | Source-checked |
| 85 | GLM-5 (Reasoning)Zhipu | 2026-02-11 | 82% | 2026-09-01 | Source-checked |
| 86 | GPT-5.4 NanoOpenAI | 2026-03-17 | 81.7% | 2026-09-01 | Source-checked |
| 87 | DeepSeek R1DeepSeek | 2025-01-20 | 81.3% | 2026-09-01 | Source-checked |
| 88 | Gemini 3 Flash PreviewGoogle | 2025-12-17 | 81.2% | 2026-09-01 | Source-checked |
| 89 | Claude 4.1 Opus (Reasoning)Anthropic | 2025-08-05 | 80.9% | 2026-09-01 | Source-checked |
| 90 | GLM 5V Turbo (Reasoning)Zhipu | 2026-04-01 | 80.9% | 2026-09-01 | Source-checked |
| 91 | GPT-5 (low)OpenAI | 2025-08-07 | 80.8% | 2026-09-01 | Source-checked |
| 92 | G9v3-39A5BOther | 2026-08-03 | 80.5% | 2026-09-01 | Source-checked |
| 93 | Claude Sonnet 4.6 (Non-reasoning, Low Effort)Anthropic | 2026-02-17 | 79.7% | 2026-09-01 | Source-checked |
| 94 | Claude 4 Opus (Reasoning)Anthropic | 2025-05-22 | 79.6% | 2026-09-01 | Source-checked |
| 1 | Gemini 3.1 ProGoogleCoT prompting. PhD-level science MCQ. | 2026-02-19 | 78.4% | Artificial Analysis GPQA Diamond evaluation2026-06-09 | Needs audit |
| 2 | Grok 4.3xAI | 2026-04-30 | 88% | GPQA Diamond paper leaderboard2026-05-15 | Source-checked |
| 97 | Kimi K2.5Moonshot | 2026-01-27 | 87.9% | 2026-09-01 | Source-checked |
| 98 | Grok 4xAI | 2025-07-10 | 87.7% | 2026-09-01 | Source-checked |
| 99 | Agnes 2.5 Pro AlphaOther | 2026-07-24 | 87.6% | 2026-09-01 | Source-checked |
| 100 | MiniMax M2.7MiniMax | 2026-04-15 | 87.4% | 2026-09-01 | Source-checked |
| 101 | Inkling (xhigh)Other | 2026-07-15 | 87.2% | 2026-09-01 | Source-checked |
| 102 | MiMo-V2-ProXiaomi | 2026-03-18 | 87% | 2026-09-01 | Source-checked |
| 103 | Motif 3 (Beta)Other | 2026-07-21 | 86.9% | 2026-09-01 | Source-checked |
| 104 | Hy3-preview (Reasoning)Other | 2026-04-23 | 86.7% | 2026-09-01 | Source-checked |
| 105 | Nemotron 3 Ultra 550B A55B (Reasoning)NVIDIA | 2026-06-04 | 86.7% | 2026-09-01 | Source-checked |
| 106 | MiMo-V2.5-ProXiaomi | 2026-04-22 | 86.6% | 2026-09-01 | Source-checked |
| 107 | Qwen3 Max ThinkingAlibaba | 2026-01-26 | 86.1% | 2026-09-01 | Source-checked |
| 108 | Qwen3.5 397B A17B (Non-reasoning)Alibaba | 2026-02-16 | 86.1% | 2026-09-01 | Source-checked |
| 109 | Qwen3.5 27B (Reasoning)Alibaba | 2026-02-24 | 85.8% | 2026-09-01 | Source-checked |
| 110 | A.X-K2Other | 2026-08-12 | 85.7% | 2026-09-01 | Source-checked |
| 111 | Qwen3.5 122B A10B (Reasoning)Alibaba | 2026-02-24 | 85.7% | 2026-09-01 | Source-checked |
| 112 | Ring-2.6-1TOther | 2026-05-08 | 85.7% | 2026-09-01 | Source-checked |
| 113 | Solar Open2 250BOther | 2026-08-12 | 85.7% | 2026-09-01 | Source-checked |
| 114 | KAT Coder Pro V2Other | 2026-03-27 | 85.5% | 2026-09-01 | Source-checked |
| 115 | Ling 3.0 FlashOther | 2026-08-04 | 85.5% | 2026-09-01 | Source-checked |
| 116 | MiMo-V2-Omni-0327Xiaomi | 2026-03-27 | 85.5% | 2026-09-01 | Source-checked |
| 117 | MiMo-V2.5Xiaomi | 2026-04-22 | 84.9% | 2026-09-01 | Source-checked |
| 118 | MiniMax M2.5MiniMax | 2026-04-01 | 84.8% | 2026-09-01 | Source-checked |
| 119 | MiMo-V2-Flash (Reasoning)Xiaomi | 2025-12-16 | 84.6% | 2026-09-01 | Source-checked |
| 120 | JT-4.1 Flash 236B A21BOther | 2026-07-09 | 84.5% | 2026-09-01 | Source-checked |
| 121 | o3-proOpenAI | 2025-06-10 | 84.5% | 2026-09-01 | Source-checked |
| 122 | Grok 4.3 (low)xAI | 2026-04-30 | 84.3% | 2026-09-01 | Source-checked |
| 123 | Kimi K3 (low)Moonshot | 2026-07-16 | 84.2% | 2026-09-01 | Source-checked |
| 124 | Qwen3.6 27B (Reasoning)Alibaba | 2026-04-22 | 84.2% | 2026-09-01 | Source-checked |
| 125 | Qwen3.6 35B A3B (Reasoning)Alibaba | 2026-04-16 | 84.1% | 2026-09-01 | Source-checked |
| 126 | Claude Opus 4.6Anthropic | 2026-02-05 | 84% | 2026-09-01 | Source-checked |
| 127 | Kimi K2 ThinkingMoonshot | 2025-11-06 | 83.8% | 2026-09-01 | Source-checked |
| 128 | MiMo-V2-Flash (Feb 2026)Xiaomi | 2025-12-16 | 83.5% | 2026-09-01 | Source-checked |
| 129 | Muse GlimmerMeta | 2026-08-10 | 83.5% | 2026-09-01 | Source-checked |
| 130 | Claude 4.5 Sonnet (Reasoning)Anthropic | 2025-09-29 | 83.4% | 2026-09-01 | Source-checked |
| 131 | Motif 3Other | 2026-08-12 | 83.4% | 2026-09-01 | Source-checked |
| 132 | MiniMax-M2.1MiniMax | 2025-12-23 | 83% | 2026-09-01 | Source-checked |
| 133 | JT-35B-FlashOther | 2026-05-14 | 82.9% | 2026-09-01 | Source-checked |
| 134 | K-EXAONE 2.0 0803Other | 2026-08-12 | 82.9% | 2026-09-01 | Source-checked |
| 135 | Qwen3.6 27B (Non-reasoning)Alibaba | 2026-04-22 | 82.9% | 2026-09-01 | Source-checked |
| 136 | MiMo-V2-OmniXiaomi | 2026-03-19 | 82.8% | 2026-09-01 | Source-checked |
| 137 | OpenAI o3OpenAI | 2025-04-16 | 82.7% | 2026-09-01 | Source-checked |
| 138 | Qwen3.6 35B A3B (Non-reasoning)Alibaba | 2026-04-16 | 81.7% | 2026-09-01 | Source-checked |
| 139 | EXAONE 4.5 33BOther | 2026-04-09 | 79.4% | 2026-09-01 | Source-checked |
| 2 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 77.2% | Artificial Analysis GPQA Diamond evaluation2026-06-09 | Needs audit |
| 3 | GPT 5.5OpenAI | 2026-04-23 | 81.2% | GPQA Diamond paper leaderboard2026-05-15 | Source-checked |
| 142 | Claude Opus 4.5Anthropic | 2025-11-24 | 81% | 2026-09-01 | Source-checked |
| 143 | Step 3.7 FlashStepFun | 2026-06-01 | 80.9% | 2026-09-01 | Source-checked |
| 3 | DeepSeek V4 Pro MaxDeepSeek | 2026-04-24 | 75.6% | Artificial Analysis GPQA Diamond evaluation2026-06-09 | Needs audit |
| 4 | Claude Opus 4.8Anthropic | 2026-05-28 | 79.8% | GPQA Diamond paper leaderboard2026-05-15 | Source-checked |
| 146 | Kimi K2.6 (Non-reasoning)Other | 2026-04-20 | 78.8% | 2026-09-01 | Source-checked |
| 147 | LongCat-2.0Meituan | 2026-06-29 | 78% | 2026-09-01 | Source-checked |
| 148 | Grok 4.20 0309 v2 (Non-reasoning)xAI | 2026-04-07 | 77.6% | 2026-09-01 | Source-checked |
| 149 | GPT-5.5 (Non-reasoning)OpenAI | 2026-04-23 | 76.8% | 2026-09-01 | Source-checked |
| 150 | MiMo-V2.5-Pro (Non-reasoning)Xiaomi | 2026-04-22 | 76.2% | 2026-09-01 | Source-checked |
| 151 | Command ACohere | 2026-03-15 | 76.1% | 2026-09-01 | Source-checked |
| 152 | North Mini CodeCohere | 2026-06-09 | 75.7% | 2026-09-01 | Source-checked |
| 153 | Gemma 4 12B (Reasoning)Google | 2026-06-07 | 75.3% | 2026-09-01 | Source-checked |
| 154 | Ling-2.6-1TOther | 2026-04-23 | 75.2% | 2026-09-01 | Source-checked |
| 155 | Mistral Medium 3.5Mistral | 2026-04-29 | 74.8% | 2026-09-01 | Source-checked |
| 156 | OpenAI o1OpenAI | 2024-09-12 | 74.7% | 2026-09-01 | Source-checked |
| 157 | Nemotron 3.5 LightningNVIDIA | 2026-08-11 | 74.3% | 2026-09-01 | Source-checked |
| 4 | Gemini 3.5 FlashGoogle | 2026-05-19 | 74.1% | Artificial Analysis GPQA Diamond evaluation2026-06-09 | Needs audit |
| 5 | Qwen 3.7 MaxAlibaba | 2026-05-20 | 73.8% | Artificial Analysis GPQA Diamond evaluation2026-06-09 | Needs audit |
| 160 | Ling 3.0 TinyOther | 2026-08-06 | 73.4% | 2026-09-01 | Source-checked |
| 161 | HyperNova 60B 2605Multiverse | 2026-05-06 | 73.3% | 2026-09-01 | Source-checked |
| 162 | Hy3-preview (Non-reasoning)Other | 2026-04-23 | 73.2% | 2026-09-01 | Source-checked |
| 163 | DeepSeek V4 Pro (Non-reasoning)DeepSeek | 2026-04-24 | 71.7% | 2026-09-01 | Source-checked |
| 164 | DeepSeek V4 Flash (Non-reasoning)DeepSeek | 2026-04-24 | 71.6% | 2026-09-01 | Source-checked |
| 6 | Mistral Large 3Mistral | 2025-12-02 | 71.2% | Artificial Analysis GPQA Diamond evaluation2026-06-09 | Needs audit |
| 7 | Llama 4 MaverickMeta | 2026-04-05 | 69.4% | Artificial Analysis GPQA Diamond evaluation2026-06-09 | Needs audit |
| 167 | GLM-5.2 (Non-reasoning)Zhipu | 2026-06-16 | 68.6% | 2026-09-01 | Source-checked |
| 168 | JT-MINIOther | 2026-04-15 | 67.6% | 2026-09-01 | Source-checked |
| 169 | DiffusionGemma 26B A4BGoogle | 2026-06-10 | 66.9% | 2026-09-01 | Source-checked |
| 170 | GLM-5 (Non-reasoning)Zhipu | 2026-02-11 | 66.6% | 2026-09-01 | Source-checked |
| 171 | Gemma 4 12B (Non-reasoning)Google | 2026-06-10 | 66.1% | 2026-09-01 | Source-checked |
| 172 | Grok 4.3 (Non-reasoning)xAI | 2026-04-30 | 65.8% | 2026-09-01 | Source-checked |
| 173 | Granite 4.2 30BOther | 2026-08-25 | 64.4% | 2026-09-01 | Source-checked |
| 174 | Granite 4.2 8BOther | 2026-08-25 | 63.1% | 2026-09-01 | Source-checked |
| 175 | Ling 2.6 FlashOther | 2026-04-21 | 59.3% | 2026-09-01 | Source-checked |
| 176 | Llama 4 ScoutMeta | 2026-04-05 | 58.7% | 2026-09-01 | Source-checked |
| 177 | Granite 4.2 3BOther | 2026-08-25 | 55.9% | 2026-09-01 | Source-checked |
| 178 | LFM2.5-8B-A1BLiquid | 2026-06-08 | 51.3% | 2026-09-01 | Source-checked |
| 179 | Claude 3 OpusAnthropic | 2024-03-04 | 48.9% | 2026-09-01 | Source-checked |
| 180 | Granite 4.1 30BOther | 2026-04-29 | 48.1% | 2026-09-01 | Source-checked |
| 181 | Nemotron 3 Nano Omni 30B A3B ReasoningNVIDIA | 2026-04-29 | 46.9% | 2026-09-01 | Source-checked |
| 182 | G9v3-3BOther | 2026-07-23 | 43.8% | 2026-09-01 | Source-checked |
| 183 | MiniCPM-V 4.6 1.3BOpenBMB | 2026-05-11 | 30.5% | 2026-09-01 | Source-checked |
| 184 | MiniCPM5-1B (Reasoning)OpenBMB | 2026-06-04 | 27.8% | 2026-09-01 | Source-checked |
| 185 | MiniCPM5-1B (Non-reasoning)OpenBMB | 2026-05-25 | 26.9% | 2026-09-01 | Source-checked |
This benchmark sits in the reasoning family. It uses 198 expert-validated multiple-choice questions in the Diamond set; designed to be 'Google-proof'. The score should travel with its task format, scoring method, source date, and benchmark version.
PhD-level domain experts score ~65-74% in their own field; skilled non-experts with web access ~34%.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
Hard graduate-level science questions (biology, physics, chemistry) written so that non-experts cannot solve them even with web access — the 'Diamond' subset is the highest-quality, expert-validated slice.
A strong GPQA Diamond result says nothing about:
Accuracy (% correct). Task format: 198 expert-validated multiple-choice questions in the Diamond set; designed to be 'Google-proof'.
GPQA Diamond is currently marked Active in the atlas. Ceiling context: PhD-level domain experts score ~65-74% in their own field; skilled non-experts with web access ~34%.
Contamination risk for GPQA Diamond is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.