Warum dieser Benchmark nützlich ist
Er liefert ein kompaktes Signal für eine bestimmte Fähigkeit. Nutze ihn als datierten Beleg neben Preisen, Datenschutz und eigener Nutzung.
Benchmarks / Wissen
MMLU
Breites Multiple-Choice-Wissen über 57 Fächer — von US-Geschichte über Hochschulphysik bis Jura — im Vier-Optionen-Prüfungsformat.
Er liefert ein kompaktes Signal für eine bestimmte Fähigkeit. Nutze ihn als datierten Beleg neben Preisen, Datenschutz und eigener Nutzung.
Lies MMLU als gesättigt Signal mit hohes kontaminationsrisiko. Vergleiche Modelle nur, wenn Quelle, Testumgebung, Prompting, Sampling und Score-Einheit übereinstimmen.
Auf die Skala dieses Benchmarks normalisiert (0–100). Die Tabelle zeigt die Rohwerte.
Scores nutzen die eigene Einheit und Skala dieses Benchmarks, keinen universellen Qualitätswert.
| # | Modell | Veröffentlichung | Score | Herkunft | Beleglage |
|---|---|---|---|---|---|
| 1 | Claude Opus 4.8Anthropic5-shot CoT. MMLU is saturated at the frontier. | 2026-05-28 | 91.4% | Anthropic Claude 4 family2026-05-28 | Prüfung nötig |
| 2 | Claude Opus 4.6AnthropicReported with 5-shot chain-of-thought prompting. | 2026-02-05 | 91.2% | Anthropic Claude 3.7/4.6 model card2026-01-01 | Quelle geprüft |
| 3 | GPT 5.5OpenAIMMLU is saturated at the frontier — high scores are table stakes, not differentiators. | 2026-04-23 | 90.8% | OpenAI model release notes2025-12-01 | Prüfung nötig |
| 4 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 90.6% | Anthropic Claude 3.7/4.6 model card2026-02-17 | Quelle geprüft |
| 5 | Gemini 3.1 ProGoogle | 2026-02-19 | 90.3% | Google Gemini 3.1 Pro2026-02-19 | Prüfung nötig |
| 6 | Gemini 1.5 Ultra (2024)GoogleReported with CoT@32 voting, not single-shot. | 2024-02-15 | 90% | Gemini technical report2023-12-06 | Quelle geprüft |
| 7 | GPT 5.4OpenAI | 2026-03-05 | 89.8% | OpenAI GPT-5.4 release2026-03-05 | Prüfung nötig |
| 8 | DeepSeek V4 Pro MaxDeepSeek | 2026-04-24 | 89.7% | DeepSeek V4 release2026-04-24 | Prüfung nötig |
| 9 | Kimi K2.6Moonshot | 2026-04-20 | 89.1% | Moonshot Kimi K2.62026-04-20 | Prüfung nötig |
| 10 | Qwen 3.7 MaxAlibaba | 2026-05-20 | 88.9% | Alibaba Qwen 3.7 Max2026-05-20 | Prüfung nötig |
| 11 | Grok 4.3xAI | 2026-04-30 | 88.5% | xAI Grok 4.32026-04-30 | Prüfung nötig |
| 12 | Gemini 3.5 FlashGoogle | 2026-05-19 | 88.4% | Google Gemini 3.5 Flash2026-05-19 | Prüfung nötig |
| 13 | Mistral Large 3Mistral | 2025-12-02 | 87.8% | Mistral Large 3 model card2025-12-02 | Prüfung nötig |
| 14 | DeepSeek R1DeepSeek | 2025-01-20 | 87.4% | DeepSeek R1 model card2025-01-20 | Prüfung nötig |
| 15 | Gemini 3 FlashGoogle | 2025-12-17 | 87.2% | Google Gemini 3 Flash2025-12-17 | Prüfung nötig |
| 16 | Claude 3 OpusAnthropic | 2024-03-04 | 86.8% | Anthropic Claude 3 model card2024-03-04 | Quelle geprüft |
| 17 | GPT-4 (2023)OpenAI | 2023-03-14 | 86.4% | GPT-4 Technical Report2023-03-15 | Quelle geprüft |
| 18 | Kimi K2.5Moonshot | 2026-01-27 | 86.4% | Moonshot Kimi K2.52026-01-27 | Prüfung nötig |
| 19 | Llama 4 MaverickMeta | 2026-04-05 | 86.1% | Meta Llama 4 model card2026-04-05 | Prüfung nötig |
| 20 | Claude Haiku 4.5Anthropic | 2025-10-01 | 85.6% | Anthropic Claude Haiku 4.52025-10-01 | Prüfung nötig |
| 21 | Llama 4 ScoutMeta | 2026-04-05 | 85.2% | Meta Llama 4 Scout model card2026-04-05 | Prüfung nötig |
| 22 | MiniMax M3MiniMax | 2026-05-31 | 84.8% | MiniMax M32026-05-31 | Prüfung nötig |
Gesättigter Glossar-Eintrag — nützlich für Kontext, nicht für Frontier-Trennung. Inline-Vendor-Zeilen können hinter dem Ledger liegen; Decken-Clustering ist die ehrliche Lesart.
Experten ~89,8 %; Spitzenmodelle liegen nahe der Decke, MMLU trennt die besten Systeme kaum noch.
The Pack · Redaktionsnewsletter
Eine kurze E-Mail, wenn eine Karte erscheint oder den Status wechselt. Kein Tracking, keine Drittanbieter-Analytics. Abmeldung mit einem Klick.