89.7%
MMLU
Modelle / DeepSeek
DeepSeek · Release 2026-04-24
DeepSeeks Frontier-Open-Weights-Modell April 2026, führend im Open-Weights-Segment bei SWE-bench Verified (80,6 %) und Terminal-Bench 2.0 (67,9 %) zu $0,40 / $1,60 pro 1M Tokens. Günstigster Frontier-Preis auf dem öffentlichen Markt.
Default-Modell für selbstgehostete Agent-Infrastruktur, kostenkritische Deployments und Teams, die Open Weights mit Frontier-Coding-Scores brauchen.
78average across 9 tested families
Wo DeepSeek V4 Pro Max auf jeder öffentlichen Benchmark-Quelle mit Zeile steht. Rang zählt jedes Modell mit neuester Zeile im selben Test — kein universeller Qualitätsscore.
| Benchmark | Family | Rang | Score | Quelle | Datum |
|---|---|---|---|---|---|
| LMArena (Chatbot Arena) | Human preference | #5/ 13 | 84 | LMArena leaderboard | 2026-05-25 |
| MCP Atlas | Tool use | #6/ 10 | 73 | MCP Atlas leaderboard | 2026-06-09 |
| SWE-bench Verified | Coding | #7/ 23 | 81 | SWE-bench official leaderboard | 2026-05-30 |
| HELM Safety | Safety | #8/ 9 | 86 | Stanford HELM leaderboard | 2026-04-24 |
| MMLU | Knowledge | #8/ 22 | 90 | DeepSeek V4 release | 2026-04-24 |
| HumanEval | Coding | #9/ 15 | 93 | DeepSeek V4 release | 2026-04-24 |
| Vals Index | Agentic | #9/ 20 | 56 | Vals AI — Vals Index | 2026-06-04 |
| MATH | Math | #12/ 20 | 84 | DeepSeek V4 release | 2026-04-24 |
| LiveBench | Reasoning | #14/ 21 | 62 | LiveBench leaderboard | 2026-04-15 |
| LongBench v2 | Long context | #14/ 19 | 49 | LongBench v2 leaderboard | 2026-04-01 |
| MMLU-Pro | Knowledge | #14/ 21 | 76 | MMLU-Pro HuggingFace leaderboard | 2026-04-24 |
| BFCL | Tool use | #16/ 20 | 75 | Berkeley Function Calling Leaderboard V4 | 2026-04-15 |
| Terminal-Bench | Agentic | #25/ 69 | 68 | Terminal-Bench 2.0 leaderboard | 2026-05-20 |
| GPQA Diamond | Reasoning | #56/ 69 | 76 | Artificial Analysis GPQA Diamond evaluation | 2026-06-09 |
Die drei höchstscorierenden kartierten Modelle pro Capability-Family. Wo DeepSeek V4 Pro Max auftaucht, ist es hervorgehoben.
Benchmark-Zeilen im öffentlichen Ledger für DeepSeek V4 Pro Max in den letzten 120 Tagen. Ältere Zeilen stehen in der vollen Tabelle unten.
| Benchmark | Family | Score | Quelle | Tage her |
|---|---|---|---|---|
| GPQA Diamond | Reasoning | 75.6% | Artificial Analysis GPQA Diamond evaluation | 44 |
| MCP Atlas | Tool use | 72.6% | MCP Atlas leaderboard | 44 |
| Vals Index | Agentic | 56.23% | Vals AI — Vals Index | 49 |
| SWE-bench Verified | Coding | 80.6% | SWE-bench official leaderboard | 54 |
| LMArena (Chatbot Arena) | Human preference | 1418 | LMArena leaderboard | 59 |
| Terminal-Bench | Agentic | 67.9% | Terminal-Bench 2.0 leaderboard | 64 |
| HELM Safety | Safety | 86.4% | Stanford HELM leaderboard | 90 |
| HumanEval | Coding | 93.4% pass@1 | DeepSeek V4 release | 90 |
| MATH | Math | 83.9% | DeepSeek V4 release | 90 |
| MMLU-Pro | Knowledge | 76.4% | MMLU-Pro HuggingFace leaderboard | 90 |
| MMLU | Knowledge | 89.7% | DeepSeek V4 release | 90 |
| BFCL | Tool use | 74.6% | Berkeley Function Calling Leaderboard V4 | 99 |
| LiveBench | Reasoning | 61.9% | LiveBench leaderboard | 99 |
| LongBench v2 | Long context | 48.6% | LongBench v2 leaderboard | 113 |
Ein hoher Composite, der eine schwache Family versteckt, ist eine Falle. Diese Balken zeigen Families ohne öffentlichen Test — und wo das Modell führt.
Hover or tab any dot for name, score, and input price.
Jeder Katalog-Benchmark für DeepSeek V4 Pro Max. Scores verlinken zur Originalquelle; Lücken heißen: noch keine öffentliche Zeile.
14 of 32 catalog benchmarks have a sourced row for DeepSeek V4 Pro Max.
75.6%
GPQA Diamond
61.9%
LiveBench
83.9%
MATH
93.4% pass@1
HumanEval
80.6%
SWE-bench Verified
67.9%
Terminal-Bench 2.0 (audited harness)
56.23%
Vals Index
48.6%
LongBench v2
86.4%
HELM Safety
1418
LMArena Elo (Style Controlled)