89.7%
MMLU
Models / DeepSeek
DeepSeek · Released 2026-04-24
DeepSeek's April 2026 frontier open-weights model, leading the open-weights tier on SWE-bench Verified (80.6%) and Terminal-Bench 2.0 (67.9%) at $0.40 / $1.60 per 1M tokens. Cheapest frontier pricing on the public market.
Default model for self-hosted agent infrastructure, cost-sensitive deployments, and teams that need open weights with frontier-tier coding scores.
78average across 9 tested families
Where DeepSeek V4 Pro Max places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.
| Benchmark | Family | Rank | Score | Source | Date |
|---|---|---|---|---|---|
| LMArena (Chatbot Arena) | Human preference | #5/ 13 | 84 | LMArena leaderboard | 2026-05-25 |
| MCP Atlas | Tool use | #6/ 10 | 73 | MCP Atlas leaderboard | 2026-06-09 |
| SWE-bench Verified | Coding | #7/ 23 | 81 | SWE-bench official leaderboard | 2026-05-30 |
| HELM Safety | Safety | #8/ 9 | 86 | Stanford HELM leaderboard | 2026-04-24 |
| MMLU | Knowledge | #8/ 22 | 90 | DeepSeek V4 release | 2026-04-24 |
| HumanEval | Coding | #9/ 15 | 93 | DeepSeek V4 release | 2026-04-24 |
| Vals Index | Agentic | #9/ 20 | 56 | Vals AI — Vals Index | 2026-06-04 |
| MATH | Math | #12/ 20 | 84 | DeepSeek V4 release | 2026-04-24 |
| LiveBench | Reasoning | #14/ 21 | 62 | LiveBench leaderboard | 2026-04-15 |
| LongBench v2 | Long context | #14/ 19 | 49 | LongBench v2 leaderboard | 2026-04-01 |
| MMLU-Pro | Knowledge | #14/ 21 | 76 | MMLU-Pro HuggingFace leaderboard | 2026-04-24 |
| BFCL | Tool use | #16/ 20 | 75 | Berkeley Function Calling Leaderboard V4 | 2026-04-15 |
| Terminal-Bench | Agentic | #25/ 69 | 68 | Terminal-Bench 2.0 leaderboard | 2026-05-20 |
| GPQA Diamond | Reasoning | #56/ 69 | 76 | Artificial Analysis GPQA Diamond evaluation | 2026-06-09 |
The three highest-scoring carded models in each capability family. Where DeepSeek V4 Pro Max shows up, it's highlighted.
Benchmark rows added to the public ledger for DeepSeek V4 Pro Max in the last 120 days. Older rows live in the full table below.
| Benchmark | Family | Score | Source | Days ago |
|---|---|---|---|---|
| GPQA Diamond | Reasoning | 75.6% | Artificial Analysis GPQA Diamond evaluation | 44 |
| MCP Atlas | Tool use | 72.6% | MCP Atlas leaderboard | 44 |
| Vals Index | Agentic | 56.23% | Vals AI — Vals Index | 49 |
| SWE-bench Verified | Coding | 80.6% | SWE-bench official leaderboard | 54 |
| LMArena (Chatbot Arena) | Human preference | 1418 | LMArena leaderboard | 59 |
| Terminal-Bench | Agentic | 67.9% | Terminal-Bench 2.0 leaderboard | 64 |
| HELM Safety | Safety | 86.4% | Stanford HELM leaderboard | 90 |
| HumanEval | Coding | 93.4% pass@1 | DeepSeek V4 release | 90 |
| MATH | Math | 83.9% | DeepSeek V4 release | 90 |
| MMLU-Pro | Knowledge | 76.4% | MMLU-Pro HuggingFace leaderboard | 90 |
| MMLU | Knowledge | 89.7% | DeepSeek V4 release | 90 |
| BFCL | Tool use | 74.6% | Berkeley Function Calling Leaderboard V4 | 99 |
| LiveBench | Reasoning | 61.9% | LiveBench leaderboard | 99 |
| LongBench v2 | Long context | 48.6% | LongBench v2 leaderboard | 113 |
A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.
Hover or tab any dot for name, score, and input price.
Every catalog benchmark for DeepSeek V4 Pro Max. Scores link to the original source; gaps mean no public row exists yet.
14 of 32 catalog benchmarks have a sourced row for DeepSeek V4 Pro Max.
75.6%
GPQA Diamond
61.9%
LiveBench
83.9%
MATH
93.4% pass@1
HumanEval
80.6%
SWE-bench Verified
67.9%
Terminal-Bench 2.0 (audited harness)
56.23%
Vals Index
48.6%
LongBench v2
86.4%
HELM Safety
1418
LMArena Elo (Style Controlled)