Models / DeepSeek

DeepSeek V4 Pro Max

DeepSeek · Released 2026-04-24

DeepSeek's April 2026 frontier open-weights model, leading the open-weights tier on SWE-bench Verified (80.6%) and Terminal-Bench 2.0 (67.9%) at $0.40 / $1.60 per 1M tokens. Cheapest frontier pricing on the public market.

Default model for self-hosted agent infrastructure, cost-sensitive deployments, and teams that need open weights with frontier-tier coding scores.

#9 of 21 on Vals Index · current snapshot
56 Vals Index · 14 benchmark rows · 10 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite74
Context256K
Input / 1M$0.40
Output / 1M$1.6
Knowledge cutoff2026-02-01
Statusflagship
KnowledgeReasoningMathCodingAgenticLong contextTool useSafetyHuman preference
Knowledge
90
Reasoning
76
Math
84
Coding
93
Agentic
68
Long context
49
Tool use
75
Safety
86
Human preference
84

78average across 9 tested families

Editor's note
V4 Pro Max is the model we use for self-hosted agent infrastructure where compute and data residency matter more than absolute best-in-class reasoning. The cost gap to Opus 4.8 is roughly 37× on input and 47× on output — and the open-weights license means you can run it on your own metal.
Benchmark placements

Where DeepSeek V4 Pro Max places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
LMArena (Chatbot Arena)Human preference#5/ 1384LMArena leaderboard2026-05-25
MCP AtlasTool use#6/ 1073MCP Atlas leaderboard2026-06-09
SWE-bench VerifiedCoding#7/ 2381SWE-bench official leaderboard2026-05-30
HELM SafetySafety#8/ 986Stanford HELM leaderboard2026-04-24
MMLUKnowledge#8/ 2290DeepSeek V4 release2026-04-24
HumanEvalCoding#9/ 1593DeepSeek V4 release2026-04-24
Vals IndexAgentic#9/ 2056Vals AI — Vals Index2026-06-04
MATHMath#12/ 2084DeepSeek V4 release2026-04-24
LiveBenchReasoning#14/ 2162LiveBench leaderboard2026-04-15
LongBench v2Long context#14/ 1949LongBench v2 leaderboard2026-04-01
MMLU-ProKnowledge#14/ 2176MMLU-Pro HuggingFace leaderboard2026-04-24
BFCLTool use#16/ 2075Berkeley Function Calling Leaderboard V42026-04-15
Terminal-BenchAgentic#25/ 6968Terminal-Bench 2.0 leaderboard2026-05-20
GPQA DiamondReasoning#56/ 6976Artificial Analysis GPQA Diamond evaluation2026-06-09
Family context

The three highest-scoring carded models in each capability family. Where DeepSeek V4 Pro Max shows up, it's highlighted.

KnowledgeDeepSeek V4 Pro Max · 90
  1. 1
    91
  2. 2
    91
  3. 3
    91
ReasoningDeepSeek V4 Pro Max · 76
  1. 1
    95
  2. 2
    94
  3. 3
    94
MathDeepSeek V4 Pro Max · 84
  1. 1
    99
  2. 2
    97
  3. 3
    96
CodingDeepSeek V4 Pro Max · 93
  1. 1
    96
  2. 2
    96
  3. 3
    96
AgenticDeepSeek V4 Pro Max · 68
  1. 1
    100
  2. 2
    100
  3. 3
    100
Long contextDeepSeek V4 Pro Max · 49
  1. 1
    75
  2. 2
    75
  3. 3
    74
Tool useDeepSeek V4 Pro Max · 75
  1. 1
    89
  2. 2
    87
  3. 3
    87
SafetyDeepSeek V4 Pro Max · 86
  1. 1
    92
  2. 2
    92
  3. 3
    91
Human preferenceDeepSeek V4 Pro Max · 84
  1. 1
    95
  2. 2
    92
  3. 3
    91
Newest receipts

Benchmark rows added to the public ledger for DeepSeek V4 Pro Max in the last 120 days. Older rows live in the full table below.

14 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Knowledge90
Reasoning76
Math84
Coding93
Agentic68
Long context49
Tool use75
Safety86
Human preference84
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover or tab any dot for name, score, and input price.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMiniMaxxAIMetaZhipuXiaomiMeituanCohere

Full benchmark ledger

Every catalog benchmark for DeepSeek V4 Pro Max. Scores link to the original source; gaps mean no public row exists yet.

14 sourced rows

14 of 32 catalog benchmarks have a sourced row for DeepSeek V4 Pro Max.

Knowledge

Reasoning

Math

Coding

Agentic

Long context

Tool use

Safety

Human preference

Changelog
  1. Released DeepSeek V4 Pro Max with open weights and frontier coding scores.
  2. DeepSeek V3.2 retired as the flagship open-weights model.
Sources