Modelle / DeepSeek

DeepSeek V4 Pro Max

DeepSeek · Release 2026-04-24

DeepSeeks Frontier-Open-Weights-Modell April 2026, führend im Open-Weights-Segment bei SWE-bench Verified (80,6 %) und Terminal-Bench 2.0 (67,9 %) zu $0,40 / $1,60 pro 1M Tokens. Günstigster Frontier-Preis auf dem öffentlichen Markt.

Default-Modell für selbstgehostete Agent-Infrastruktur, kostenkritische Deployments und Teams, die Open Weights mit Frontier-Coding-Scores brauchen.

#9 von 21 im Vals Index · aktueller Snapshot
56 Vals Index · 14 Benchmark-Zeilen · 10 QuellenPosition ist benchmark-spezifisch — kein Cross-Family- oder Cross-Source-Ranking.
Composite74
Kontext256K
Input / 1M$0.40
Output / 1M$1.6
Wissensstichtag2026-02-01
Statusflagship
KnowledgeReasoningMathCodingAgenticLong contextTool useSafetyHuman preference
Knowledge
90
Reasoning
76
Math
84
Coding
93
Agentic
68
Long context
49
Tool use
75
Safety
86
Human preference
84

78average across 9 tested families

Redaktionsnotiz
V4 Pro Max ist das Modell, das wir für selbstgehostete Agent-Infrastruktur nutzen, wo Compute und Data-Residency wichtiger sind als absolutes Best-in-Class-Reasoning. Die Preislücke zu Opus 4.8 beträgt etwa 37× beim Input und 47× beim Output — und die Open-Weights-Lizenz erlaubt den Betrieb auf eigener Hardware.
Benchmark-Platzierungen

Wo DeepSeek V4 Pro Max auf jeder öffentlichen Benchmark-Quelle mit Zeile steht. Rang zählt jedes Modell mit neuester Zeile im selben Test — kein universeller Qualitätsscore.

BenchmarkFamilyRangScoreQuelleDatum
LMArena (Chatbot Arena)Human preference#5/ 1384LMArena leaderboard2026-05-25
MCP AtlasTool use#6/ 1073MCP Atlas leaderboard2026-06-09
SWE-bench VerifiedCoding#7/ 2381SWE-bench official leaderboard2026-05-30
HELM SafetySafety#8/ 986Stanford HELM leaderboard2026-04-24
MMLUKnowledge#8/ 2290DeepSeek V4 release2026-04-24
HumanEvalCoding#9/ 1593DeepSeek V4 release2026-04-24
Vals IndexAgentic#9/ 2056Vals AI — Vals Index2026-06-04
MATHMath#12/ 2084DeepSeek V4 release2026-04-24
LiveBenchReasoning#14/ 2162LiveBench leaderboard2026-04-15
LongBench v2Long context#14/ 1949LongBench v2 leaderboard2026-04-01
MMLU-ProKnowledge#14/ 2176MMLU-Pro HuggingFace leaderboard2026-04-24
BFCLTool use#16/ 2075Berkeley Function Calling Leaderboard V42026-04-15
Terminal-BenchAgentic#25/ 6968Terminal-Bench 2.0 leaderboard2026-05-20
GPQA DiamondReasoning#56/ 6976Artificial Analysis GPQA Diamond evaluation2026-06-09
Family-Kontext

Die drei höchstscorierenden kartierten Modelle pro Capability-Family. Wo DeepSeek V4 Pro Max auftaucht, ist es hervorgehoben.

KnowledgeDeepSeek V4 Pro Max · 90
  1. 1
    91
  2. 2
    91
  3. 3
    91
ReasoningDeepSeek V4 Pro Max · 76
  1. 1
    95
  2. 2
    94
  3. 3
    94
MathDeepSeek V4 Pro Max · 84
  1. 1
    99
  2. 2
    97
  3. 3
    96
CodingDeepSeek V4 Pro Max · 93
  1. 1
    96
  2. 2
    96
  3. 3
    96
AgenticDeepSeek V4 Pro Max · 68
  1. 1
    100
  2. 2
    100
  3. 3
    100
Long contextDeepSeek V4 Pro Max · 49
  1. 1
    75
  2. 2
    75
  3. 3
    74
Tool useDeepSeek V4 Pro Max · 75
  1. 1
    89
  2. 2
    87
  3. 3
    87
SafetyDeepSeek V4 Pro Max · 86
  1. 1
    92
  2. 2
    92
  3. 3
    91
Human preferenceDeepSeek V4 Pro Max · 84
  1. 1
    95
  2. 2
    92
  3. 3
    91
Neueste Belege

Benchmark-Zeilen im öffentlichen Ledger für DeepSeek V4 Pro Max in den letzten 120 Tagen. Ältere Zeilen stehen in der vollen Tabelle unten.

14 neu
Family-Abdeckung

Ein hoher Composite, der eine schwache Family versteckt, ist eine Falle. Diese Balken zeigen Families ohne öffentlichen Test — und wo das Modell führt.

Knowledge90
Reasoning76
Math84
Coding93
Agentic68
Long context49
Tool use75
Safety86
Human preference84
Preis vs. Performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover or tab any dot for name, score, and input price.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMiniMaxxAIMetaZhipuXiaomiMeituanCohere

Volles Benchmark-Ledger

Jeder Katalog-Benchmark für DeepSeek V4 Pro Max. Scores verlinken zur Originalquelle; Lücken heißen: noch keine öffentliche Zeile.

14 Quell-Zeilen

14 of 32 catalog benchmarks have a sourced row for DeepSeek V4 Pro Max.

Knowledge

Reasoning

Math

Coding

Agentic

Long context

Tool use

Safety

Human preference

Changelog
  1. DeepSeek V4 Pro Max mit Open Weights und Frontier-Coding-Scores veröffentlicht.
  2. DeepSeek V3.2 als Flaggschiff-Open-Weights-Modell ausgemustert.
Quellen