VerdictPal · editorial desk · 2026VerdictPal
Compare · Two models, head-to-head

Pick two models. See the diff.

Composite scores, benchmark ranks, capability grids, and pricing per million tokens. All sourced from the same data that powers the model atlas.

Mistral Large 3 vs DeepSeek V4 Pro Max

Mistral Large 32stats won
Tied0stats
DeepSeek V4 Pro Max5stats won
Spec showdown

Capacity, price, freshness.

Mistral Large 3vsDeepSeek V4 Pro Max

DeepSeek V4 Pro Max leads on composite score (74 vs 48), with deeper benchmark coverage across more families. DeepSeek V4 Pro Max has the higher Vals/AA intelligence score (56 vs 16).

  1. Composite score
    Mistral Large 348DeepSeek V4 Pro Max74Winner on this row
  2. Intelligence (Vals/AA)
    Mistral Large 316DeepSeek V4 Pro Max56Winner on this row
  3. Context window
    Mistral Large 3262KWinner on this rowDeepSeek V4 Pro Max256K
  4. Max output
    Mistral Large 38KDeepSeek V4 Pro Max66KWinner on this row
  5. Input price / 1M (lower wins)
    Mistral Large 3$0.50DeepSeek V4 Pro Max$0.40Winner on this row
  6. Output speed (AA default API)
    Mistral Large 365 t/sWinner on this rowDeepSeek V4 Pro Max
  7. Time to first token (AA default API)
    Mistral Large 30.67sDeepSeek V4 Pro MaxWinner on this row
Composite score
48
vs
74Winner on this row
Intelligence (Vals/AA)
16
vs
56Winner on this row
Context window
262KWinner on this row
vs
256K
Max output
8K
vs
66KWinner on this row
Input price / 1M (lower wins)
$0.50
vs
$0.40Winner on this row
Output speed (AA default API)
65 t/sWinner on this row
vs
Time to first token (AA default API)
0.67s
vs
Winner on this row
Capability grid

What each model can do.

Modalities

Mistral Large 3DeepSeek V4 Pro Max
Text
Image
Audio
Video
Code
Tool use
Reasoning

Access

Mistral Large 3DeepSeek V4 Pro Max
API
Consumer app
Open weights
Managed cloud
On-device
Family radar

Top score per benchmark family.

Knowledge
88
vs
90Winner on this row
Reasoning
71
vs
76Winner on this row
Math
38
vs
84Winner on this row
Coding
92
vs
93Winner on this row
Agentic
12
vs
68Winner on this row
Long context
51Winner on this row
vs
49
Tool use
80Winner on this row
vs
75
Benchmark ranks

Every shared benchmark, side by side.

Knowledge

MMLU#13/22 · #8/22
87.8%
vs
89.7%Winner on this row
MMLU-Pro#31/34 · #27/34
72.8%
vs
76.4%Winner on this row

Reasoning

GPQA Diamond#135/151 · #124/151
71.2%
vs
75.6%Winner on this row
LiveBench#11/21 · #14/21
62.8%Winner on this row
vs
61.9%

Coding

HumanEval#12/15 · #9/15
92.4% pass@1
vs
93.4% pass@1Winner on this row
SWE-bench Verified#21/23 · #7/23
76.2%
vs
80.6%Winner on this row

Agentic

Terminal-Bench#141/150 · #35/150
11.99%
vs
67.9%Winner on this row

Long context

LongBench v2#11/19 · #14/19
50.8%Winner on this row
vs
48.6%

Tool use

BFCL#8/20 · #16/20
80.1%Winner on this row
vs
74.6%
MCP Atlas#8/10 · #6/10
70.4%
vs
72.6%Winner on this row
Dossier facts

Dates, sources, fineprint.

Mistral Large 3DeepSeek V4 Pro Max
ProviderMistralDeepSeek
FamilyMistral LargeDeepSeek
Release date2025-12-022026-04-24
Knowledge cutoff2025-06-012026-02-01
Context window262K256K
Max output8K66K
Input price /1M$0.50$0.40
Output price /1M$1.5$1.6
Statussolidflagship
Composite score4874
Intelligence score1656
Intel sourceaavals
Benchmark rows2014
Source count610
Freshnessfreshfresh
Pricing checked2026-06-092026-06-01
Ecosystem

Which tools wrap each model.

Verdict

Which model to actually pick.

DeepSeek V4 Pro Max leads on composite score (74 vs 48), with deeper benchmark coverage across more families. DeepSeek V4 Pro Max has the higher Vals/AA intelligence score (56 vs 16).