VerdictPal · editorial desk · 2026VerdictPal
Compare · Two models, head-to-head

Pick two models. See the diff.

Composite scores, benchmark ranks, capability grids, and pricing per million tokens. All sourced from the same data that powers the model atlas.

DeepSeek V4 Pro Max vs Qwen 3.7 Max

DeepSeek V4 Pro Max4stats won
Tied1stats
Qwen 3.7 Max0stats won
Spec showdown

Capacity, price, freshness.

DeepSeek V4 Pro MaxvsQwen 3.7 Max

Composite scores are within 5 points — 74 for DeepSeek V4 Pro Max vs 73 for Qwen 3.7 Max. The choice comes down to capability profile and price. DeepSeek V4 Pro Max has the higher Vals/AA intelligence score (56 vs 48).

  1. Composite score
    DeepSeek V4 Pro Max74Winner on this rowQwen 3.7 Max73
  2. Intelligence (Vals/AA)
    DeepSeek V4 Pro Max56Winner on this rowQwen 3.7 Max48
  3. Context window
    DeepSeek V4 Pro Max256KWinner on this rowQwen 3.7 Max128K
  4. Max output
    DeepSeek V4 Pro Max66KWinner on this rowQwen 3.7 Max33K
  5. Input price / 1M (lower wins)
    DeepSeek V4 Pro Max$0.40Qwen 3.7 Max$0.40
Composite score
74Winner on this row
vs
73
Intelligence (Vals/AA)
56Winner on this row
vs
48
Context window
256KWinner on this row
vs
128K
Max output
66KWinner on this row
vs
33K
Input price / 1M (lower wins)
$0.40
vs
$0.40
Capability grid

What each model can do.

Modalities

DeepSeek V4 Pro MaxQwen 3.7 Max
Text
Image
Audio
Video
Code
Tool use
Reasoning

Access

DeepSeek V4 Pro MaxQwen 3.7 Max
API
Consumer app
Open weights
Managed cloud
On-device
Family radar

Top score per benchmark family.

Knowledge
90Winner on this row
vs
89
Reasoning
76Winner on this row
vs
74
Math
84Winner on this row
vs
82
Coding
93
vs
93
Agentic
68
vs
70Winner on this row
Long context
49Winner on this row
vs
48
Tool use
75
vs
78Winner on this row
Human preference
84Winner on this row
vs
82
Benchmark ranks

Every shared benchmark, side by side.

Knowledge

MMLU#8/22 · #10/22
89.7%Winner on this row
vs
88.9%
MMLU-Pro#14/21 · #16/21
76.4%Winner on this row
vs
75.1%

Reasoning

GPQA Diamond#56/69 · #60/69
75.6%Winner on this row
vs
73.8%
LiveBench#14/21 · #15/21
61.9%Winner on this row
vs
61.6%

Math

MATH#12/20 · #15/20
83.9%Winner on this row
vs
81.6%

Coding

HumanEval#9/15 · #10/15
93.4% pass@1Winner on this row
vs
93.1% pass@1
SWE-bench Verified#7/23 · #14/23
80.6%Winner on this row
vs
79.4%

Agentic

Terminal-Bench#25/69 · #22/69
67.9%
vs
69.7%Winner on this row
Vals Index#9/20 · #14/20
56.23%Winner on this row
vs
48.04%

Long context

LongBench v2#14/19 · #16/19
48.6%Winner on this row
vs
47.8%

Tool use

BFCL#16/20 · #10/20
74.6%
vs
78.4%Winner on this row
MCP Atlas#6/10 · #7/10
72.6%Winner on this row
vs
71.8%

Human preference

LMArena (Chatbot Arena)#5/13 · #7/13
1418Winner on this row
vs
1409
Dossier facts

Dates, sources, fineprint.

DeepSeek V4 Pro MaxQwen 3.7 Max
ProviderDeepSeekAlibaba
FamilyDeepSeekQwen
Release date2026-04-242026-05-20
Knowledge cutoff2026-02-012026-02-01
Context window256K128K
Max output66K33K
Input price /1M$0.40$0.40
Output price /1M$1.6$1.2
Statusflagshipflagship
Composite score7473
Intelligence score5648
Intel sourcevalsvals
Benchmark rows1414
Source count109
Freshnessfreshfresh
Pricing checked2026-06-012026-06-01
Ecosystem

Which tools wrap each model.

Qwen 3.7 Max

Not listed in any tool yet.

Verdict

Which model to actually pick.

Composite scores are within 5 points — 74 for DeepSeek V4 Pro Max vs 73 for Qwen 3.7 Max. The choice comes down to capability profile and price. DeepSeek V4 Pro Max has the higher Vals/AA intelligence score (56 vs 48).