VerdictPal · editorial desk · updated 1 Sep 2026VerdictPal
Compare · Two models, head-to-head

Pick two models. See the diff.

Composite scores, benchmark ranks, capability grids, and pricing per million tokens. All sourced from the same data that powers the model atlas.

Qwen3.8 2.4T A95B vs Qwen3.8 Max

Qwen3.8 2.4T A95B0stats won
Tied3stats
Qwen3.8 Max4stats won
Spec showdown

Capacity, price, freshness.

Qwen3.8 2.4T A95BvsQwen3.8 Max

Composite scores are within 5 points — 64 for Qwen3.8 2.4T A95B vs 66 for Qwen3.8 Max. The choice comes down to capability profile and price. Qwen3.8 Max's 1M context window dwarfs Qwen3.8 2.4T A95B's 262K — a meaningful difference for long-document or codebase work.

  1. Composite score
    Qwen3.8 2.4T A95B64Qwen3.8 Max66Winner on this row
  2. Intelligence (Vals/AA)
    Qwen3.8 2.4T A95B58Qwen3.8 Max58
  3. Context window
    Qwen3.8 2.4T A95B262KQwen3.8 Max1MWinner on this row
  4. Max output
    Qwen3.8 2.4T A95BQwen3.8 Max
  5. Input price / 1M (lower wins)
    Qwen3.8 2.4T A95B$2.0Qwen3.8 Max$2.0
  6. Output speed (AA default API)
    Qwen3.8 2.4T A95B40 t/sQwen3.8 Max40 t/sWinner on this row
  7. Time to first token (AA default API)
    Qwen3.8 2.4T A95B1.76sQwen3.8 Max1.46sWinner on this row
Composite score
64
vs
66Winner on this row
Intelligence (Vals/AA)
58
vs
58
Context window
262K
vs
1MWinner on this row
Max output
vs
Input price / 1M (lower wins)
$2.0
vs
$2.0
Output speed (AA default API)
40 t/s
vs
40 t/sWinner on this row
Time to first token (AA default API)
1.76s
vs
1.46sWinner on this row
Capability grid

What each model can do.

Modalities

Qwen3.8 2.4T A95BQwen3.8 Max
Text
Image
Audio
Video
Code
Tool use
Reasoning

Access

Qwen3.8 2.4T A95BQwen3.8 Max
API
Consumer app
Open weights
Managed cloud
On-device
Family profile

Top score per benchmark family.

Reasoning
94Winner on this row
vs
93
Coding
52
vs
53Winner on this row
Agentic
82Winner on this row
vs
81
Long context
75Winner on this row
vs
74
Benchmark ranks

Every shared benchmark, side by side.

Reasoning

Artificial Analysis Intelligence Index#12/184 · #11/184
57.7%
vs
58.1%Winner on this row
GPQA Diamond#9/185 · #16/185
93.5%Winner on this row
vs
92.7%
Humanity's Last Exam#24/180 · #17/180
42.4%
vs
43%Winner on this row

Coding

SciCode#34/180 · #28/180
51.6%
vs
52.9%Winner on this row

Agentic

Terminal-Bench#16/184 · #20/184
82.02%Winner on this row
vs
81.27%
τ³-Bench Banking#4/106 · #1/106
49.07%
vs
51.34%Winner on this row

Long context

AA-LCR#47/179 · #56/179
75.33%Winner on this row
vs
74.33%

Performance

AA output speed#78/185 · #77/185
40 t/s
vs
40 t/sWinner on this row
AA time to first token#33/185 · #45/185
1.76s
vs
1.46sWinner on this row
Dossier facts

Dates, sources, fineprint.

Qwen3.8 2.4T A95BQwen3.8 Max
ProviderAlibabaAlibaba
FamilyQwenQwen
Release date2026-08-122026-08-03
Knowledge cutoffunpublishedunpublished
Context window262K1M
Max output
Input price /1M$2.0$2.0
Output price /1M$6.0$6.0
Statussolidflagship
Composite score6466
Intelligence score5858
Intel sourceaaaa
Benchmark rows910
Source count11
Freshnessfreshfresh
Pricing checked2026-08-152026-08-12
Ecosystem

Which tools wrap each model.

Qwen3.8 2.4T A95B

Not listed in any tool yet.

Qwen3.8 Max

Not listed in any tool yet.

Verdict

Which model to actually pick.

Composite scores are within 5 points — 64 for Qwen3.8 2.4T A95B vs 66 for Qwen3.8 Max. The choice comes down to capability profile and price. Qwen3.8 Max's 1M context window dwarfs Qwen3.8 2.4T A95B's 262K — a meaningful difference for long-document or codebase work.