VerdictPal · editorial desk · 2026VerdictPal
Compare · Two models, head-to-head

Pick two models. See the diff.

Composite scores, benchmark ranks, capability grids, and pricing per million tokens. All sourced from the same data that powers the model atlas.

Mistral Large 3 vs Llama 4 Maverick

Mistral Large 32stats won
Tied0stats
Llama 4 Maverick5stats won
Spec showdown

Capacity, price, freshness.

Mistral Large 3vsLlama 4 Maverick

Composite scores are within 5 points — 48 for Mistral Large 3 vs 49 for Llama 4 Maverick. The choice comes down to capability profile and price. Mistral Large 3 has the higher Vals/AA intelligence score (16 vs 14). Llama 4 Maverick is more affordable at $0.17/1M input — about 3× cheaper than $0.50/1M. Llama 4 Maverick's 1M context window dwarfs Mistral Large 3's 262K — a meaningful difference for long-document or codebase work.

  1. Composite score
    Mistral Large 348Llama 4 Maverick49Winner on this row
  2. Intelligence (Vals/AA)
    Mistral Large 316Winner on this rowLlama 4 Maverick14
  3. Context window
    Mistral Large 3262KLlama 4 Maverick1MWinner on this row
  4. Max output
    Mistral Large 38KLlama 4 Maverick16KWinner on this row
  5. Input price / 1M (lower wins)
    Mistral Large 3$0.50Llama 4 Maverick$0.17Winner on this row
  6. Output speed (AA default API)
    Mistral Large 364 t/sLlama 4 Maverick110 t/sWinner on this row
  7. Time to first token (AA default API)
    Mistral Large 30.59sWinner on this rowLlama 4 Maverick0.60s
Composite score
48
vs
49Winner on this row
Intelligence (Vals/AA)
16Winner on this row
vs
14
Context window
262K
vs
1MWinner on this row
Max output
8K
vs
16KWinner on this row
Input price / 1M (lower wins)
$0.50
vs
$0.17Winner on this row
Output speed (AA default API)
64 t/s
vs
110 t/sWinner on this row
Time to first token (AA default API)
0.59sWinner on this row
vs
0.60s
Capability grid

What each model can do.

Modalities

Mistral Large 3Llama 4 Maverick
Text
Image
Audio
Video
Code
Tool use
Reasoning

Access

Mistral Large 3Llama 4 Maverick
API
Consumer app
Open weights
Managed cloud
On-device
Family radar

Top score per benchmark family.

Knowledge
88Winner on this row
vs
86
Reasoning
71Winner on this row
vs
69
Math
38
vs
89Winner on this row
Coding
92
vs
92
Agentic
12Winner on this row
vs
8
Long context
51Winner on this row
vs
50
Tool use
80Winner on this row
vs
76
Benchmark ranks

Every shared benchmark, side by side.

Knowledge

MMLU#13/22 · #19/22
87.8%Winner on this row
vs
86.1%
MMLU-Pro#18/21 · #11/21
72.8%
vs
80.9%Winner on this row

Reasoning

Artificial Analysis Intelligence Index#58/67 · #59/67
15.9%Winner on this row
vs
14.3%
GPQA Diamond#62/69 · #63/69
71.2%Winner on this row
vs
69.4%
Humanity's Last Exam#64/65 · #61/65
4.1%
vs
4.8%Winner on this row
IFBench#57/57 · #53/57
36.19%
vs
42.99%Winner on this row
LiveBench#11/21 · #13/21
62.8%Winner on this row
vs
62.1%

Math

AIME#15/19 · #16/19
38%Winner on this row
vs
19.33%

Coding

HumanEval#12/15 · #14/15
92.4% pass@1Winner on this row
vs
91.8% pass@1
LiveCodeBench#9/13 · #10/13
46.5%Winner on this row
vs
39.7%
SciCode#55/65 · #58/65
36.2%Winner on this row
vs
33.1%
SWE-bench Verified#21/23 · #22/23
76.2%Winner on this row
vs
74.8%

Agentic

Terminal-Bench#64/69 · #66/69
11.99%Winner on this row
vs
7.87%
τ³-Bench Banking#61/64 · #63/64
5.77%Winner on this row
vs
3.92%

Long context

AA-LCR#57/64 · #56/64
34.67%
vs
46%Winner on this row
LongBench v2#11/19 · #12/19
50.8%Winner on this row
vs
49.8%

Tool use

BFCL#8/20 · #15/20
80.1%Winner on this row
vs
75.9%

Performance

AA output speed#44/69 · #24/69
64 t/s
vs
110 t/sWinner on this row
AA time to first token#55/69 · #54/69
0.59sWinner on this row
vs
0.60s
Dossier facts

Dates, sources, fineprint.

Mistral Large 3Llama 4 Maverick
ProviderMistralMeta
FamilyMistral LargeLlama 4
Release date2025-12-022026-04-05
Knowledge cutoff2025-06-012024-08-01
Context window262K1M
Max output8K16K
Input price /1M$0.50$0.17
Output price /1M$1.5$0.60
Statussolidflagship
Composite score4849
Intelligence score1614
Intel sourceaaaa
Benchmark rows2021
Source count66
Freshnessfreshfresh
Pricing checked2026-06-092026-06-09
Ecosystem

Which tools wrap each model.

Llama 4 Maverick

Not listed in any tool yet.

Verdict

Which model to actually pick.

Composite scores are within 5 points — 48 for Mistral Large 3 vs 49 for Llama 4 Maverick. The choice comes down to capability profile and price. Mistral Large 3 has the higher Vals/AA intelligence score (16 vs 14). Llama 4 Maverick is more affordable at $0.17/1M input — about 3× cheaper than $0.50/1M. Llama 4 Maverick's 1M context window dwarfs Mistral Large 3's 262K — a meaningful difference for long-document or codebase work.