VerdictPal · editorial desk · 2026VerdictPal
Compare · Two models, head-to-head

Pick two models. See the diff.

Composite scores, benchmark ranks, capability grids, and pricing per million tokens. All sourced from the same data that powers the model atlas.

Gemini 3.5 Flash vs Claude Sonnet 4.6

Gemini 3.5 Flash3stats won
Tied1stats
Claude Sonnet 4.63stats won
Spec showdown

Capacity, price, freshness.

Gemini 3.5 FlashvsClaude Sonnet 4.6

Composite scores are within 5 points — 67 for Gemini 3.5 Flash vs 66 for Claude Sonnet 4.6. The choice comes down to capability profile and price. Claude Sonnet 4.6 has the higher Vals/AA intelligence score (60 vs 49). Gemini 3.5 Flash is significantly cheaper at $0.30/1M input vs $3.0/1M input.

  1. Composite score
    Gemini 3.5 Flash67Winner on this rowClaude Sonnet 4.666
  2. Intelligence (Vals/AA)
    Gemini 3.5 Flash49Claude Sonnet 4.660Winner on this row
  3. Context window
    Gemini 3.5 Flash1MClaude Sonnet 4.61M
  4. Max output
    Gemini 3.5 Flash66KClaude Sonnet 4.6128KWinner on this row
  5. Input price / 1M (lower wins)
    Gemini 3.5 Flash$0.30Winner on this rowClaude Sonnet 4.6$3.0
  6. Output speed (AA default API)
    Gemini 3.5 Flash287 t/sWinner on this rowClaude Sonnet 4.660 t/s
  7. Time to first token (AA default API)
    Gemini 3.5 Flash12.11sClaude Sonnet 4.61.18sWinner on this row
Composite score
67Winner on this row
vs
66
Intelligence (Vals/AA)
49
vs
60Winner on this row
Context window
1M
vs
1M
Max output
66K
vs
128KWinner on this row
Input price / 1M (lower wins)
$0.30Winner on this row
vs
$3.0
Output speed (AA default API)
287 t/sWinner on this row
vs
60 t/s
Time to first token (AA default API)
12.11s
vs
1.18sWinner on this row
Capability grid

What each model can do.

Modalities

Gemini 3.5 FlashClaude Sonnet 4.6
Text
Image
Audio
Video
Code
Tool use
Reasoning

Access

Gemini 3.5 FlashClaude Sonnet 4.6
API
Consumer app
Open weights
Managed cloud
On-device
Family radar

Top score per benchmark family.

Knowledge
88
vs
91Winner on this row
Reasoning
76
vs
77Winner on this row
Coding
95Winner on this row
vs
94
Agentic
76
vs
80Winner on this row
Multimodal
69Winner on this row
vs
66
Long context
69Winner on this row
vs
58
Tool use
84
vs
89Winner on this row
Human preference
91Winner on this row
vs
79
Benchmark ranks

Every shared benchmark, side by side.

Knowledge

MMLU#12/22 · #4/22
88.4%
vs
90.6%Winner on this row

Reasoning

ARC-AGI#7/14 · #8/14
72.1%Winner on this row
vs
58.3%
Artificial Analysis Intelligence Index#16/67 · #43/67
50.2%Winner on this row
vs
35.9%
GPQA Diamond#59/69 · #54/69
74.1%
vs
77.2%Winner on this row
Humanity's Last Exam#11/65 · #54/65
41%Winner on this row
vs
13.2%
IFBench#11/57 · #54/57
76.33%Winner on this row
vs
41.16%
LiveBench#8/21 · #5/21
64.2%
vs
67.2%Winner on this row

Coding

HumanEval#6/15 · #7/15
94.6% pass@1Winner on this row
vs
94.1% pass@1
SciCode#18/65 · #37/65
53.1%Winner on this row
vs
46.9%
SWE-bench Verified#20/23 · #13/23
77.2%
vs
79.6%Winner on this row

Agentic

Terminal-Bench#16/69 · #24/69
76.2%Winner on this row
vs
68.4%
τ³-Bench Banking#35/64 · #12/64
25.36%
vs
79.53%Winner on this row

Multimodal

MMMU#4/8 · #5/8
68.9%Winner on this row
vs
66.2%

Long context

AA-LCR#25/64 · #51/64
69.33%Winner on this row
vs
57.67%
LongBench v2#7/19 · #8/19
53.8%Winner on this row
vs
52.6%

Tool use

BFCL#6/20 · #1/20
82.3%
vs
88.7%Winner on this row
MCP Atlas#1/10 · #4/10
83.6%Winner on this row
vs
75.4%

Human preference

LMArena (Chatbot Arena)#3/13 · #10/13
1455Winner on this row
vs
1396

Performance

AA output speed#6/69 · #48/69
287 t/sWinner on this row
vs
60 t/s
AA time to first token#19/69 · #39/69
12.11s
vs
1.18sWinner on this row
Dossier facts

Dates, sources, fineprint.

Gemini 3.5 FlashClaude Sonnet 4.6
ProviderGoogleAnthropic
FamilyGeminiClaude Sonnet
Release date2026-05-192026-02-17
Knowledge cutoff2026-03-012025-12-01
Context window1M1M
Max output66K128K
Input price /1M$0.30$3.0
Output price /1M$1.2$15
Statusflagshipflagship
Composite score6766
Intelligence score4960
Intel sourcevalsvals
Benchmark rows2724
Source count1111
Freshnessfreshfresh
Pricing checked2026-06-012026-06-01
Ecosystem

Which tools wrap each model.

Verdict

Which model to actually pick.

Composite scores are within 5 points — 67 for Gemini 3.5 Flash vs 66 for Claude Sonnet 4.6. The choice comes down to capability profile and price. Claude Sonnet 4.6 has the higher Vals/AA intelligence score (60 vs 49). Gemini 3.5 Flash is significantly cheaper at $0.30/1M input vs $3.0/1M input.