VerdictPal · editorial desk · 2026VerdictPal
Compare · Two models, head-to-head

Pick two models. See the diff.

Composite scores, benchmark ranks, capability grids, and pricing per million tokens. All sourced from the same data that powers the model atlas.

Grok 4.3 vs Gemini 3.1 Pro

Grok 4.34stats won
Tied0stats
Gemini 3.1 Pro3stats won
Spec showdown

Capacity, price, freshness.

Grok 4.3vsGemini 3.1 Pro

Gemini 3.1 Pro leads on composite score (76 vs 60), with deeper benchmark coverage across more families. Gemini 3.1 Pro has the higher Vals/AA intelligence score (53 vs 47).

  1. Composite score
    Grok 4.360Gemini 3.1 Pro76Winner on this row
  2. Intelligence (Vals/AA)
    Grok 4.347Gemini 3.1 Pro53Winner on this row
  3. Context window
    Grok 4.32MWinner on this rowGemini 3.1 Pro1M
  4. Max output
    Grok 4.3128KWinner on this rowGemini 3.1 Pro66K
  5. Input price / 1M (lower wins)
    Grok 4.3$5.0Winner on this rowGemini 3.1 Pro$7.0
  6. Output speed (AA default API)
    Grok 4.3124 t/sWinner on this rowGemini 3.1 Pro
  7. Time to first token (AA default API)
    Grok 4.331.32sGemini 3.1 ProWinner on this row
Composite score
60
vs
76Winner on this row
Intelligence (Vals/AA)
47
vs
53Winner on this row
Context window
2MWinner on this row
vs
1M
Max output
128KWinner on this row
vs
66K
Input price / 1M (lower wins)
$5.0Winner on this row
vs
$7.0
Output speed (AA default API)
124 t/sWinner on this row
vs
Time to first token (AA default API)
31.32s
vs
Winner on this row
Capability grid

What each model can do.

Modalities

Grok 4.3Gemini 3.1 Pro
Text
Image
Audio
Video
Code
Tool use
Reasoning

Access

Grok 4.3Gemini 3.1 Pro
API
Consumer app
Open weights
Managed cloud
On-device
Family radar

Top score per benchmark family.

Knowledge
89
vs
90Winner on this row
Reasoning
88Winner on this row
vs
78
Coding
78
vs
81Winner on this row
Agentic
64
vs
65Winner on this row
Long context
64Winner on this row
vs
55
Tool use
77
vs
82Winner on this row
Human preference
88Winner on this row
vs
80
Benchmark ranks

Every shared benchmark, side by side.

Knowledge

MMLU#11/22 · #5/22
88.5%
vs
90.3%Winner on this row

Reasoning

GPQA Diamond#31/69 · #52/69
88%Winner on this row
vs
78.4%
LiveBench#17/21 · #7/21
60.8%
vs
65.8%Winner on this row

Coding

SWE-bench Verified#19/23 · #8/23
77.6%
vs
80.6%Winner on this row

Agentic

Terminal-Bench#33/69 · #30/69
63.8%
vs
65.4%Winner on this row
Vals Index#15/20 · #11/20
46.63%
vs
53.42%Winner on this row

Long context

LongBench v2#17/19 · #4/19
46.2%
vs
55.4%Winner on this row

Tool use

BFCL#12/20 · #7/20
77.2%
vs
81.6%Winner on this row

Human preference

LMArena (Chatbot Arena)#4/13 · #9/13
1441Winner on this row
vs
1402
Dossier facts

Dates, sources, fineprint.

Grok 4.3Gemini 3.1 Pro
ProviderxAIGoogle
FamilyGrokGemini
Release date2026-04-302026-02-19
Knowledge cutoff2026-03-152026-01-31
Context window2M1M
Max output128K66K
Input price /1M$5.0$7.0
Output price /1M$25$21
Statussolidflagship
Composite score6076
Intelligence score4753
Intel sourcevalsvals
Benchmark rows1716
Source count99
Freshnessfreshfresh
Pricing checked2026-06-012026-06-01
Ecosystem

Which tools wrap each model.

Grok 4.3

Not listed in any tool yet.

Verdict

Which model to actually pick.

Gemini 3.1 Pro leads on composite score (76 vs 60), with deeper benchmark coverage across more families. Gemini 3.1 Pro has the higher Vals/AA intelligence score (53 vs 47).