Models / Google

Gemini 3.1 Pro

Gemini · Released 2026-02-19

Google's February 2026 frontier model. AA Intelligence Index 57.2 (#4). SWE-bench Verified 80.6%, Terminal-Bench 2.0 78.4%, ARC-AGI-2 77.1% (Google's top non-flash model). 1M context, $7 / $21 per 1M tokens.

Google's strongest non-flash frontier model. The model to choose when 1M context, multimodal input, and tool-use all matter in the same workflow.

#11 of 21 on Vals Index · current snapshot
53 Vals Index · 16 benchmark rows · 9 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite76
Context1M
Input / 1M$7.0
Output / 1M$21
Knowledge cutoff2026-01-31
Statusflagship
KnowledgeReasoningMathCodingAgenticMultimodalLong contextTool useSafetyHuman preference
Knowledge
90
Reasoning
78
Math
88
Coding
81
Agentic
65
Multimodal
72
Long context
55
Tool use
82
Safety
89
Human preference
80

78average across 10 tested families

Editor's note
Gemini 3.1 Pro is the Google pick when 3.5 Flash is too cheap to be serious. The 1M context and native multimodal input are the real differentiators — ARC-AGI-2 77.1% puts it in the top 5 globally on the hardest reasoning test.
Benchmark placements

Where Gemini 3.1 Pro places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
MMMUMultimodal#1/ 872MMMU leaderboard2026-02-19
ARC-AGIReasoning#4/ 1477ARC Prize leaderboard2026-05-15
LongBench v2Long context#4/ 1955LongBench v2 leaderboard2026-04-01
MCP AtlasTool use#5/ 1074MCP Atlas leaderboard2026-06-09
MMLUKnowledge#5/ 2290Google Gemini 3.1 Pro2026-02-19
HELM SafetySafety#6/ 989Stanford HELM leaderboard2026-02-19
BFCLTool use#7/ 2082Berkeley Function Calling Leaderboard V42026-04-15
LiveBenchReasoning#7/ 2166LiveBench leaderboard2026-04-15
MATHMath#8/ 2087Google Gemini 3.1 Pro2026-02-19
SWE-bench VerifiedCoding#8/ 2381SWE-bench official leaderboard2026-05-30
AIMEMath#9/ 1988Google Gemini 3.1 Pro2026-02-19
LMArena (Chatbot Arena)Human preference#9/ 1380LMArena leaderboard2026-05-25
Vals IndexAgentic#11/ 2053Vals AI — Vals Index2026-06-04
MMLU-ProKnowledge#12/ 2180MMLU-Pro HuggingFace leaderboard2026-02-19
Terminal-BenchAgentic#30/ 6965Terminal-Bench 2.0 leaderboard2026-05-20
GPQA DiamondReasoning#52/ 6978Artificial Analysis GPQA Diamond evaluation2026-06-09
Family context

The three highest-scoring carded models in each capability family. Where Gemini 3.1 Pro shows up, it's highlighted.

KnowledgeGemini 3.1 Pro · 90
  1. 1
    91
  2. 2
    91
  3. 3
    91
ReasoningGemini 3.1 Pro · 78
  1. 1
    95
  2. 2
    94
  3. 3
    94
MathGemini 3.1 Pro · 88
  1. 1
    99
  2. 2
    97
  3. 3
    96
CodingGemini 3.1 Pro · 81
  1. 1
    96
  2. 2
    96
  3. 3
    96
AgenticGemini 3.1 Pro · 65
  1. 1
    100
  2. 2
    100
  3. 3
    100
MultimodalGemini 3.1 Pro · 72
  1. 1
    72
  2. 2
    72
  3. 3
    71
Long contextGemini 3.1 Pro · 55
  1. 1
    75
  2. 2
    75
  3. 3
    74
Tool useGemini 3.1 Pro · 82
  1. 1
    89
  2. 2
    87
  3. 3
    87
SafetyGemini 3.1 Pro · 89
  1. 1
    92
  2. 2
    92
  3. 3
    91
Human preferenceGemini 3.1 Pro · 80
  1. 1
    95
  2. 2
    92
  3. 3
    91
Newest receipts

Benchmark rows added to the public ledger for Gemini 3.1 Pro in the last 120 days. Older rows live in the full table below.

10 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Knowledge90
Reasoning78
Math88
Coding81
Agentic65
Multimodal72
Long context55
Tool use82
Safety89
Human preference80
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover or tab any dot for name, score, and input price.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMiniMaxxAIMetaZhipuXiaomiMeituanCohere

Full benchmark ledger

Every catalog benchmark for Gemini 3.1 Pro. Scores link to the original source; gaps mean no public row exists yet.

17 sourced rows

16 of 32 catalog benchmarks have a sourced row for Gemini 3.1 Pro.

Knowledge

Reasoning

Math

Coding

Agentic

Multimodal

Long context

Tool use

Safety

Human preference

Changelog
  1. Released Gemini 3.1 Pro — AA Index 57.2, ARC-AGI-2 77.1%.
Sources