Modelle / Google

Gemini 3.1 Pro

Gemini · Release 2026-02-19

Googles Frontier-Modell Februar 2026. AA Intelligence Index 57,2 (#4). SWE-bench Verified 80,6 %, Terminal-Bench 2.0 78,4 %, ARC-AGI-2 77,1 % (Googles Top-Modell ohne Flash). 1M-Kontext, $7 / $21 pro 1M Tokens.

Googles stärkstes Frontier-Modell ohne Flash. Das Modell der Wahl, wenn 1M-Kontext, multimodaler Input und Tool-Use im selben Workflow zählen.

#11 von 21 im Vals Index · aktueller Snapshot
53 Vals Index · 16 Benchmark-Zeilen · 9 QuellenPosition ist benchmark-spezifisch — kein Cross-Family- oder Cross-Source-Ranking.
Composite76
Kontext1M
Input / 1M$7.0
Output / 1M$21
Wissensstichtag2026-01-31
Statusflagship
KnowledgeReasoningMathCodingAgenticMultimodalLong contextTool useSafetyHuman preference
Knowledge
90
Reasoning
78
Math
88
Coding
81
Agentic
65
Multimodal
72
Long context
55
Tool use
82
Safety
89
Human preference
80

78average across 10 tested families

Redaktionsnotiz
Gemini 3.1 Pro ist die Google-Wahl, wenn 3.5 Flash zu günstig ist, um ernst genommen zu werden. Der 1M-Kontext und der native multimodale Input sind die echten Differentiatoren — ARC-AGI-2 77,1 % bringt es in die globale Top 5 beim härtesten Reasoning-Test.
Benchmark-Platzierungen

Wo Gemini 3.1 Pro auf jeder öffentlichen Benchmark-Quelle mit Zeile steht. Rang zählt jedes Modell mit neuester Zeile im selben Test — kein universeller Qualitätsscore.

BenchmarkFamilyRangScoreQuelleDatum
MMMUMultimodal#1/ 872MMMU leaderboard2026-02-19
ARC-AGIReasoning#4/ 1477ARC Prize leaderboard2026-05-15
LongBench v2Long context#4/ 1955LongBench v2 leaderboard2026-04-01
MCP AtlasTool use#5/ 1074MCP Atlas leaderboard2026-06-09
MMLUKnowledge#5/ 2290Google Gemini 3.1 Pro2026-02-19
HELM SafetySafety#6/ 989Stanford HELM leaderboard2026-02-19
BFCLTool use#7/ 2082Berkeley Function Calling Leaderboard V42026-04-15
LiveBenchReasoning#7/ 2166LiveBench leaderboard2026-04-15
MATHMath#8/ 2087Google Gemini 3.1 Pro2026-02-19
SWE-bench VerifiedCoding#8/ 2381SWE-bench official leaderboard2026-05-30
AIMEMath#9/ 1988Google Gemini 3.1 Pro2026-02-19
LMArena (Chatbot Arena)Human preference#9/ 1380LMArena leaderboard2026-05-25
Vals IndexAgentic#11/ 2053Vals AI — Vals Index2026-06-04
MMLU-ProKnowledge#12/ 2180MMLU-Pro HuggingFace leaderboard2026-02-19
Terminal-BenchAgentic#30/ 6965Terminal-Bench 2.0 leaderboard2026-05-20
GPQA DiamondReasoning#52/ 6978Artificial Analysis GPQA Diamond evaluation2026-06-09
Family-Kontext

Die drei höchstscorierenden kartierten Modelle pro Capability-Family. Wo Gemini 3.1 Pro auftaucht, ist es hervorgehoben.

KnowledgeGemini 3.1 Pro · 90
  1. 1
    91
  2. 2
    91
  3. 3
    91
ReasoningGemini 3.1 Pro · 78
  1. 1
    95
  2. 2
    94
  3. 3
    94
MathGemini 3.1 Pro · 88
  1. 1
    99
  2. 2
    97
  3. 3
    96
CodingGemini 3.1 Pro · 81
  1. 1
    96
  2. 2
    96
  3. 3
    96
AgenticGemini 3.1 Pro · 65
  1. 1
    100
  2. 2
    100
  3. 3
    100
MultimodalGemini 3.1 Pro · 72
  1. 1
    72
  2. 2
    72
  3. 3
    71
Long contextGemini 3.1 Pro · 55
  1. 1
    75
  2. 2
    75
  3. 3
    74
Tool useGemini 3.1 Pro · 82
  1. 1
    89
  2. 2
    87
  3. 3
    87
SafetyGemini 3.1 Pro · 89
  1. 1
    92
  2. 2
    92
  3. 3
    91
Human preferenceGemini 3.1 Pro · 80
  1. 1
    95
  2. 2
    92
  3. 3
    91
Neueste Belege

Benchmark-Zeilen im öffentlichen Ledger für Gemini 3.1 Pro in den letzten 120 Tagen. Ältere Zeilen stehen in der vollen Tabelle unten.

10 neu
Family-Abdeckung

Ein hoher Composite, der eine schwache Family versteckt, ist eine Falle. Diese Balken zeigen Families ohne öffentlichen Test — und wo das Modell führt.

Knowledge90
Reasoning78
Math88
Coding81
Agentic65
Multimodal72
Long context55
Tool use82
Safety89
Human preference80
Preis vs. Performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover or tab any dot for name, score, and input price.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMiniMaxxAIMetaZhipuXiaomiMeituanCohere

Volles Benchmark-Ledger

Jeder Katalog-Benchmark für Gemini 3.1 Pro. Scores verlinken zur Originalquelle; Lücken heißen: noch keine öffentliche Zeile.

17 Quell-Zeilen

16 of 32 catalog benchmarks have a sourced row for Gemini 3.1 Pro.

Knowledge

Reasoning

Math

Coding

Agentic

Multimodal

Long context

Tool use

Safety

Human preference

Changelog
  1. Gemini 3.1 Pro veröffentlicht — AA-Index 57,2, ARC-AGI-2 77,1 %.
Quellen