Models / Google

Gemini 3.5 Flash

Gemini · Released 2026-05-19

Google's fastest frontier-tier model. Leads the MCP Atlas tool-use benchmark (83.6%) and pairs a 1M-token context with sub-200ms first-token latency. Terminal-Bench 2.0 76.2%, ARC-AGI-2 72.1%. $0.30 / $1.20 per 1M tokens.

Default model for high-throughput, low-latency tool use: search agents, batch classification, real-time customer-facing apps. The cost-per-1M gap to Opus 4.8 is roughly 50× on input and 60× on output.

#13 of 21 on Vals Index · current snapshot
49 Vals Index · 27 benchmark rows · 11 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite67
Context1M
Input / 1M$0.30
Output / 1M$1.2
Knowledge cutoff2026-03-01
Output speed287 t/s
TTFT (default API)12.11s
Statusflagship
KnowledgeReasoningCodingAgenticMultimodalLong contextTool useHuman preference
Knowledge
88
Reasoning
76
Coding
95
Agentic
76
Multimodal
69
Long context
69
Tool use
84
Human preference
91

81average across 8 tested families

Editor's note
3.5 Flash is the model we recommend for high-volume tool use. The cost-per-1M gap to Opus 4.8 is roughly 50× on input and 60× on output — a meaningful line item for any team running agents at scale. The MCP Atlas 83.6% makes it the strongest tool-use model in the public market.
Benchmark placements

Where Gemini 3.5 Flash places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
MCP AtlasTool use#1/ 1084MCP Atlas leaderboard2026-05-15
LMArena (Chatbot Arena)Human preference#3/ 1391LMArena leaderboard2026-05-25
MMMUMultimodal#4/ 869MMMU leaderboard2026-05-19
AA output speedPerformance#6/ 69Artificial Analysis2026-07-21
BFCLTool use#6/ 2082Berkeley Function Calling Leaderboard V42026-04-15
DeepSWECoding#6/ 637DeepSWE leaderboard v1.12026-06-20
HumanEvalCoding#6/ 1595CodeSOTA HumanEval leaderboard2026-05-19
ARC-AGIReasoning#7/ 1472ARC Prize leaderboard2026-05-15
LongBench v2Long context#7/ 1954LongBench v2 leaderboard2026-04-01
LiveBenchReasoning#8/ 2164LiveBench leaderboard2026-04-15
Humanity's Last ExamReasoning#11/ 6541Artificial Analysis2026-07-21
IFBenchReasoning#11/ 5776Artificial Analysis2026-07-21
MMLUKnowledge#12/ 2288Google Gemini 3.5 Flash2026-05-19
Artificial Analysis Intelligence IndexReasoning#16/ 6750Artificial Analysis2026-07-21
Terminal-BenchAgentic#16/ 6976Terminal-Bench 2.0 leaderboard2026-05-20
SciCodeCoding#18/ 6553Artificial Analysis2026-07-21
AA time to first tokenPerformance#19/ 69Artificial Analysis2026-07-21
SWE-bench VerifiedCoding#20/ 2377SWE-bench Verified leaderboard2026-06-09
AA-LCRLong context#25/ 6469Artificial Analysis2026-07-21
τ³-Bench BankingAgentic#35/ 6425Artificial Analysis2026-07-21
GPQA DiamondReasoning#59/ 6974Artificial Analysis GPQA Diamond evaluation2026-06-09
Family context

The three highest-scoring carded models in each capability family. Where Gemini 3.5 Flash shows up, it's highlighted.

KnowledgeGemini 3.5 Flash · 88
  1. 1
    91
  2. 2
    91
  3. 3
    91
ReasoningGemini 3.5 Flash · 76
  1. 1
    95
  2. 2
    94
  3. 3
    94
CodingGemini 3.5 Flash · 95
  1. 1
    96
  2. 2
    96
  3. 3
    96
AgenticGemini 3.5 Flash · 76
  1. 1
    100
  2. 2
    100
  3. 3
    100
MultimodalGemini 3.5 Flash · 69
  1. 1
    72
  2. 2
    72
  3. 3
    71
Long contextGemini 3.5 Flash · 69
  1. 1
    75
  2. 2
    75
  3. 3
    74
Tool useGemini 3.5 Flash · 84
  1. 1
    89
  2. 2
    87
  3. 3
    87
Human preferenceGemini 3.5 Flash · 91
  1. 1
    95
  2. 2
    92
  3. 3
    91
Newest receipts

Benchmark rows added to the public ledger for Gemini 3.5 Flash in the last 120 days. Older rows live in the full table below.

26 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Knowledge88
Reasoning76
Coding95
Agentic76
Multimodal69
Long context69
Tool use84
Human preference91
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover or tab any dot for name, score, and input price.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMiniMaxxAIMetaZhipuXiaomiMeituanCohere

Full benchmark ledger

Every catalog benchmark for Gemini 3.5 Flash. Scores link to the original source; gaps mean no public row exists yet.

27 sourced rows

21 of 32 catalog benchmarks have a sourced row for Gemini 3.5 Flash.

Knowledge

Reasoning

Coding

Agentic

Multimodal

Long context

Tool use

Performance

Human preference

Changelog
  1. Released Gemini 3.5 Flash — MCP Atlas leader (83.6%), 1M context, cheapest frontier pricing.
Sources