Google's fastest frontier-tier model. Leads the MCP Atlas tool-use benchmark (83.6%) and pairs a 1M-token context with sub-200ms first-token latency. Terminal-Bench 2.0 76.2%, ARC-AGI-2 72.1%. $0.30 / $1.20 per 1M tokens.
Default model for high-throughput, low-latency tool use: search agents, batch classification, real-time customer-facing apps. The cost-per-1M gap to Opus 4.8 is roughly 50× on input and 60× on output.
#16 of 24 on Vals Index · current snapshot
49 Vals Index · 30 benchmark rows · 11 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Best published score in each covered benchmark family.
79/100 avg
050100
Knowledge
88
Reasoning
76
Math
39
Coding
95
Agentic
76
Multimodal
84
Long context
81
Tool use
84
Human preference
91
9 tested benchmark families
Editor's note
3.5 Flash is the model we recommend for high-volume tool use. The cost-per-1M gap to Opus 4.8 is roughly 50× on input and 60× on output — a meaningful line item for any team running agents at scale. The MCP Atlas 83.6% makes it the strongest tool-use model in the public market.
Benchmark placements
Per-benchmark positions
Where Gemini 3.5 Flash places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.