88.4%
MMLU
Models / Google
Gemini · Released 2026-05-19
Google's fastest frontier-tier model. Leads the MCP Atlas tool-use benchmark (83.6%) and pairs a 1M-token context with sub-200ms first-token latency. Terminal-Bench 2.0 76.2%, ARC-AGI-2 72.1%. $0.30 / $1.20 per 1M tokens.
Default model for high-throughput, low-latency tool use: search agents, batch classification, real-time customer-facing apps. The cost-per-1M gap to Opus 4.8 is roughly 50× on input and 60× on output.
81average across 8 tested families
Where Gemini 3.5 Flash places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.
The three highest-scoring carded models in each capability family. Where Gemini 3.5 Flash shows up, it's highlighted.
Benchmark rows added to the public ledger for Gemini 3.5 Flash in the last 120 days. Older rows live in the full table below.
A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.
Hover or tab any dot for name, score, and input price.
Every catalog benchmark for Gemini 3.5 Flash. Scores link to the original source; gaps mean no public row exists yet.
21 of 32 catalog benchmarks have a sourced row for Gemini 3.5 Flash.
88.4%
MMLU
72.1%
ARC-AGI-2 semi-private
50.2%
AA Intelligence Index v4.1
74.1%
GPQA Diamond
41%
HLE (AA run)
76.33%
IFBench (AA run)
64.2%
LiveBench
76.2%
Terminal-Bench 2.0 (audited harness)
25.36%
τ³-Bench Banking (AA run)
68.9%
MMMU
69.33%
AA-LCR (AA run)
53.8%
LongBench v2
287 t/s
Median output tokens/s (1k prompt, default provider)
12.11s
Median time to first token (1k prompt, default provider)
1455
LMArena Elo (Style Controlled)