Models / Google

Gemma 4 31B IT

Gemma · Released 2026-03-01

Google's March 2026 open-weights Gemma 4 instruct model. SAGE 55% on Vals (#2). Small enough to self-host on a single GPU cluster.

Self-hosted Google open weights for privacy-sensitive workflows that can't call cloud APIs.

#129 of 168 on AA Index · current snapshot
30 AA Index · 10 benchmark rows · 1 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite48
Context131K
Input / 1M
Output / 1M
Knowledge cutoffopen
Output speed36 t/s
TTFT (default API)0.97s
Statussolid

Family profile

Best published score in each covered benchmark family.

60/100 avg
Reasoning
86
Coding
43
Agentic
43
Long context
68

4 tested benchmark families

Benchmark placements

Where Gemma 4 31B IT places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
IFBenchReasoning#25/ 12776Artificial Analysis2026-09-01
AA time to first tokenPerformance#65/ 186Artificial Analysis2026-09-01
τ³-Bench BankingAgentic#72/ 10615Artificial Analysis2026-09-01
AA output speedPerformance#85/ 187Artificial Analysis2026-09-01
GPQA DiamondReasoning#89/ 18586Artificial Analysis2026-09-01
SciCodeCoding#93/ 18043Artificial Analysis2026-09-01
AA-LCRLong context#101/ 17968Artificial Analysis2026-09-01
Terminal-BenchAgentic#102/ 18443Artificial Analysis2026-09-01
Humanity's Last ExamReasoning#107/ 18024Artificial Analysis2026-09-01
Artificial Analysis Intelligence IndexReasoning#147/ 18730Artificial Analysis2026-09-01
Family context

The three highest-scoring models with pages in each capability family. Where Gemma 4 31B IT shows up, it's highlighted.

Newest receipts

Benchmark rows added to the public ledger for Gemma 4 31B IT in the last 120 days. Older rows live in the full table below.

10 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Reasoning86
Coding43
Agentic43
Long context68
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover any dot for name, score, and input price. Keyboard: tab through the top twelve, or use the ranking below.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMetaMiniMaxxAIZhipuXiaomiOtherMeituanCohereNVIDIA

Full benchmark ledger

Every catalog benchmark for Gemma 4 31B IT. Scores link to the original source; gaps mean no public row exists yet.

10 sourced rows

10 of 37 catalog benchmarks have a sourced row for Gemma 4 31B IT.

Reasoning

Coding

Agentic

Long context

Performance

Sources