Models / xAI

Grok 4.20 Reasoning

Grok 4 · Released 2026-03-05

xAI's March 2026 reasoning Grok. Vals Index 39.11 (#15). Reasoning-first variant of the Grok 4 family.

xAI reasoning route when Grok 4.3 is too expensive or too general.

#23 of 24 on Vals Index · current snapshot
39 Vals Index · 12 benchmark rows · 4 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite54
Context256K
Input / 1M
Output / 1M
Knowledge cutoff2026-01-01
Statussolid

Family profile

Best published score in each covered benchmark family.

60/100 avg
Reasoning
91
Coding
46
Agentic
39
Long context
62

4 tested benchmark families

Benchmark placements

Where Grok 4.20 Reasoning places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
IFBenchReasoning#6/ 12781Artificial Analysis2026-09-01
GPQA DiamondReasoning#32/ 18591Artificial Analysis2026-09-01
Humanity's Last ExamReasoning#55/ 18035Artificial Analysis2026-09-01
SciCodeCoding#77/ 18046Artificial Analysis2026-09-01
Artificial Analysis Intelligence IndexReasoning#86/ 18738Artificial Analysis2026-09-01
Terminal-BenchAgentic#123/ 18438Artificial Analysis2026-09-01
AA-LCRLong context#133/ 17962Artificial Analysis2026-09-01
AA time to first tokenPerformance#148/ 186Artificial Analysis2026-09-01
AA output speedPerformance#149/ 187Artificial Analysis2026-09-01
Family context

The three highest-scoring models with pages in each capability family. Where Grok 4.20 Reasoning shows up, it's highlighted.

Newest receipts

Benchmark rows added to the public ledger for Grok 4.20 Reasoning in the last 120 days. Older rows live in the full table below.

11 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Reasoning91
Coding46
Agentic39
Long context62
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover any dot for name, score, and input price. Keyboard: tab through the top twelve, or use the ranking below.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMetaMiniMaxxAIZhipuXiaomiOtherMeituanCohereNVIDIA

Full benchmark ledger

Every catalog benchmark for Grok 4.20 Reasoning. Scores link to the original source; gaps mean no public row exists yet.

13 sourced rows

12 of 37 catalog benchmarks have a sourced row for Grok 4.20 Reasoning.

Reasoning

Coding

Agentic

Long context

Performance

Sources