Models / xAI

Grok 4.3

Grok · Released 2026-04-30

xAI's April 2026 frontier model with a 2M-token context window, real-time X search built in, and Vals Index 46.63%. Not the smartest model on the board, but the only one with a 2M context and a live social graph.

Default model for real-time social signals, very long context (>1M), and freshness-sensitive research where minutes matter.

#19 of 24 on Vals Index · current snapshot
47 Vals Index · 17 benchmark rows · 9 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite60
Context2M
Input / 1M$5.0
Output / 1M$25
Knowledge cutoff2026-03-15
Statussolid

Family profile

Best published score in each covered benchmark family.

79/100 avg
Knowledge
89
Reasoning
88
Coding
78
Agentic
64
Long context
66
Tool use
77
Human preference
88

7 tested benchmark families

Editor's note
Grok 4.3 is the only model on the public market with a real >1M context. The 2M window is real — but in our long-document tests, citation fidelity past 1.5M tokens drops sharply. The X search integration is a genuine differentiator for breaking-news research.
Benchmark placements

Where Grok 4.3 places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
LMArena (Chatbot Arena)Human preference#4/ 1388LMArena leaderboard2026-05-25
IFBenchReasoning#5/ 12781Artificial Analysis2026-09-01
MMLUKnowledge#11/ 2289xAI Grok 4.32026-04-30
BFCLTool use#12/ 2077Berkeley Function Calling Leaderboard V42026-04-15
LiveBenchReasoning#17/ 2161LiveBench leaderboard2026-04-15
LongBench v2Long context#17/ 1946LongBench v2 leaderboard2026-04-01
Vals IndexAgentic#18/ 2347Vals AI — Vals Index2026-06-04
SWE-bench VerifiedCoding#19/ 2378SWE-bench Verified leaderboard2026-06-09
Humanity's Last ExamReasoning#46/ 18037Artificial Analysis2026-09-01
Terminal-BenchAgentic#60/ 18464Terminal-Bench 2.0 leaderboard2026-06-09
SciCodeCoding#61/ 18047Artificial Analysis2026-09-01
GPQA DiamondReasoning#63/ 18588GPQA Diamond paper leaderboard2026-05-15
τ³-Bench BankingAgentic#82/ 10612Artificial Analysis2026-09-01
Artificial Analysis Intelligence IndexReasoning#88/ 18738Artificial Analysis2026-09-01
AA-LCRLong context#119/ 17966Artificial Analysis2026-09-01
AA time to first tokenPerformance#149/ 186Artificial Analysis2026-09-01
AA output speedPerformance#150/ 187Artificial Analysis2026-09-01
Family context

The three highest-scoring models with pages in each capability family. Where Grok 4.3 shows up, it's highlighted.

Newest receipts

Benchmark rows added to the public ledger for Grok 4.3 in the last 120 days. Older rows live in the full table below.

13 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Knowledge89
Reasoning88
Coding78
Agentic64
Long context66
Tool use77
Human preference88
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover any dot for name, score, and input price. Keyboard: tab through the top twelve, or use the ranking below.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMetaMiniMaxxAIZhipuXiaomiOtherMeituanCohereNVIDIA

Full benchmark ledger

Every catalog benchmark for Grok 4.3. Scores link to the original source; gaps mean no public row exists yet.

17 sourced rows

17 of 37 catalog benchmarks have a sourced row for Grok 4.3.

Knowledge

Reasoning

Coding

Agentic

Long context

Tool use

Performance

Human preference

Changelog
  1. Released Grok 4.3 with 2M context and real-time X search.
  2. Grok 4.20 retired from general availability.
Sources