Models / xAI

Grok 4.5

Grok · Released 2026-07-08

xAI's July 2026 frontier release. Artificial Analysis Intelligence Index 53.8 (high-effort setting) — above Grok 4.3 and in the current top tier with Claude Opus 4.8 and GPT-5.5. API list price $2 / $6 per 1M tokens.

Frontier xAI model for long-context and freshness-sensitive work. Prefer sourced ledger rows over marketing claims; AA scores sync from the public snapshot.

#11 of 168 on AA Index · current snapshot
56 AA Index · 10 benchmark rows · 1 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite62
Context2M
Input / 1M$2.0
Output / 1M$6.0
Knowledge cutoffopen
Output speed48 t/s
TTFT (default API)11.48s
Statussolid

Family profile

Best published score in each covered benchmark family.

71/100 avg
Knowledge
52
Reasoning
93
Coding
54
Agentic
82
Long context
74

5 tested benchmark families

Editor's note
Tool scaffolded from the 2026-07-09 AA snapshot (slug grok-4-5, high effort). Editorial depth and failure modes still need a desk pass — scores below are ledger-backed, not hands-on claims.
Benchmark placements

Where Grok 4.5 places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
AA-OmniscienceKnowledge#7/ 1252Artificial Analysis — AA-Omniscience2026-09-01
GPQA DiamondReasoning#12/ 18593Artificial Analysis2026-09-01
τ³-Bench BankingAgentic#13/ 10642Artificial Analysis2026-09-01
AA time to first tokenPerformance#14/ 186Artificial Analysis2026-09-01
Artificial Analysis Intelligence IndexReasoning#15/ 18756Artificial Analysis2026-09-01
SciCodeCoding#18/ 18054Artificial Analysis2026-09-01
Terminal-BenchAgentic#19/ 18482Artificial Analysis2026-09-01
Humanity's Last ExamReasoning#21/ 18043Artificial Analysis2026-09-01
AA-LCRLong context#58/ 17974Artificial Analysis2026-09-01
AA output speedPerformance#69/ 187Artificial Analysis2026-09-01
Family context

The three highest-scoring models with pages in each capability family. Where Grok 4.5 shows up, it's highlighted.

Newest receipts

Benchmark rows added to the public ledger for Grok 4.5 in the last 120 days. Older rows live in the full table below.

10 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Knowledge52
Reasoning93
Coding54
Agentic82
Long context74
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover any dot for name, score, and input price. Keyboard: tab through the top twelve, or use the ranking below.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMetaMiniMaxxAIZhipuXiaomiOtherMeituanCohereNVIDIA

Full benchmark ledger

Every catalog benchmark for Grok 4.5. Scores link to the original source; gaps mean no public row exists yet.

10 sourced rows

10 of 37 catalog benchmarks have a sourced row for Grok 4.5.

Knowledge

Reasoning

Coding

Agentic

Long context

Performance

Changelog
  1. Grok 4.5 appears on Artificial Analysis (high-effort index 53.8).
Sources