Models / xAI

Grok 4.6

Grok · Released 2026-08-12

xAI's August 2026 flagship. Artificial Analysis Intelligence Index 60.9 on the high-effort endpoint — tied with GPT-5.6 Sol (max) on the same snapshot. API list price $2 / $6 per 1M tokens.

Current xAI frontier model for agentic coding and knowledge-work runs. Prefer sourced ledger rows over launch claims; AA scores sync from the public snapshot.

#3 of 166 on AA Index · current snapshot
61 AA Index · 12 benchmark rows · 2 sourcesPosition is benchmark-specific — not a cross-family or cross-source ranking.
Composite54
ContextUnlisted
Input / 1M$2.0
Output / 1M$6.0
Knowledge cutoffopen
Output speed51 t/s
TTFT (default API)35.04s
Statusflagship

Family profile

Best published score in each covered benchmark family.

72/100 avg
Knowledge
48
Reasoning
95
Coding
54
Agentic
88
Long context
75

5 tested benchmark families

Editor's note
Recorded from the 2026-08-12 AA snapshot (slug grok-4-6, high effort, index 60.9) plus the xAI launch post. Editorial depth and failure modes still need a desk pass — scores below are ledger-backed, not hands-on claims. Context window is unlisted on AA.
Benchmark placements

Where Grok 4.6 places on each public benchmark source that publishes a row. Rank counts every model with a latest row in the same test — not a universal quality score.

BenchmarkFamilyRankScoreSourceDate
GPQA DiamondReasoning#1/ 18595Artificial Analysis2026-09-01
τ³-Bench BankingAgentic#2/ 10651Artificial Analysis2026-09-01
Terminal-BenchAgentic#2/ 18488Artificial Analysis2026-09-01
Artificial Analysis Intelligence IndexReasoning#7/ 18461Artificial Analysis2026-09-01
AA time to first tokenPerformance#8/ 185Artificial Analysis2026-09-01
Terminal-Bench-ScienceAgentic#8/ 97Terminal-Bench-Science 0.1 announcement2026-08-27
AA-OmniscienceKnowledge#10/ 1248Artificial Analysis — AA-Omniscience2026-09-01
CritPtReasoning#12/ 1217Artificial Analysis — CritPt2026-09-01
Humanity's Last ExamReasoning#19/ 18043Artificial Analysis2026-09-01
SciCodeCoding#21/ 18054Artificial Analysis2026-09-01
AA-LCRLong context#48/ 17975Artificial Analysis2026-09-01
AA output speedPerformance#62/ 185Artificial Analysis2026-09-01
Family context

The three highest-scoring models with pages in each capability family. Where Grok 4.6 shows up, it's highlighted.

Newest receipts

Benchmark rows added to the public ledger for Grok 4.6 in the last 120 days. Older rows live in the full table below.

12 recent
Family coverage

A high composite that hides a weak family is a trap. These bars surface the families where this model hasn't been publicly tested, and where it leads.

Knowledge48
Reasoning95
Coding54
Agentic88
Long context75
Price vs. performance
0255075100$0.10$0.30$1$3$10$30input USD / 1M tokens · log scale →← intelligence

Hover or tab any dot for name, score, and input price.

AnthropicOpenAIMistralAlibabaDeepSeekGoogleMoonshotMetaMiniMaxxAIZhipuXiaomiOtherMeituanCohereNVIDIA

Full benchmark ledger

Every catalog benchmark for Grok 4.6. Scores link to the original source; gaps mean no public row exists yet.

12 sourced rows

12 of 37 catalog benchmarks have a sourced row for Grok 4.6.

Knowledge

Reasoning

Coding

Agentic

Long context

Performance

Changelog
  1. Grok 4.6 appears on Artificial Analysis (high-effort index 60.9, $2 / $6 per 1M tokens).
Sources