Benchmarks / Reasoning

Artificial Analysis Intelligence Index

AA Intelligence Index

AA Intelligence Index v4.1 — weighted composite emphasizing agentic workloads: GDPval-AA v2 (20%), Terminal-Bench 2.1 (16%), τ³-Bench Banking (14%), Humanity's Last Exam (12%), AA-Omniscience (12%), SciCode (8%), GPQA Diamond (6%), AA-LCR (6%), CritPt (6%). IFBench removed for saturation.

Artificial AnalysisReasoningFlagshipActiveMedium contamination riskSince 2026
What this does not measure
  • Latency, cost-per-token, and throughput — the index is intelligence-only. Use the AA API alongside for $/1M tokens.
  • Multilingual or multimodal ability — AA benchmarks image, speech, and multilingual performance on separate indices.
  • Any single sub-test in isolation — a high composite can hide weak agentic or coding slices.
Analysis

Why this benchmark is useful

Single-number model rankings hide tradeoffs. The AA Intelligence Index is a weighted composite of ten public sub-evals — useful as a dated frontier snapshot when you read the slices, not as a universal quality score.

Scope

Coverage map

Task family
Reasoning
Format
Composite of ten versioned sub-evaluations with public weights; methodology page lists per-eval question counts and scoring.
Scoring
Weighted percentage across four equal categories. AA estimates ±1% CI on the composite for models with sufficient repeats.
Maintainer
Artificial Analysis
Reading guide

How to read the scores

Open the methodology page for v4.1 weights, then drill into each sub-eval card. A high composite can mask weak agentic, coding, or long-context slices.

Blind spots

What it does not cover

  • Latency, cost-per-token, and throughput — the index is intelligence-only. Use the AA API alongside for $/1M tokens.
  • Multilingual or multimodal ability — AA benchmarks image, speech, and multilingual performance on separate indices.
  • Any single sub-test in isolation — a high composite can hide weak agentic or coding slices.
Scores

Evidence ledger

67 rows
6767 rows
59.9%best score
67source-checked
1sources
2026-07-21source date
67

api · api

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Claude Fable 559.9% ·
  2. GPT-5.6 Sol58.9% ·
  3. Kimi K357.1% ·
  4. Claude Opus 4.855.7% ·
  5. GPT-5.6 Terra55% ·

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Claude Fable 5Anthropic2026-06-0959.9%
2026-07-21
Source-checked
2GPT-5.6 SolOpenAI2026-07-0958.9%
2026-07-21
Source-checked
3Kimi K3Moonshot2026-07-1657.1%
2026-07-21
Source-checked
4Claude Opus 4.8Anthropic2026-05-2855.7%
2026-07-21
Source-checked
5GPT-5.6 TerraOpenAI2026-07-0955%
2026-07-21
Source-checked
6GPT 5.5OpenAI2026-04-2354.8%
2026-07-21
Source-checked
7Grok 4.5xAI2026-07-0853.8%
2026-07-21
Source-checked
8Claude Opus 4.7Anthropic2026-04-1653.5%
2026-07-21
Source-checked
9Claude Sonnet 5Anthropic2026-06-3053.4%
2026-07-21
Source-checked
10GPT-5.5 (high)OpenAI2026-04-2353.1%
2026-07-21
Source-checked
11GPT 5.4OpenAI2026-03-0551.4%
2026-07-21
Source-checked
12GPT-5.6 LunaOpenAI2026-07-0951.2%
2026-07-21
Source-checked
13GLM-5.2Zhipu2026-06-1351.1%
2026-07-21
Source-checked
14Muse Spark 1.1Meta2026-07-0950.6%
2026-07-21
Source-checked
15GPT-5.5 (medium)OpenAI2026-04-2350.4%
2026-07-21
Source-checked
16Gemini 3.5 FlashGoogle2026-05-1950.2%
2026-07-21
Source-checked
17Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)Anthropic2026-02-1747.2%
2026-07-21
Source-checked
18Gemini 3.1 Pro PreviewGoogle2026-02-1946.5%
2026-07-21
Source-checked
19Qwen3.7 MaxAlibaba2026-05-1946%
2026-07-21
Source-checked
20Gemini 3.5 Flash (medium)Google2026-05-1945.4%
2026-07-21
Source-checked
21MiniMax M3MiniMax2026-05-3144.4%
2026-07-21
Source-checked
22DeepSeek V4 Pro (Reasoning, Max Effort)DeepSeek2026-04-2444.3%
2026-07-21
Source-checked
23GPT-5.3 Codex (xhigh)OpenAI2026-02-0544.3%
2026-07-21
Source-checked
24Kimi K2.6Moonshot2026-04-2044.2%
2026-07-21
Source-checked
25Claude Opus 4.6 (Adaptive Reasoning, Max Effort)Anthropic2026-02-0543.7%
2026-07-21
Source-checked
26GPT-5.5 (low)OpenAI2026-04-2343.5%
2026-07-21
Source-checked
27Muse SparkOther2026-01-0143.1%
2026-07-21
Source-checked
28Claude Opus 4.7 (Non-reasoning, High Effort)Anthropic2026-04-1642.7%
2026-07-21
Source-checked
29GPT-5.2OpenAI2025-12-1142.2%
2026-07-21
Source-checked
30MiMo-V2.5-ProXiaomi2026-04-2242.2%
2026-07-21
Source-checked
31Kimi K2.7 CodeMoonshot2026-06-1241.9%
2026-07-21
Source-checked
32GLM-5.1Zhipu2026-04-0740.2%
2026-07-21
Source-checked
33GPT-5.4 MiniOpenAI2026-03-1740%
2026-07-21
Source-checked
34Gemini 3 ProGoogle2025-11-1839.6%
2026-07-21
Source-checked
35Qwen 3.7 PlusAlibaba2026-06-0239%
2026-07-21
Source-checked
36GPT-5.4 NanoOpenAI2026-03-1738.2%
2026-07-21
Source-checked
37MiniMax M2.7MiniMax2026-04-1538.1%
2026-07-21
Source-checked
38Claude Opus 4.6Anthropic2026-02-0537.8%
2026-07-21
Source-checked
39Nemotron 3 Ultra 550B A55B (Reasoning)NVIDIA2026-06-0437.8%
2026-07-21
Source-checked
40Grok 4.3xAI2026-04-3037.6%
2026-07-21
Source-checked
41Grok 4.20 ReasoningxAI2026-03-0537%
2026-07-21
Source-checked
42GPT-5.1OpenAI2025-11-1336.9%
2026-07-21
Source-checked
43Claude Sonnet 4.6Anthropic2026-02-1735.9%
2026-07-21
Source-checked
44Kimi K2.5Moonshot2026-01-2735.4%
2026-07-21
Source-checked
45Claude Opus 4.5Anthropic2025-11-2434.7%
2026-07-21
Source-checked
46MiniMax M2.5MiniMax2026-04-0133.7%
2026-07-21
Source-checked
47LongCat-2.0Meituan2026-06-2933.5%
2026-07-21
Source-checked
48OpenAI o3OpenAI2025-04-1630.4%
2026-07-21
Source-checked
49Step 3.7 FlashStepFun2026-06-0130.3%
2026-07-21
Source-checked
50Gemma 4 31B ITGoogle2026-03-0129.4%
2026-07-21
Source-checked
51Gemini 3 Flash PreviewGoogle2025-12-1727.4%
2026-07-21
Source-checked
52Gemini 3.1 Flash LiteGoogle2026-05-0725%
2026-07-21
Source-checked
53OpenAI o1OpenAI2024-09-1223.4%
2026-07-21
Source-checked
54Gemma 4 12B (Reasoning)Google2026-06-0722%
2026-07-21
Source-checked
55DeepSeek R1DeepSeek2025-01-2020.1%
2026-07-21
Source-checked
56North Mini CodeCohere2026-06-0919.8%
2026-07-21
Source-checked
57HyperNova 60B 2605Multiverse2026-05-0617.8%
2026-07-21
Source-checked
58Mistral Large 3Mistral2025-12-0215.9%
2026-07-21
Source-checked
59Llama 4 MaverickMeta2026-04-0514.3%
2026-07-21
Source-checked
60Gemma 4 12B (Non-reasoning)Google2026-06-1013.2%
2026-07-21
Source-checked
61MiniCPM5-1B (Reasoning)OpenBMB2026-06-0412%
2026-07-21
Source-checked
62Claude 3 OpusAnthropic2024-03-0411.8%
2026-07-21
Source-checked
63Llama 4 ScoutMeta2026-04-0510%
2026-07-21
Source-checked
64LFM2.5-8B-A1BLiquid2026-06-088.3%
2026-07-21
Source-checked
65Command ACohere2026-03-157.7%
2026-07-21
Source-checked
66DeepSeek Coder V2DeepSeek2024-05-155.1%
2026-07-21
Source-checked
67Gemini 1.0 UltraGoogle2024-02-084.6%
2026-07-21
Source-checked
Method

What it covers

Scores in the ledger are snapshotted from artificialanalysis.ai/models with source date and ingest timestamp on each row.

Score ceiling

Where it breaks down

v4.1 upgraded GDPval, Terminal-Bench, and τ-bench; removed IFBench for saturation. Do not compare v4.1 headline scores to pre-2026 AA index versions.

Tools that report it

Receipts

Sources and further reading

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.