Benchmarks / Reasoning

Artificial Analysis Intelligence Index

AA Intelligence Index

AA Intelligence Index v4.2 — weighted composite across four categories: Agents 30% (AA-Briefcase 15%, GDPval-AA v2 10%, τ³-Banking 5%), Coding 20% (Terminal-Bench 2.1 10%, SciCode 10%), Scientific Reasoning 20% (HLE 10%, CritPt 10%), General 30% (AA-Omniscience Accuracy 10% + Non-Hallucination 5%, GDP.pdf 10%, AA-LCR v1.1 5%).

What this does not measure
  • Latency, cost-per-token, and throughput — the index is intelligence-only. Use the AA API alongside for $/1M tokens.
  • Multilingual or multimodal ability — AA benchmarks image, speech, and multilingual performance on separate indices.
  • Any single sub-test in isolation — a high composite can hide weak agentic or coding slices.
Analysis

Why this benchmark is useful

Single-number model rankings hide tradeoffs. The AA Intelligence Index is a weighted composite of public sub-evals — useful as a dated frontier snapshot when you read the slices, not as a universal quality score.

Scope

Coverage map

Task family
Reasoning
Format
Composite of ten versioned sub-evaluations with public category weights. Methodology page lists per-eval question counts, repeats, and scoring. AA-Briefcase and GDP.pdf are new in v4.2.
Scoring
Weighted percentage across four categories (Agents 30%, Coding 20%, Scientific Reasoning 20%, General 30%). AA estimates a 95% CI under ±1% on the composite for models with enough repeats.
Maintainer
Artificial Analysis
Reading guide

How to read the scores

Open the methodology page for v4.2 weights, then drill into each sub-eval. On 5 Sep 2026 the live board is Fable 5.1 at 57, GPT-6 Astra at 55, Opus 5 at 54, Muse Spark 1.3 at 53, then Grok 4.6 and GPT-5.6 Sol at 51.

Blind spots

What it does not cover

  • Latency, cost-per-token, and throughput — the index is intelligence-only. Use the AA API alongside for $/1M tokens.
  • Multilingual or multimodal ability — AA benchmarks image, speech, and multilingual performance on separate indices.
  • Any single sub-test in isolation — a high composite can hide weak agentic or coding slices.
  • v4.1 headlines — GPQA Diamond left the index in v4.2; AA-Briefcase and GDP.pdf entered. A drop from a September v4.1 score is usually recomposition, not a weaker model.
Scores

Evidence ledger

187 rows
187187 rows
62.5%best score
187source-checked
1sources
2026-09-05to 2026-09-01
187

api · api

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Claude Fable 5.157% ·
  2. GPT-6 Astra55% ·
  3. Claude Opus 5 (Adaptive Reasoning, Max Effort)54% ·
  4. Muse Spark 1.353% ·
  5. GPT-5.6 Sol51% ·

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)Anthropic2026-07-2462.5%
2026-09-01
Source-checked
2Claude Fable 5Anthropic2026-06-0962.1%
2026-09-01
Source-checked
3Claude Opus 5 (Adaptive Reasoning, High Effort)Anthropic2026-07-2461.5%
2026-09-01
Source-checked
4Kimi K3Moonshot2026-07-1659.7%
2026-09-01
Source-checked
5GLM-5.3Zhipu2026-08-1459.5%
2026-09-01
Source-checked
6Claude Opus 5 (Adaptive Reasoning, Medium Effort)Anthropic2026-07-2458.6%
2026-09-01
Source-checked
7Qwen3.8 MaxAlibaba2026-08-0358.1%
2026-09-01
Source-checked
8Qwen3.8 2.4T A95BAlibaba2026-08-1257.7%
2026-09-01
Source-checked
9GLM-5.3-FlashZhipu2026-08-2657.5%
2026-09-01
Source-checked
10Claude Opus 4.8Anthropic2026-05-2857.3%
2026-09-01
Source-checked
1Claude Fable 5.1AnthropicPublic AA model page, 5 Sep 2026, Intelligence Index v4.2. Max + default fallback. The 1 Sep v4.1 headline was 66 — the drop is index recomposition, not a model regression.2026-09-0157%
2026-09-05
Source-checked
12Muse Spark 1.2Meta2026-08-0556.8%
2026-09-01
Source-checked
13GPT-5.6 TerraOpenAI2026-07-0956.6%
2026-09-01
Source-checked
14GPT 5.5OpenAI2026-04-2356.3%
2026-09-01
Source-checked
15Grok 4.5xAI2026-07-0855.8%
2026-09-01
Source-checked
16Qwen3.8-Flash-NextAlibaba2026-08-2655.8%
2026-09-01
Source-checked
17Claude Sonnet 5Anthropic2026-06-3055.3%
2026-09-01
Source-checked
2GPT-6 AstraOpenAIPublic AA model page, 5 Sep 2026, max effort, Intelligence Index v4.2.2026-09-0355%
2026-09-05
Source-checked
19Claude Opus 4.7Anthropic2026-04-1655%
2026-09-01
Source-checked
20GPT-5.5 (high)OpenAI2026-04-2354.7%
2026-09-01
Source-checked
4Claude Opus 5 (Adaptive Reasoning, Max Effort)AnthropicPublic AA model page, 5 Sep 2026, max effort, Intelligence Index v4.2. Replaces the 1 Sep v4.1 snapshot row of 63.1.2026-07-2454%
2026-09-05
Source-checked
22DeepSeek V4 Pro 0813 (Reasoning, Max Effort)DeepSeek2026-08-1353.2%
2026-09-01
Source-checked
23Muse Spark 1.1Meta2026-07-0953.2%
2026-09-01
Source-checked
24GPT 5.4OpenAI2026-03-0553.1%
2026-09-01
Source-checked
9Muse Spark 1.3MetaPublic AA model page, 5 Sep 2026, max effort, Intelligence Index v4.2. Rank #9 / 202 in that price class.2026-09-0253%
2026-09-05
Source-checked
26GLM-5.2Zhipu2026-06-1352.6%
2026-09-01
Source-checked
27Claude Opus 5 (Adaptive Reasoning, Low Effort)Anthropic2026-07-2452.5%
2026-09-01
Source-checked
28GPT-5.6 LunaOpenAI2026-07-0952.3%
2026-09-01
Source-checked
29Gemini 3.5 FlashGoogle2026-05-1952%
2026-09-01
Source-checked
30Qwen3.8 27BAlibaba2026-08-1452%
2026-09-01
Source-checked
31DeepSeek V4 Flash 0731 (Reasoning, Max Effort)DeepSeek2026-07-3151.8%
2026-09-01
Source-checked
32Gemini 3.6 Flash (high)Google2026-07-2151.6%
2026-09-01
Source-checked
33DeepSeek V4 Flash Vision (Reasoning, Max Effort)DeepSeek2026-08-2151.5%
2026-09-01
Source-checked
34GPT-5.5 (medium)OpenAI2026-04-2351.4%
2026-09-01
Source-checked
14GPT-5.6 SolOpenAIPublic AA model page, 5 Sep 2026, max effort, Intelligence Index v4.2. Replaces the 1 Sep v4.1 snapshot row of 60.9.2026-07-0951%
2026-09-05
Source-checked
15Grok 4.6xAIPublic AA model page, 5 Sep 2026, high effort, Intelligence Index v4.2. Replaces the 1 Sep v4.1 snapshot row of 60.9.2026-08-1251%
2026-09-05
Source-checked
37Agnes 2.5 Pro BetaOther2026-08-2649.1%
2026-09-01
Source-checked
38Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)Anthropic2026-02-1748.4%
2026-09-01
Source-checked
39Kimi K3 (low)Moonshot2026-07-1648.3%
2026-09-01
Source-checked
40Gemini 3.1 Pro PreviewGoogle2026-02-1947.7%
2026-09-01
Source-checked
41Motif 3Other2026-08-1247.4%
2026-09-01
Source-checked
27Gemini 3.8 FlashGooglePublic AA model page, 5 Sep 2026, high effort, Intelligence Index v4.2.2026-09-0247%
2026-09-05
Source-checked
43Gemini 3.5 Flash (medium)Google2026-05-1946.7%
2026-09-01
Source-checked
44Qwen3.7 MaxAlibaba2026-05-1946.7%
2026-09-01
Source-checked
45GPT-5.3 Codex (xhigh)OpenAI2026-02-0545.5%
2026-09-01
Source-checked
46MiniMax M3MiniMax2026-05-3145.4%
2026-09-01
Source-checked
47Motif 3 (Beta)Other2026-07-2145.3%
2026-09-01
Source-checked
48Kimi K2.6Moonshot2026-04-2045.1%
2026-09-01
Source-checked
37Gemini 3.7 FlashGooglePublic AA model page, 5 Sep 2026, high effort, Intelligence Index v4.2. Replaces the 1 Sep v4.1 snapshot row of 56.2026-08-1345%
2026-09-05
Source-checked
50Claude Opus 4.6 (Adaptive Reasoning, Max Effort)Anthropic2026-02-0544.9%
2026-09-01
Source-checked
51GPT-5.5 (low)OpenAI2026-04-2344.5%
2026-09-01
Source-checked
52Muse SparkOther2026-01-0144.3%
2026-09-01
Source-checked
53Apodex 1.1Other2026-08-3044%
2026-09-01
Source-checked
54Claude Opus 4.7 (Non-reasoning, High Effort)Anthropic2026-04-1643.9%
2026-09-01
Source-checked
55DeepSeek V4 Pro (Reasoning, High Effort)DeepSeek2026-04-2443.7%
2026-09-01
Source-checked
56GPT-5.2OpenAI2025-12-1143.3%
2026-09-01
Source-checked
57Kimi K2.7 CodeMoonshot2026-06-1243%
2026-09-01
Source-checked
58MiMo-V2.5-ProXiaomi2026-04-2242.9%
2026-09-01
Source-checked
59Inkling (xhigh)Other2026-07-1542.3%
2026-09-01
Source-checked
60Hy3Other2026-07-0642.2%
2026-09-01
Source-checked
61Claude Opus 4.5 (Reasoning)Anthropic2025-11-2441.9%
2026-09-01
Source-checked
62Nex-N2-ProOther2026-06-0241.7%
2026-09-01
Source-checked
63Solar Pro 4Other2026-08-0641.6%
2026-09-01
Source-checked
64MiMo-V2-ProXiaomi2026-03-1841.4%
2026-09-01
Source-checked
65GPT-5.2 Codex (xhigh)OpenAI2025-12-1141.2%
2026-09-01
Source-checked
66Inkling SmallOther2026-07-3041.2%
2026-09-01
Source-checked
67Qwen3.6 Max PreviewAlibaba2026-04-2041.1%
2026-09-01
Source-checked
68GLM-5.1Zhipu2026-04-0741%
2026-09-01
Source-checked
69GPT-5.4 MiniOpenAI2026-03-1740.9%
2026-09-01
Source-checked
70Grok Build 0.1 0616xAI2026-06-1640.7%
2026-09-01
Source-checked
71Gemini 3 ProGoogle2025-11-1840.6%
2026-09-01
Source-checked
72GLM-5 (Reasoning)Zhipu2026-02-1140.6%
2026-09-01
Source-checked
73Qwen3.6 PlusAlibaba2026-04-0240.5%
2026-09-01
Source-checked
74GPT-5.4 (low)OpenAI2026-03-0540.2%
2026-09-01
Source-checked
75JT-4.1 Flash 236B A21BOther2026-07-0939.9%
2026-09-01
Source-checked
76Agnes 2.5 Pro AlphaOther2026-07-2439.7%
2026-09-01
Source-checked
77GPT-5.4 NanoOpenAI2026-03-1739.7%
2026-09-01
Source-checked
78Qwen 3.7 PlusAlibaba2026-06-0239.4%
2026-09-01
Source-checked
79GLM-5-TurboZhipu2026-03-1539.1%
2026-09-01
Source-checked
80DeepSeek V4 Flash (Reasoning, High Effort)DeepSeek2026-04-2439%
2026-09-01
Source-checked
81GPT-5.2 (medium)OpenAI2025-12-1138.9%
2026-09-01
Source-checked
82MiniMax M2.7MiniMax2026-04-1538.9%
2026-09-01
Source-checked
83Claude Opus 4.6Anthropic2026-02-0538.8%
2026-09-01
Source-checked
84Gemini 3 Flash Preview (Reasoning)Google2025-12-1738.7%
2026-09-01
Source-checked
85Nemotron 3 Ultra 550B A55B (Reasoning)NVIDIA2026-06-0438.3%
2026-09-01
Source-checked
86Grok 4.20 ReasoningxAI2026-03-0538%
2026-09-01
Source-checked
87MiMo-V2.5Xiaomi2026-04-2238%
2026-09-01
Source-checked
88Grok 4.3xAI2026-04-3037.9%
2026-09-01
Source-checked
89Ling 3.0 FlashOther2026-08-0437.8%
2026-09-01
Source-checked
90Qwen3.6 27B (Reasoning)Alibaba2026-04-2237.7%
2026-09-01
Source-checked
91GPT-5.1OpenAI2025-11-1337.5%
2026-09-01
Source-checked
92Claude 4.5 Sonnet (Reasoning)Anthropic2025-09-2937.4%
2026-09-01
Source-checked
93Gemini 3.5 Flash-LiteGoogle2026-07-2137.4%
2026-09-01
Source-checked
94Grok 4.20 0309 (Reasoning)xAI2026-03-1037.4%
2026-09-01
Source-checked
95Solar Open2 250BOther2026-08-1237.4%
2026-09-01
Source-checked
96MiMo-V2-Omni-0327Xiaomi2026-03-2737.3%
2026-09-01
Source-checked
97GPT-5 Codex (high)OpenAI2025-09-2337%
2026-09-01
Source-checked
98Grok 4.3 (medium)xAI2026-04-3036.9%
2026-09-01
Source-checked
99Claude Sonnet 4.6Anthropic2026-02-1736.8%
2026-09-01
Source-checked
100Grok 4.3 (low)xAI2026-04-3036.3%
2026-09-01
Source-checked
101Kimi K2.5Moonshot2026-01-2736%
2026-09-01
Source-checked
102MiMo-V2-OmniXiaomi2026-03-1935.9%
2026-09-01
Source-checked
103Gemini 3.5 Flash (minimal)Google2026-05-1935.8%
2026-09-01
Source-checked
104GPT-5.5 (Non-reasoning)OpenAI2026-04-2335.8%
2026-09-01
Source-checked
105Claude Opus 4.5Anthropic2025-11-2435.6%
2026-09-01
Source-checked
106GPT-5.1 Codex (high)OpenAI2025-11-1335.6%
2026-09-01
Source-checked
107Kimi K2.6 (Non-reasoning)Other2026-04-2035.4%
2026-09-01
Source-checked
108GLM 5V Turbo (Reasoning)Zhipu2026-04-0135.3%
2026-09-01
Source-checked
109GPT-5 (high)OpenAI2025-08-0735.3%
2026-09-01
Source-checked
110Claude Sonnet 4.6 (Non-reasoning, Low Effort)Anthropic2026-02-1735.1%
2026-09-01
Source-checked
111Muse GlimmerMeta2026-08-1035.1%
2026-09-01
Source-checked
112A.X-K2Other2026-08-1235%
2026-09-01
Source-checked
113GLM-5.2 (Non-reasoning)Zhipu2026-06-1634.8%
2026-09-01
Source-checked
114GPT-5 (medium)OpenAI2025-08-0734.6%
2026-09-01
Source-checked
115Qwen3.5 27B (Reasoning)Alibaba2026-02-2434.6%
2026-09-01
Source-checked
116Claude 4.1 Opus (Reasoning)Anthropic2025-08-0534.5%
2026-09-01
Source-checked
117GLM-4.7 (Reasoning)Zhipu2025-12-2234.5%
2026-09-01
Source-checked
118MiniMax M2.5MiniMax2026-04-0134.5%
2026-09-01
Source-checked
119Hy3-preview (Reasoning)Other2026-04-2334.4%
2026-09-01
Source-checked
120GPT-5.5 Instant (May 2026)OpenAI2026-05-0534.3%
2026-09-01
Source-checked
121Qwen3.5 397B A17B (Reasoning)Alibaba2026-02-1634.3%
2026-09-01
Source-checked
122Grok 4xAI2025-07-1034.1%
2026-09-01
Source-checked
123G9v3-39A5BOther2026-08-0334%
2026-09-01
Source-checked
124LongCat-2.0Meituan2026-06-2934%
2026-09-01
Source-checked
125MiMo-V2-Flash (Feb 2026)Xiaomi2025-12-1634%
2026-09-01
Source-checked
126Gemini 3 Pro Preview (low)Google2025-11-1833.9%
2026-09-01
Source-checked
127KAT Coder Pro V2Other2026-03-2733.7%
2026-09-01
Source-checked
128Kimi K2 ThinkingMoonshot2025-11-0633.5%
2026-09-01
Source-checked
129o3-proOpenAI2025-06-1033.3%
2026-09-01
Source-checked
130GLM-5 (Non-reasoning)Zhipu2026-02-1133.2%
2026-09-01
Source-checked
131DeepSeek V3.2 (Reasoning)DeepSeek2025-12-0132.8%
2026-09-01
Source-checked
132Qwen3.5 122B A10B (Reasoning)Alibaba2026-02-2432.8%
2026-09-01
Source-checked
133Qwen3.5 397B A17B (Non-reasoning)Alibaba2026-02-1632.7%
2026-09-01
Source-checked
134Qwen3 Max ThinkingAlibaba2026-01-2632.5%
2026-09-01
Source-checked
135MiniMax-M2.1MiniMax2025-12-2332.1%
2026-09-01
Source-checked
136Qwen3.6 35B A3B (Reasoning)Alibaba2026-04-1632.1%
2026-09-01
Source-checked
137DeepSeek V4 Pro (Non-reasoning)DeepSeek2026-04-2431.9%
2026-09-01
Source-checked
138GPT-5 (low)OpenAI2025-08-0731.9%
2026-09-01
Source-checked
139MiMo-V2-Flash (Reasoning)Xiaomi2025-12-1631.9%
2026-09-01
Source-checked
140Claude 4 Opus (Reasoning)Anthropic2025-05-2231.7%
2026-09-01
Source-checked
141Ring-2.6-1TOther2026-05-0831.7%
2026-09-01
Source-checked
142Qwen3.6 27B (Non-reasoning)Alibaba2026-04-2231.3%
2026-09-01
Source-checked
143OpenAI o3OpenAI2025-04-1631.1%
2026-09-01
Source-checked
144K-EXAONE 2.0 0803Other2026-08-1231%
2026-09-01
Source-checked
145Step 3.7 FlashStepFun2026-06-0130.9%
2026-09-01
Source-checked
146Mistral Medium 3.5Mistral2026-04-2930.4%
2026-09-01
Source-checked
147Gemma 4 31B ITGoogle2026-03-0129.7%
2026-09-01
Source-checked
148DeepSeek V4 Flash (Non-reasoning)DeepSeek2026-04-2429.3%
2026-09-01
Source-checked
149GPT-5.5 Instant (June 2026)OpenAI2026-06-2529.2%
2026-09-01
Source-checked
150JT-35B-FlashOther2026-05-1429%
2026-09-01
Source-checked
151MiMo-V2.5-Pro (Non-reasoning)Xiaomi2026-04-2228.4%
2026-09-01
Source-checked
152Gemini 3 Flash PreviewGoogle2025-12-1727.9%
2026-09-01
Source-checked
153Hy3-preview (Non-reasoning)Other2026-04-2326.6%
2026-09-01
Source-checked
154Ling-2.6-1TOther2026-04-2326.6%
2026-09-01
Source-checked
155Gemini 3.1 Flash LiteGoogle2026-05-0725.6%
2026-09-01
Source-checked
156Grok 4.3 (Non-reasoning)xAI2026-04-3025%
2026-09-01
Source-checked
157Qwen3.6 35B A3B (Non-reasoning)Alibaba2026-04-1624.6%
2026-09-01
Source-checked
158Ling 3.0 TinyOther2026-08-0624.5%
2026-09-01
Source-checked
159OpenAI o1OpenAI2024-09-1223.9%
2026-09-01
Source-checked
160Granite 4.2 30BOther2026-08-2523.7%
2026-09-01
Source-checked
161Nemotron 3.5 LightningNVIDIA2026-08-1123.6%
2026-09-01
Source-checked
162Command ACohere2026-03-1522.8%
2026-09-01
Source-checked
163Gemma 4 12B (Reasoning)Google2026-06-0722.2%
2026-09-01
Source-checked
164Grok 4.20 0309 v2 (Non-reasoning)xAI2026-04-0722.2%
2026-09-01
Source-checked
165EXAONE 4.5 33BOther2026-04-0920.5%
2026-09-01
Source-checked
166DeepSeek R1DeepSeek2025-01-2020.4%
2026-09-01
Source-checked
167North Mini CodeCohere2026-06-0920.2%
2026-09-01
Source-checked
168Granite 4.2 8BOther2026-08-2519.6%
2026-09-01
Source-checked
169JT-MINIOther2026-04-1518.8%
2026-09-01
Source-checked
170HyperNova 60B 2605Multiverse2026-05-0618.3%
2026-09-01
Source-checked
171G9v3-3BOther2026-07-2316.2%
2026-09-01
Source-checked
172Mistral Large 3Mistral2025-12-0215.9%
2026-09-01
Source-checked
173Nemotron 3 Nano Omni 30B A3B ReasoningNVIDIA2026-04-2915%
2026-09-01
Source-checked
174Llama 4 MaverickMeta2026-04-0514.5%
2026-09-01
Source-checked
175Granite 4.2 3BOther2026-08-2514.3%
2026-09-01
Source-checked
176Ling 2.6 FlashOther2026-04-2114.2%
2026-09-01
Source-checked
177DiffusionGemma 26B A4BGoogle2026-06-1013.5%
2026-09-01
Source-checked
178Gemma 4 12B (Non-reasoning)Google2026-06-1013.2%
2026-09-01
Source-checked
179MiniCPM5-1B (Reasoning)OpenBMB2026-06-0411.9%
2026-09-01
Source-checked
180Claude 3 OpusAnthropic2024-03-0411.8%
2026-09-01
Source-checked
181MiniCPM5-1B (Non-reasoning)OpenBMB2026-05-2511.7%
2026-09-01
Source-checked
182Llama 4 ScoutMeta2026-04-0510.3%
2026-09-01
Source-checked
183Granite 4.1 30BOther2026-04-298.7%
2026-09-01
Source-checked
184LFM2.5-8B-A1BLiquid2026-06-088.1%
2026-09-01
Source-checked
185DeepSeek Coder V2DeepSeek2024-05-154.7%
2026-09-01
Source-checked
186Gemini 1.0 UltraGoogle2024-02-084.3%
2026-09-01
Source-checked
187MiniCPM-V 4.6 1.3BOpenBMB2026-05-113.8%
2026-09-01
Source-checked
Method

What it covers

Frontier rows dated 2026-09-05 were read from public AA model pages on Intelligence Index v4.2. The committed API snapshot is still 2026-09-01 (v4.1). AA-Briefcase and GDP.pdf do not yet have standalone VerdictPal explainers — they lack five independently extracted model rows.

Score ceiling

Where it breaks down

v4.2 added AA-Briefcase (15%) and GDP.pdf (10%) and removed GPQA Diamond. Do not compare v4.2 headlines to v4.1 or pre-2026 AA index versions. Sept 1 API-snapshot rows in the ledger are v4.1; hand-updated frontier rows dated 5 Sep 2026 are v4.2.

Tools that report it

Receipts

Sources and further reading

Questions about this benchmark

Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.

What does Artificial Analysis Intelligence Index measure?

AA Intelligence Index v4.2 — weighted composite across four categories: Agents 30% (AA-Briefcase 15%, GDPval-AA v2 10%, τ³-Banking 5%), Coding 20% (Terminal-Bench 2.1 10%, SciCode 10%), Scientific Reasoning 20% (HLE 10%, CritPt 10%), General 30% (AA-Omniscience Accuracy 10% + Non-Hallucination 5%, GDP.pdf 10%, AA-LCR v1.1 5%).

What does a high Artificial Analysis Intelligence Index score not prove?

A strong Artificial Analysis Intelligence Index result says nothing about:

  • Latency, cost-per-token, and throughput — the index is intelligence-only. Use the AA API alongside for $/1M tokens.
  • Multilingual or multimodal ability — AA benchmarks image, speech, and multilingual performance on separate indices.
  • Any single sub-test in isolation — a high composite can hide weak agentic or coding slices.
  • v4.1 headlines — GPQA Diamond left the index in v4.2; AA-Briefcase and GDP.pdf entered. A drop from a September v4.1 score is usually recomposition, not a weaker model.

How is Artificial Analysis Intelligence Index scored?

Weighted percentage across four categories (Agents 30%, Coding 20%, Scientific Reasoning 20%, General 30%). AA estimates a 95% CI under ±1% on the composite for models with enough repeats. Task format: Composite of ten versioned sub-evaluations with public category weights. Methodology page lists per-eval question counts, repeats, and scoring. AA-Briefcase and GDP.pdf are new in v4.2.

Is Artificial Analysis Intelligence Index saturated?

Artificial Analysis Intelligence Index is currently marked Active in the atlas.

Can Artificial Analysis Intelligence Index results be contaminated by training data?

Contamination risk for Artificial Analysis Intelligence Index is graded Medium contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.