Why this benchmark is useful
Single-number model rankings hide tradeoffs. The AA Intelligence Index is a weighted composite of public sub-evals — useful as a dated frontier snapshot when you read the slices, not as a universal quality score.
AA Intelligence Index
AA Intelligence Index v4.2 — weighted composite across four categories: Agents 30% (AA-Briefcase 15%, GDPval-AA v2 10%, τ³-Banking 5%), Coding 20% (Terminal-Bench 2.1 10%, SciCode 10%), Scientific Reasoning 20% (HLE 10%, CritPt 10%), General 30% (AA-Omniscience Accuracy 10% + Non-Hallucination 5%, GDP.pdf 10%, AA-LCR v1.1 5%).
Single-number model rankings hide tradeoffs. The AA Intelligence Index is a weighted composite of public sub-evals — useful as a dated frontier snapshot when you read the slices, not as a universal quality score.
Open the methodology page for v4.2 weights, then drill into each sub-eval. On 5 Sep 2026 the live board is Fable 5.1 at 57, GPT-6 Astra at 55, Opus 5 at 54, Muse Spark 1.3 at 53, then Grok 4.6 and GPT-5.6 Sol at 51.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)Anthropic | 2026-07-24 | 62.5% | 2026-09-01 | Source-checked |
| 2 | Claude Fable 5Anthropic | 2026-06-09 | 62.1% | 2026-09-01 | Source-checked |
| 3 | Claude Opus 5 (Adaptive Reasoning, High Effort)Anthropic | 2026-07-24 | 61.5% | 2026-09-01 | Source-checked |
| 4 | Kimi K3Moonshot | 2026-07-16 | 59.7% | 2026-09-01 | Source-checked |
| 5 | GLM-5.3Zhipu | 2026-08-14 | 59.5% | 2026-09-01 | Source-checked |
| 6 | Claude Opus 5 (Adaptive Reasoning, Medium Effort)Anthropic | 2026-07-24 | 58.6% | 2026-09-01 | Source-checked |
| 7 | Qwen3.8 MaxAlibaba | 2026-08-03 | 58.1% | 2026-09-01 | Source-checked |
| 8 | Qwen3.8 2.4T A95BAlibaba | 2026-08-12 | 57.7% | 2026-09-01 | Source-checked |
| 9 | GLM-5.3-FlashZhipu | 2026-08-26 | 57.5% | 2026-09-01 | Source-checked |
| 10 | Claude Opus 4.8Anthropic | 2026-05-28 | 57.3% | 2026-09-01 | Source-checked |
| 1 | Claude Fable 5.1AnthropicPublic AA model page, 5 Sep 2026, Intelligence Index v4.2. Max + default fallback. The 1 Sep v4.1 headline was 66 — the drop is index recomposition, not a model regression. | 2026-09-01 | 57% | 2026-09-05 | Source-checked |
| 12 | Muse Spark 1.2Meta | 2026-08-05 | 56.8% | 2026-09-01 | Source-checked |
| 13 | GPT-5.6 TerraOpenAI | 2026-07-09 | 56.6% | 2026-09-01 | Source-checked |
| 14 | GPT 5.5OpenAI | 2026-04-23 | 56.3% | 2026-09-01 | Source-checked |
| 15 | Grok 4.5xAI | 2026-07-08 | 55.8% | 2026-09-01 | Source-checked |
| 16 | Qwen3.8-Flash-NextAlibaba | 2026-08-26 | 55.8% | 2026-09-01 | Source-checked |
| 17 | Claude Sonnet 5Anthropic | 2026-06-30 | 55.3% | 2026-09-01 | Source-checked |
| 2 | GPT-6 AstraOpenAIPublic AA model page, 5 Sep 2026, max effort, Intelligence Index v4.2. | 2026-09-03 | 55% | 2026-09-05 | Source-checked |
| 19 | Claude Opus 4.7Anthropic | 2026-04-16 | 55% | 2026-09-01 | Source-checked |
| 20 | GPT-5.5 (high)OpenAI | 2026-04-23 | 54.7% | 2026-09-01 | Source-checked |
| 4 | Claude Opus 5 (Adaptive Reasoning, Max Effort)AnthropicPublic AA model page, 5 Sep 2026, max effort, Intelligence Index v4.2. Replaces the 1 Sep v4.1 snapshot row of 63.1. | 2026-07-24 | 54% | 2026-09-05 | Source-checked |
| 22 | DeepSeek V4 Pro 0813 (Reasoning, Max Effort)DeepSeek | 2026-08-13 | 53.2% | 2026-09-01 | Source-checked |
| 23 | Muse Spark 1.1Meta | 2026-07-09 | 53.2% | 2026-09-01 | Source-checked |
| 24 | GPT 5.4OpenAI | 2026-03-05 | 53.1% | 2026-09-01 | Source-checked |
| 9 | Muse Spark 1.3MetaPublic AA model page, 5 Sep 2026, max effort, Intelligence Index v4.2. Rank #9 / 202 in that price class. | 2026-09-02 | 53% | 2026-09-05 | Source-checked |
| 26 | GLM-5.2Zhipu | 2026-06-13 | 52.6% | 2026-09-01 | Source-checked |
| 27 | Claude Opus 5 (Adaptive Reasoning, Low Effort)Anthropic | 2026-07-24 | 52.5% | 2026-09-01 | Source-checked |
| 28 | GPT-5.6 LunaOpenAI | 2026-07-09 | 52.3% | 2026-09-01 | Source-checked |
| 29 | Gemini 3.5 FlashGoogle | 2026-05-19 | 52% | 2026-09-01 | Source-checked |
| 30 | Qwen3.8 27BAlibaba | 2026-08-14 | 52% | 2026-09-01 | Source-checked |
| 31 | DeepSeek V4 Flash 0731 (Reasoning, Max Effort)DeepSeek | 2026-07-31 | 51.8% | 2026-09-01 | Source-checked |
| 32 | Gemini 3.6 Flash (high)Google | 2026-07-21 | 51.6% | 2026-09-01 | Source-checked |
| 33 | DeepSeek V4 Flash Vision (Reasoning, Max Effort)DeepSeek | 2026-08-21 | 51.5% | 2026-09-01 | Source-checked |
| 34 | GPT-5.5 (medium)OpenAI | 2026-04-23 | 51.4% | 2026-09-01 | Source-checked |
| 14 | GPT-5.6 SolOpenAIPublic AA model page, 5 Sep 2026, max effort, Intelligence Index v4.2. Replaces the 1 Sep v4.1 snapshot row of 60.9. | 2026-07-09 | 51% | 2026-09-05 | Source-checked |
| 15 | Grok 4.6xAIPublic AA model page, 5 Sep 2026, high effort, Intelligence Index v4.2. Replaces the 1 Sep v4.1 snapshot row of 60.9. | 2026-08-12 | 51% | 2026-09-05 | Source-checked |
| 37 | Agnes 2.5 Pro BetaOther | 2026-08-26 | 49.1% | 2026-09-01 | Source-checked |
| 38 | Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)Anthropic | 2026-02-17 | 48.4% | 2026-09-01 | Source-checked |
| 39 | Kimi K3 (low)Moonshot | 2026-07-16 | 48.3% | 2026-09-01 | Source-checked |
| 40 | Gemini 3.1 Pro PreviewGoogle | 2026-02-19 | 47.7% | 2026-09-01 | Source-checked |
| 41 | Motif 3Other | 2026-08-12 | 47.4% | 2026-09-01 | Source-checked |
| 27 | Gemini 3.8 FlashGooglePublic AA model page, 5 Sep 2026, high effort, Intelligence Index v4.2. | 2026-09-02 | 47% | 2026-09-05 | Source-checked |
| 43 | Gemini 3.5 Flash (medium)Google | 2026-05-19 | 46.7% | 2026-09-01 | Source-checked |
| 44 | Qwen3.7 MaxAlibaba | 2026-05-19 | 46.7% | 2026-09-01 | Source-checked |
| 45 | GPT-5.3 Codex (xhigh)OpenAI | 2026-02-05 | 45.5% | 2026-09-01 | Source-checked |
| 46 | MiniMax M3MiniMax | 2026-05-31 | 45.4% | 2026-09-01 | Source-checked |
| 47 | Motif 3 (Beta)Other | 2026-07-21 | 45.3% | 2026-09-01 | Source-checked |
| 48 | Kimi K2.6Moonshot | 2026-04-20 | 45.1% | 2026-09-01 | Source-checked |
| 37 | Gemini 3.7 FlashGooglePublic AA model page, 5 Sep 2026, high effort, Intelligence Index v4.2. Replaces the 1 Sep v4.1 snapshot row of 56. | 2026-08-13 | 45% | 2026-09-05 | Source-checked |
| 50 | Claude Opus 4.6 (Adaptive Reasoning, Max Effort)Anthropic | 2026-02-05 | 44.9% | 2026-09-01 | Source-checked |
| 51 | GPT-5.5 (low)OpenAI | 2026-04-23 | 44.5% | 2026-09-01 | Source-checked |
| 52 | Muse SparkOther | 2026-01-01 | 44.3% | 2026-09-01 | Source-checked |
| 53 | Apodex 1.1Other | 2026-08-30 | 44% | 2026-09-01 | Source-checked |
| 54 | Claude Opus 4.7 (Non-reasoning, High Effort)Anthropic | 2026-04-16 | 43.9% | 2026-09-01 | Source-checked |
| 55 | DeepSeek V4 Pro (Reasoning, High Effort)DeepSeek | 2026-04-24 | 43.7% | 2026-09-01 | Source-checked |
| 56 | GPT-5.2OpenAI | 2025-12-11 | 43.3% | 2026-09-01 | Source-checked |
| 57 | Kimi K2.7 CodeMoonshot | 2026-06-12 | 43% | 2026-09-01 | Source-checked |
| 58 | MiMo-V2.5-ProXiaomi | 2026-04-22 | 42.9% | 2026-09-01 | Source-checked |
| 59 | Inkling (xhigh)Other | 2026-07-15 | 42.3% | 2026-09-01 | Source-checked |
| 60 | Hy3Other | 2026-07-06 | 42.2% | 2026-09-01 | Source-checked |
| 61 | Claude Opus 4.5 (Reasoning)Anthropic | 2025-11-24 | 41.9% | 2026-09-01 | Source-checked |
| 62 | Nex-N2-ProOther | 2026-06-02 | 41.7% | 2026-09-01 | Source-checked |
| 63 | Solar Pro 4Other | 2026-08-06 | 41.6% | 2026-09-01 | Source-checked |
| 64 | MiMo-V2-ProXiaomi | 2026-03-18 | 41.4% | 2026-09-01 | Source-checked |
| 65 | GPT-5.2 Codex (xhigh)OpenAI | 2025-12-11 | 41.2% | 2026-09-01 | Source-checked |
| 66 | Inkling SmallOther | 2026-07-30 | 41.2% | 2026-09-01 | Source-checked |
| 67 | Qwen3.6 Max PreviewAlibaba | 2026-04-20 | 41.1% | 2026-09-01 | Source-checked |
| 68 | GLM-5.1Zhipu | 2026-04-07 | 41% | 2026-09-01 | Source-checked |
| 69 | GPT-5.4 MiniOpenAI | 2026-03-17 | 40.9% | 2026-09-01 | Source-checked |
| 70 | Grok Build 0.1 0616xAI | 2026-06-16 | 40.7% | 2026-09-01 | Source-checked |
| 71 | Gemini 3 ProGoogle | 2025-11-18 | 40.6% | 2026-09-01 | Source-checked |
| 72 | GLM-5 (Reasoning)Zhipu | 2026-02-11 | 40.6% | 2026-09-01 | Source-checked |
| 73 | Qwen3.6 PlusAlibaba | 2026-04-02 | 40.5% | 2026-09-01 | Source-checked |
| 74 | GPT-5.4 (low)OpenAI | 2026-03-05 | 40.2% | 2026-09-01 | Source-checked |
| 75 | JT-4.1 Flash 236B A21BOther | 2026-07-09 | 39.9% | 2026-09-01 | Source-checked |
| 76 | Agnes 2.5 Pro AlphaOther | 2026-07-24 | 39.7% | 2026-09-01 | Source-checked |
| 77 | GPT-5.4 NanoOpenAI | 2026-03-17 | 39.7% | 2026-09-01 | Source-checked |
| 78 | Qwen 3.7 PlusAlibaba | 2026-06-02 | 39.4% | 2026-09-01 | Source-checked |
| 79 | GLM-5-TurboZhipu | 2026-03-15 | 39.1% | 2026-09-01 | Source-checked |
| 80 | DeepSeek V4 Flash (Reasoning, High Effort)DeepSeek | 2026-04-24 | 39% | 2026-09-01 | Source-checked |
| 81 | GPT-5.2 (medium)OpenAI | 2025-12-11 | 38.9% | 2026-09-01 | Source-checked |
| 82 | MiniMax M2.7MiniMax | 2026-04-15 | 38.9% | 2026-09-01 | Source-checked |
| 83 | Claude Opus 4.6Anthropic | 2026-02-05 | 38.8% | 2026-09-01 | Source-checked |
| 84 | Gemini 3 Flash Preview (Reasoning)Google | 2025-12-17 | 38.7% | 2026-09-01 | Source-checked |
| 85 | Nemotron 3 Ultra 550B A55B (Reasoning)NVIDIA | 2026-06-04 | 38.3% | 2026-09-01 | Source-checked |
| 86 | Grok 4.20 ReasoningxAI | 2026-03-05 | 38% | 2026-09-01 | Source-checked |
| 87 | MiMo-V2.5Xiaomi | 2026-04-22 | 38% | 2026-09-01 | Source-checked |
| 88 | Grok 4.3xAI | 2026-04-30 | 37.9% | 2026-09-01 | Source-checked |
| 89 | Ling 3.0 FlashOther | 2026-08-04 | 37.8% | 2026-09-01 | Source-checked |
| 90 | Qwen3.6 27B (Reasoning)Alibaba | 2026-04-22 | 37.7% | 2026-09-01 | Source-checked |
| 91 | GPT-5.1OpenAI | 2025-11-13 | 37.5% | 2026-09-01 | Source-checked |
| 92 | Claude 4.5 Sonnet (Reasoning)Anthropic | 2025-09-29 | 37.4% | 2026-09-01 | Source-checked |
| 93 | Gemini 3.5 Flash-LiteGoogle | 2026-07-21 | 37.4% | 2026-09-01 | Source-checked |
| 94 | Grok 4.20 0309 (Reasoning)xAI | 2026-03-10 | 37.4% | 2026-09-01 | Source-checked |
| 95 | Solar Open2 250BOther | 2026-08-12 | 37.4% | 2026-09-01 | Source-checked |
| 96 | MiMo-V2-Omni-0327Xiaomi | 2026-03-27 | 37.3% | 2026-09-01 | Source-checked |
| 97 | GPT-5 Codex (high)OpenAI | 2025-09-23 | 37% | 2026-09-01 | Source-checked |
| 98 | Grok 4.3 (medium)xAI | 2026-04-30 | 36.9% | 2026-09-01 | Source-checked |
| 99 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 36.8% | 2026-09-01 | Source-checked |
| 100 | Grok 4.3 (low)xAI | 2026-04-30 | 36.3% | 2026-09-01 | Source-checked |
| 101 | Kimi K2.5Moonshot | 2026-01-27 | 36% | 2026-09-01 | Source-checked |
| 102 | MiMo-V2-OmniXiaomi | 2026-03-19 | 35.9% | 2026-09-01 | Source-checked |
| 103 | Gemini 3.5 Flash (minimal)Google | 2026-05-19 | 35.8% | 2026-09-01 | Source-checked |
| 104 | GPT-5.5 (Non-reasoning)OpenAI | 2026-04-23 | 35.8% | 2026-09-01 | Source-checked |
| 105 | Claude Opus 4.5Anthropic | 2025-11-24 | 35.6% | 2026-09-01 | Source-checked |
| 106 | GPT-5.1 Codex (high)OpenAI | 2025-11-13 | 35.6% | 2026-09-01 | Source-checked |
| 107 | Kimi K2.6 (Non-reasoning)Other | 2026-04-20 | 35.4% | 2026-09-01 | Source-checked |
| 108 | GLM 5V Turbo (Reasoning)Zhipu | 2026-04-01 | 35.3% | 2026-09-01 | Source-checked |
| 109 | GPT-5 (high)OpenAI | 2025-08-07 | 35.3% | 2026-09-01 | Source-checked |
| 110 | Claude Sonnet 4.6 (Non-reasoning, Low Effort)Anthropic | 2026-02-17 | 35.1% | 2026-09-01 | Source-checked |
| 111 | Muse GlimmerMeta | 2026-08-10 | 35.1% | 2026-09-01 | Source-checked |
| 112 | A.X-K2Other | 2026-08-12 | 35% | 2026-09-01 | Source-checked |
| 113 | GLM-5.2 (Non-reasoning)Zhipu | 2026-06-16 | 34.8% | 2026-09-01 | Source-checked |
| 114 | GPT-5 (medium)OpenAI | 2025-08-07 | 34.6% | 2026-09-01 | Source-checked |
| 115 | Qwen3.5 27B (Reasoning)Alibaba | 2026-02-24 | 34.6% | 2026-09-01 | Source-checked |
| 116 | Claude 4.1 Opus (Reasoning)Anthropic | 2025-08-05 | 34.5% | 2026-09-01 | Source-checked |
| 117 | GLM-4.7 (Reasoning)Zhipu | 2025-12-22 | 34.5% | 2026-09-01 | Source-checked |
| 118 | MiniMax M2.5MiniMax | 2026-04-01 | 34.5% | 2026-09-01 | Source-checked |
| 119 | Hy3-preview (Reasoning)Other | 2026-04-23 | 34.4% | 2026-09-01 | Source-checked |
| 120 | GPT-5.5 Instant (May 2026)OpenAI | 2026-05-05 | 34.3% | 2026-09-01 | Source-checked |
| 121 | Qwen3.5 397B A17B (Reasoning)Alibaba | 2026-02-16 | 34.3% | 2026-09-01 | Source-checked |
| 122 | Grok 4xAI | 2025-07-10 | 34.1% | 2026-09-01 | Source-checked |
| 123 | G9v3-39A5BOther | 2026-08-03 | 34% | 2026-09-01 | Source-checked |
| 124 | LongCat-2.0Meituan | 2026-06-29 | 34% | 2026-09-01 | Source-checked |
| 125 | MiMo-V2-Flash (Feb 2026)Xiaomi | 2025-12-16 | 34% | 2026-09-01 | Source-checked |
| 126 | Gemini 3 Pro Preview (low)Google | 2025-11-18 | 33.9% | 2026-09-01 | Source-checked |
| 127 | KAT Coder Pro V2Other | 2026-03-27 | 33.7% | 2026-09-01 | Source-checked |
| 128 | Kimi K2 ThinkingMoonshot | 2025-11-06 | 33.5% | 2026-09-01 | Source-checked |
| 129 | o3-proOpenAI | 2025-06-10 | 33.3% | 2026-09-01 | Source-checked |
| 130 | GLM-5 (Non-reasoning)Zhipu | 2026-02-11 | 33.2% | 2026-09-01 | Source-checked |
| 131 | DeepSeek V3.2 (Reasoning)DeepSeek | 2025-12-01 | 32.8% | 2026-09-01 | Source-checked |
| 132 | Qwen3.5 122B A10B (Reasoning)Alibaba | 2026-02-24 | 32.8% | 2026-09-01 | Source-checked |
| 133 | Qwen3.5 397B A17B (Non-reasoning)Alibaba | 2026-02-16 | 32.7% | 2026-09-01 | Source-checked |
| 134 | Qwen3 Max ThinkingAlibaba | 2026-01-26 | 32.5% | 2026-09-01 | Source-checked |
| 135 | MiniMax-M2.1MiniMax | 2025-12-23 | 32.1% | 2026-09-01 | Source-checked |
| 136 | Qwen3.6 35B A3B (Reasoning)Alibaba | 2026-04-16 | 32.1% | 2026-09-01 | Source-checked |
| 137 | DeepSeek V4 Pro (Non-reasoning)DeepSeek | 2026-04-24 | 31.9% | 2026-09-01 | Source-checked |
| 138 | GPT-5 (low)OpenAI | 2025-08-07 | 31.9% | 2026-09-01 | Source-checked |
| 139 | MiMo-V2-Flash (Reasoning)Xiaomi | 2025-12-16 | 31.9% | 2026-09-01 | Source-checked |
| 140 | Claude 4 Opus (Reasoning)Anthropic | 2025-05-22 | 31.7% | 2026-09-01 | Source-checked |
| 141 | Ring-2.6-1TOther | 2026-05-08 | 31.7% | 2026-09-01 | Source-checked |
| 142 | Qwen3.6 27B (Non-reasoning)Alibaba | 2026-04-22 | 31.3% | 2026-09-01 | Source-checked |
| 143 | OpenAI o3OpenAI | 2025-04-16 | 31.1% | 2026-09-01 | Source-checked |
| 144 | K-EXAONE 2.0 0803Other | 2026-08-12 | 31% | 2026-09-01 | Source-checked |
| 145 | Step 3.7 FlashStepFun | 2026-06-01 | 30.9% | 2026-09-01 | Source-checked |
| 146 | Mistral Medium 3.5Mistral | 2026-04-29 | 30.4% | 2026-09-01 | Source-checked |
| 147 | Gemma 4 31B ITGoogle | 2026-03-01 | 29.7% | 2026-09-01 | Source-checked |
| 148 | DeepSeek V4 Flash (Non-reasoning)DeepSeek | 2026-04-24 | 29.3% | 2026-09-01 | Source-checked |
| 149 | GPT-5.5 Instant (June 2026)OpenAI | 2026-06-25 | 29.2% | 2026-09-01 | Source-checked |
| 150 | JT-35B-FlashOther | 2026-05-14 | 29% | 2026-09-01 | Source-checked |
| 151 | MiMo-V2.5-Pro (Non-reasoning)Xiaomi | 2026-04-22 | 28.4% | 2026-09-01 | Source-checked |
| 152 | Gemini 3 Flash PreviewGoogle | 2025-12-17 | 27.9% | 2026-09-01 | Source-checked |
| 153 | Hy3-preview (Non-reasoning)Other | 2026-04-23 | 26.6% | 2026-09-01 | Source-checked |
| 154 | Ling-2.6-1TOther | 2026-04-23 | 26.6% | 2026-09-01 | Source-checked |
| 155 | Gemini 3.1 Flash LiteGoogle | 2026-05-07 | 25.6% | 2026-09-01 | Source-checked |
| 156 | Grok 4.3 (Non-reasoning)xAI | 2026-04-30 | 25% | 2026-09-01 | Source-checked |
| 157 | Qwen3.6 35B A3B (Non-reasoning)Alibaba | 2026-04-16 | 24.6% | 2026-09-01 | Source-checked |
| 158 | Ling 3.0 TinyOther | 2026-08-06 | 24.5% | 2026-09-01 | Source-checked |
| 159 | OpenAI o1OpenAI | 2024-09-12 | 23.9% | 2026-09-01 | Source-checked |
| 160 | Granite 4.2 30BOther | 2026-08-25 | 23.7% | 2026-09-01 | Source-checked |
| 161 | Nemotron 3.5 LightningNVIDIA | 2026-08-11 | 23.6% | 2026-09-01 | Source-checked |
| 162 | Command ACohere | 2026-03-15 | 22.8% | 2026-09-01 | Source-checked |
| 163 | Gemma 4 12B (Reasoning)Google | 2026-06-07 | 22.2% | 2026-09-01 | Source-checked |
| 164 | Grok 4.20 0309 v2 (Non-reasoning)xAI | 2026-04-07 | 22.2% | 2026-09-01 | Source-checked |
| 165 | EXAONE 4.5 33BOther | 2026-04-09 | 20.5% | 2026-09-01 | Source-checked |
| 166 | DeepSeek R1DeepSeek | 2025-01-20 | 20.4% | 2026-09-01 | Source-checked |
| 167 | North Mini CodeCohere | 2026-06-09 | 20.2% | 2026-09-01 | Source-checked |
| 168 | Granite 4.2 8BOther | 2026-08-25 | 19.6% | 2026-09-01 | Source-checked |
| 169 | JT-MINIOther | 2026-04-15 | 18.8% | 2026-09-01 | Source-checked |
| 170 | HyperNova 60B 2605Multiverse | 2026-05-06 | 18.3% | 2026-09-01 | Source-checked |
| 171 | G9v3-3BOther | 2026-07-23 | 16.2% | 2026-09-01 | Source-checked |
| 172 | Mistral Large 3Mistral | 2025-12-02 | 15.9% | 2026-09-01 | Source-checked |
| 173 | Nemotron 3 Nano Omni 30B A3B ReasoningNVIDIA | 2026-04-29 | 15% | 2026-09-01 | Source-checked |
| 174 | Llama 4 MaverickMeta | 2026-04-05 | 14.5% | 2026-09-01 | Source-checked |
| 175 | Granite 4.2 3BOther | 2026-08-25 | 14.3% | 2026-09-01 | Source-checked |
| 176 | Ling 2.6 FlashOther | 2026-04-21 | 14.2% | 2026-09-01 | Source-checked |
| 177 | DiffusionGemma 26B A4BGoogle | 2026-06-10 | 13.5% | 2026-09-01 | Source-checked |
| 178 | Gemma 4 12B (Non-reasoning)Google | 2026-06-10 | 13.2% | 2026-09-01 | Source-checked |
| 179 | MiniCPM5-1B (Reasoning)OpenBMB | 2026-06-04 | 11.9% | 2026-09-01 | Source-checked |
| 180 | Claude 3 OpusAnthropic | 2024-03-04 | 11.8% | 2026-09-01 | Source-checked |
| 181 | MiniCPM5-1B (Non-reasoning)OpenBMB | 2026-05-25 | 11.7% | 2026-09-01 | Source-checked |
| 182 | Llama 4 ScoutMeta | 2026-04-05 | 10.3% | 2026-09-01 | Source-checked |
| 183 | Granite 4.1 30BOther | 2026-04-29 | 8.7% | 2026-09-01 | Source-checked |
| 184 | LFM2.5-8B-A1BLiquid | 2026-06-08 | 8.1% | 2026-09-01 | Source-checked |
| 185 | DeepSeek Coder V2DeepSeek | 2024-05-15 | 4.7% | 2026-09-01 | Source-checked |
| 186 | Gemini 1.0 UltraGoogle | 2024-02-08 | 4.3% | 2026-09-01 | Source-checked |
| 187 | MiniCPM-V 4.6 1.3BOpenBMB | 2026-05-11 | 3.8% | 2026-09-01 | Source-checked |
Frontier rows dated 2026-09-05 were read from public AA model pages on Intelligence Index v4.2. The committed API snapshot is still 2026-09-01 (v4.1). AA-Briefcase and GDP.pdf do not yet have standalone VerdictPal explainers — they lack five independently extracted model rows.
v4.2 added AA-Briefcase (15%) and GDP.pdf (10%) and removed GPQA Diamond. Do not compare v4.2 headlines to v4.1 or pre-2026 AA index versions. Sept 1 API-snapshot rows in the ledger are v4.1; hand-updated frontier rows dated 5 Sep 2026 are v4.2.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
AA Intelligence Index v4.2 — weighted composite across four categories: Agents 30% (AA-Briefcase 15%, GDPval-AA v2 10%, τ³-Banking 5%), Coding 20% (Terminal-Bench 2.1 10%, SciCode 10%), Scientific Reasoning 20% (HLE 10%, CritPt 10%), General 30% (AA-Omniscience Accuracy 10% + Non-Hallucination 5%, GDP.pdf 10%, AA-LCR v1.1 5%).
A strong Artificial Analysis Intelligence Index result says nothing about:
Weighted percentage across four categories (Agents 30%, Coding 20%, Scientific Reasoning 20%, General 30%). AA estimates a 95% CI under ±1% on the composite for models with enough repeats. Task format: Composite of ten versioned sub-evaluations with public category weights. Methodology page lists per-eval question counts, repeats, and scoring. AA-Briefcase and GDP.pdf are new in v4.2.
Artificial Analysis Intelligence Index is currently marked Active in the atlas.
Contamination risk for Artificial Analysis Intelligence Index is graded Medium contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.