Why this benchmark is useful
Single-number model rankings hide tradeoffs. The AA Intelligence Index is a weighted composite of ten public sub-evals — useful as a dated frontier snapshot when you read the slices, not as a universal quality score.
Benchmarks / Reasoning
AA Intelligence Index
AA Intelligence Index v4.1 — weighted composite emphasizing agentic workloads: GDPval-AA v2 (20%), Terminal-Bench 2.1 (16%), τ³-Bench Banking (14%), Humanity's Last Exam (12%), AA-Omniscience (12%), SciCode (8%), GPQA Diamond (6%), AA-LCR (6%), CritPt (6%). IFBench removed for saturation.
Single-number model rankings hide tradeoffs. The AA Intelligence Index is a weighted composite of ten public sub-evals — useful as a dated frontier snapshot when you read the slices, not as a universal quality score.
Open the methodology page for v4.1 weights, then drill into each sub-eval card. A high composite can mask weak agentic, coding, or long-context slices.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Claude Fable 5Anthropic | 2026-06-09 | 59.9% | 2026-07-21 | Source-checked |
| 2 | GPT-5.6 SolOpenAI | 2026-07-09 | 58.9% | 2026-07-21 | Source-checked |
| 3 | Kimi K3Moonshot | 2026-07-16 | 57.1% | 2026-07-21 | Source-checked |
| 4 | Claude Opus 4.8Anthropic | 2026-05-28 | 55.7% | 2026-07-21 | Source-checked |
| 5 | GPT-5.6 TerraOpenAI | 2026-07-09 | 55% | 2026-07-21 | Source-checked |
| 6 | GPT 5.5OpenAI | 2026-04-23 | 54.8% | 2026-07-21 | Source-checked |
| 7 | Grok 4.5xAI | 2026-07-08 | 53.8% | 2026-07-21 | Source-checked |
| 8 | Claude Opus 4.7Anthropic | 2026-04-16 | 53.5% | 2026-07-21 | Source-checked |
| 9 | Claude Sonnet 5Anthropic | 2026-06-30 | 53.4% | 2026-07-21 | Source-checked |
| 10 | GPT-5.5 (high)OpenAI | 2026-04-23 | 53.1% | 2026-07-21 | Source-checked |
| 11 | GPT 5.4OpenAI | 2026-03-05 | 51.4% | 2026-07-21 | Source-checked |
| 12 | GPT-5.6 LunaOpenAI | 2026-07-09 | 51.2% | 2026-07-21 | Source-checked |
| 13 | GLM-5.2Zhipu | 2026-06-13 | 51.1% | 2026-07-21 | Source-checked |
| 14 | Muse Spark 1.1Meta | 2026-07-09 | 50.6% | 2026-07-21 | Source-checked |
| 15 | GPT-5.5 (medium)OpenAI | 2026-04-23 | 50.4% | 2026-07-21 | Source-checked |
| 16 | Gemini 3.5 FlashGoogle | 2026-05-19 | 50.2% | 2026-07-21 | Source-checked |
| 17 | Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)Anthropic | 2026-02-17 | 47.2% | 2026-07-21 | Source-checked |
| 18 | Gemini 3.1 Pro PreviewGoogle | 2026-02-19 | 46.5% | 2026-07-21 | Source-checked |
| 19 | Qwen3.7 MaxAlibaba | 2026-05-19 | 46% | 2026-07-21 | Source-checked |
| 20 | Gemini 3.5 Flash (medium)Google | 2026-05-19 | 45.4% | 2026-07-21 | Source-checked |
| 21 | MiniMax M3MiniMax | 2026-05-31 | 44.4% | 2026-07-21 | Source-checked |
| 22 | DeepSeek V4 Pro (Reasoning, Max Effort)DeepSeek | 2026-04-24 | 44.3% | 2026-07-21 | Source-checked |
| 23 | GPT-5.3 Codex (xhigh)OpenAI | 2026-02-05 | 44.3% | 2026-07-21 | Source-checked |
| 24 | Kimi K2.6Moonshot | 2026-04-20 | 44.2% | 2026-07-21 | Source-checked |
| 25 | Claude Opus 4.6 (Adaptive Reasoning, Max Effort)Anthropic | 2026-02-05 | 43.7% | 2026-07-21 | Source-checked |
| 26 | GPT-5.5 (low)OpenAI | 2026-04-23 | 43.5% | 2026-07-21 | Source-checked |
| 27 | Muse SparkOther | 2026-01-01 | 43.1% | 2026-07-21 | Source-checked |
| 28 | Claude Opus 4.7 (Non-reasoning, High Effort)Anthropic | 2026-04-16 | 42.7% | 2026-07-21 | Source-checked |
| 29 | GPT-5.2OpenAI | 2025-12-11 | 42.2% | 2026-07-21 | Source-checked |
| 30 | MiMo-V2.5-ProXiaomi | 2026-04-22 | 42.2% | 2026-07-21 | Source-checked |
| 31 | Kimi K2.7 CodeMoonshot | 2026-06-12 | 41.9% | 2026-07-21 | Source-checked |
| 32 | GLM-5.1Zhipu | 2026-04-07 | 40.2% | 2026-07-21 | Source-checked |
| 33 | GPT-5.4 MiniOpenAI | 2026-03-17 | 40% | 2026-07-21 | Source-checked |
| 34 | Gemini 3 ProGoogle | 2025-11-18 | 39.6% | 2026-07-21 | Source-checked |
| 35 | Qwen 3.7 PlusAlibaba | 2026-06-02 | 39% | 2026-07-21 | Source-checked |
| 36 | GPT-5.4 NanoOpenAI | 2026-03-17 | 38.2% | 2026-07-21 | Source-checked |
| 37 | MiniMax M2.7MiniMax | 2026-04-15 | 38.1% | 2026-07-21 | Source-checked |
| 38 | Claude Opus 4.6Anthropic | 2026-02-05 | 37.8% | 2026-07-21 | Source-checked |
| 39 | Nemotron 3 Ultra 550B A55B (Reasoning)NVIDIA | 2026-06-04 | 37.8% | 2026-07-21 | Source-checked |
| 40 | Grok 4.3xAI | 2026-04-30 | 37.6% | 2026-07-21 | Source-checked |
| 41 | Grok 4.20 ReasoningxAI | 2026-03-05 | 37% | 2026-07-21 | Source-checked |
| 42 | GPT-5.1OpenAI | 2025-11-13 | 36.9% | 2026-07-21 | Source-checked |
| 43 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 35.9% | 2026-07-21 | Source-checked |
| 44 | Kimi K2.5Moonshot | 2026-01-27 | 35.4% | 2026-07-21 | Source-checked |
| 45 | Claude Opus 4.5Anthropic | 2025-11-24 | 34.7% | 2026-07-21 | Source-checked |
| 46 | MiniMax M2.5MiniMax | 2026-04-01 | 33.7% | 2026-07-21 | Source-checked |
| 47 | LongCat-2.0Meituan | 2026-06-29 | 33.5% | 2026-07-21 | Source-checked |
| 48 | OpenAI o3OpenAI | 2025-04-16 | 30.4% | 2026-07-21 | Source-checked |
| 49 | Step 3.7 FlashStepFun | 2026-06-01 | 30.3% | 2026-07-21 | Source-checked |
| 50 | Gemma 4 31B ITGoogle | 2026-03-01 | 29.4% | 2026-07-21 | Source-checked |
| 51 | Gemini 3 Flash PreviewGoogle | 2025-12-17 | 27.4% | 2026-07-21 | Source-checked |
| 52 | Gemini 3.1 Flash LiteGoogle | 2026-05-07 | 25% | 2026-07-21 | Source-checked |
| 53 | OpenAI o1OpenAI | 2024-09-12 | 23.4% | 2026-07-21 | Source-checked |
| 54 | Gemma 4 12B (Reasoning)Google | 2026-06-07 | 22% | 2026-07-21 | Source-checked |
| 55 | DeepSeek R1DeepSeek | 2025-01-20 | 20.1% | 2026-07-21 | Source-checked |
| 56 | North Mini CodeCohere | 2026-06-09 | 19.8% | 2026-07-21 | Source-checked |
| 57 | HyperNova 60B 2605Multiverse | 2026-05-06 | 17.8% | 2026-07-21 | Source-checked |
| 58 | Mistral Large 3Mistral | 2025-12-02 | 15.9% | 2026-07-21 | Source-checked |
| 59 | Llama 4 MaverickMeta | 2026-04-05 | 14.3% | 2026-07-21 | Source-checked |
| 60 | Gemma 4 12B (Non-reasoning)Google | 2026-06-10 | 13.2% | 2026-07-21 | Source-checked |
| 61 | MiniCPM5-1B (Reasoning)OpenBMB | 2026-06-04 | 12% | 2026-07-21 | Source-checked |
| 62 | Claude 3 OpusAnthropic | 2024-03-04 | 11.8% | 2026-07-21 | Source-checked |
| 63 | Llama 4 ScoutMeta | 2026-04-05 | 10% | 2026-07-21 | Source-checked |
| 64 | LFM2.5-8B-A1BLiquid | 2026-06-08 | 8.3% | 2026-07-21 | Source-checked |
| 65 | Command ACohere | 2026-03-15 | 7.7% | 2026-07-21 | Source-checked |
| 66 | DeepSeek Coder V2DeepSeek | 2024-05-15 | 5.1% | 2026-07-21 | Source-checked |
| 67 | Gemini 1.0 UltraGoogle | 2024-02-08 | 4.6% | 2026-07-21 | Source-checked |
Scores in the ledger are snapshotted from artificialanalysis.ai/models with source date and ingest timestamp on each row.
v4.1 upgraded GDPval, Terminal-Bench, and τ-bench; removed IFBench for saturation. Do not compare v4.1 headline scores to pre-2026 AA index versions.
The Pack · Editorial newsletter
One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.