Benchmarks / Long context

Artificial Analysis Long Context Reasoning

AA-LCR

Long-context synthesis and reasoning: 100 open-answer questions requiring models to integrate evidence across long inputs. 6% weight in AA Intelligence Index v4.1.

What this does not measure
  • Retrieval from external databases — the context is in-prompt only.
  • Multimodal long documents — AA-LCR is text-only in the current AA harness.
  • Infinite context — the eval uses a fixed long-context window, not arbitrary document length.
Analysis

Why this benchmark is useful

Long-context research breaks when models lose thread across distant evidence. AA-LCR is one of the few public suites that grades synthesis across long inputs instead of retrieval trivia alone.

Scope

Coverage map

Task family
Long context
Format
100 open-answer questions with long context passages; equality-checker LLM grading, pass@1.
Scoring
pass@1 accuracy with 3 repeats.
Maintainer
Artificial Analysis
Reading guide

How to read the scores

Read pass@1 with three repeats as a conservative headline. Compare only rows from the same AA harness version and long-context window.

Blind spots

What it does not cover

  • Retrieval from external databases — the context is in-prompt only.
  • Multimodal long documents — AA-LCR is text-only in the current AA harness.
  • Infinite context — the eval uses a fixed long-context window, not arbitrary document length.
Scores

Evidence ledger

179 rows
179179 rows
83.33%best score
179source-checked
1sources
2026-09-01to 2026-06-17
179

api · api

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Muse Spark 1.283.33% ·
  2. Kimi K382.67% ·
  3. Muse Spark 1.181.33% ·
  4. Gemini 3.5 Flash81% ·
  5. MiniMax M380.33% ·

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Muse Spark 1.2Meta2026-08-0583.33%
2026-09-01
Source-checked
2Kimi K3Moonshot2026-07-1682.67%
2026-09-01
Source-checked
3Muse Spark 1.1Meta2026-07-0981.33%
2026-09-01
Source-checked
4Gemini 3.5 FlashGoogle2026-05-1981%
2026-09-01
Source-checked
5MiniMax M3MiniMax2026-05-3180.33%
2026-09-01
Source-checked
6Gemini 3.7 FlashGoogle2026-08-1380%
2026-09-01
Source-checked
7Muse GlimmerMeta2026-08-1080%
2026-09-01
Source-checked
8Gemini 3.5 Flash (medium)Google2026-05-1979.67%
2026-09-01
Source-checked
9GPT-5.6 TerraOpenAI2026-07-0979.67%
2026-09-01
Source-checked
10GPT-5.2OpenAI2025-12-1179.33%
2026-09-01
Source-checked
11GPT-5.2 Codex (xhigh)OpenAI2025-12-1179.33%
2026-09-01
Source-checked
12Gemini 3.1 Pro PreviewGoogle2026-02-1979%
2026-09-01
Source-checked
13Gemini 3.6 Flash (high)Google2026-07-2179%
2026-09-01
Source-checked
14GPT-5.5 (high)OpenAI2026-04-2379%
2026-09-01
Source-checked
15Claude Opus 5 (Adaptive Reasoning, Medium Effort)Anthropic2026-07-2478.67%
2026-09-01
Source-checked
16GPT-5.3 Codex (xhigh)OpenAI2026-02-0578.33%
2026-09-01
Source-checked
17GPT-5.6 LunaOpenAI2026-07-0978.33%
2026-09-01
Source-checked
18Agnes 2.5 Pro BetaOther2026-08-2678%
2026-09-01
Source-checked
19DeepSeek V4 Flash Vision (Reasoning, Max Effort)DeepSeek2026-08-2178%
2026-09-01
Source-checked
20GLM-5.3-FlashZhipu2026-08-2678%
2026-09-01
Source-checked
21GPT 5.4OpenAI2026-03-0577.67%
2026-09-01
Source-checked
22GPT-5.5 (low)OpenAI2026-04-2377.67%
2026-09-01
Source-checked
23GPT-5.6 SolOpenAI2026-07-0977.67%
2026-09-01
Source-checked
24MiMo-V2.5-ProXiaomi2026-04-2277.67%
2026-09-01
Source-checked
25Qwen3.8 27BAlibaba2026-08-1477.33%
2026-09-01
Source-checked
26Claude Opus 5 (Adaptive Reasoning, Low Effort)Anthropic2026-07-2477%
2026-09-01
Source-checked
27Claude Sonnet 5Anthropic2026-06-3077%
2026-09-01
Source-checked
28Kimi K3 (low)Moonshot2026-07-1677%
2026-09-01
Source-checked
29Muse SparkOther2026-01-0177%
2026-09-01
Source-checked
30Qwen3.8-Flash-NextAlibaba2026-08-2677%
2026-09-01
Source-checked
31Claude Fable 5Anthropic2026-06-0976.67%
2026-09-01
Source-checked
32GLM-5.2Zhipu2026-06-1376.67%
2026-09-01
Source-checked
33GPT-5.1OpenAI2025-11-1376.67%
2026-09-01
Source-checked
34GPT-5.5 (medium)OpenAI2026-04-2376.67%
2026-09-01
Source-checked
35Kimi K2.6Moonshot2026-04-2076.67%
2026-09-01
Source-checked
36Claude Opus 5 (Adaptive Reasoning, High Effort)Anthropic2026-07-2476.33%
2026-09-01
Source-checked
37Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)Anthropic2026-07-2476.33%
2026-09-01
Source-checked
38GLM-5.3Zhipu2026-08-1476.33%
2026-09-01
Source-checked
39GPT-5 (high)OpenAI2025-08-0776.33%
2026-09-01
Source-checked
40GPT-5 (medium)OpenAI2025-08-0776.33%
2026-09-01
Source-checked
41Nex-N2-ProOther2026-06-0276.33%
2026-09-01
Source-checked
42Claude Opus 4.5 (Reasoning)Anthropic2025-11-2476%
2026-09-01
Source-checked
43Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic2026-07-2475.67%
2026-09-01
Source-checked
44Claude Opus 4.7Anthropic2026-04-1675.33%
2026-09-01
Source-checked
45DeepSeek V4 Pro 0813 (Reasoning, Max Effort)DeepSeek2026-08-1375.33%
2026-09-01
Source-checked
46MiniMax M2.7MiniMax2026-04-1575.33%
2026-09-01
Source-checked
47Qwen3.8 2.4T A95BAlibaba2026-08-1275.33%
2026-09-01
Source-checked
48Grok 4.6xAI2026-08-1275%
2026-09-01
Source-checked
49Kimi K2.7 CodeMoonshot2026-06-1275%
2026-09-01
Source-checked
50Apodex 1.1Other2026-08-3074.67%
2026-09-01
Source-checked
51Gemini 3.5 Flash-LiteGoogle2026-07-2174.67%
2026-09-01
Source-checked
52Hy3Other2026-07-0674.67%
2026-09-01
Source-checked
53Qwen3.7 MaxAlibaba2026-05-1974.67%
2026-09-01
Source-checked
54Claude Opus 4.6 (Adaptive Reasoning, Max Effort)Anthropic2026-02-0574.33%
2026-09-01
Source-checked
55DeepSeek V4 Flash 0731 (Reasoning, Max Effort)DeepSeek2026-07-3174.33%
2026-09-01
Source-checked
56Qwen3.8 MaxAlibaba2026-08-0374.33%
2026-09-01
Source-checked
57Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)Anthropic2026-02-1774%
2026-09-01
Source-checked
58Grok 4.5xAI2026-07-0874%
2026-09-01
Source-checked
59GPT-5.4 (low)OpenAI2026-03-0573.67%
2026-09-01
Source-checked
60Claude 4.1 Opus (Reasoning)Anthropic2025-08-0573.33%
2026-09-01
Source-checked
61Inkling (xhigh)Other2026-07-1573.33%
2026-09-01
Source-checked
62OpenAI o3OpenAI2025-04-1673.33%
2026-09-01
Source-checked
63Qwen3.6 27B (Reasoning)Alibaba2026-04-2273.33%
2026-09-01
Source-checked
64Agnes 2.5 Pro AlphaOther2026-07-2473%
2026-09-01
Source-checked
65Gemini 3 Flash Preview (Reasoning)Google2025-12-1773%
2026-09-01
Source-checked
66Gemini 3 ProGoogle2025-11-1873%
2026-09-01
Source-checked
67GPT-5.4 MiniOpenAI2026-03-1773%
2026-09-01
Source-checked
68Kimi K2.5Moonshot2026-01-2773%
2026-09-01
Source-checked
69Qwen3.5 397B A17B (Reasoning)Alibaba2026-02-1672.67%
2026-09-01
Source-checked
70Claude Opus 4.7 (Non-reasoning, High Effort)Anthropic2026-04-1672.33%
2026-09-01
Source-checked
71Motif 3Other2026-08-1272.33%
2026-09-01
Source-checked
72Qwen3.5 27B (Reasoning)Alibaba2026-02-2472.33%
2026-09-01
Source-checked
73Qwen3.6 PlusAlibaba2026-04-0272.33%
2026-09-01
Source-checked
74GPT-5.4 NanoOpenAI2026-03-1772%
2026-09-01
Source-checked
75MiniMax M2.5MiniMax2026-04-0172%
2026-09-01
Source-checked
76Qwen3.6 Max PreviewAlibaba2026-04-2072%
2026-09-01
Source-checked
77MiMo-V2-OmniXiaomi2026-03-1971.67%
2026-09-01
Source-checked
78Gemini 3.1 Flash LiteGoogle2026-05-0771.33%
2026-09-01
Source-checked
79GPT-5 Codex (high)OpenAI2025-09-2371%
2026-09-01
Source-checked
80Nemotron 3 Ultra 550B A55B (Reasoning)NVIDIA2026-06-0471%
2026-09-01
Source-checked
81DeepSeek V3.2 (Reasoning)DeepSeek2025-12-0170.67%
2026-09-01
Source-checked
82GLM-5 (Reasoning)Zhipu2026-02-1170.67%
2026-09-01
Source-checked
83Motif 3 (Beta)Other2026-07-2170.67%
2026-09-01
Source-checked
84Solar Pro 4Other2026-08-0670.67%
2026-09-01
Source-checked
85Gemini 3 Pro Preview (low)Google2025-11-1870.33%
2026-09-01
Source-checked
86JT-4.1 Flash 236B A21BOther2026-07-0970.33%
2026-09-01
Source-checked
87Kimi K2 ThinkingMoonshot2025-11-0670.33%
2026-09-01
Source-checked
88Qwen3 Max ThinkingAlibaba2026-01-2670.33%
2026-09-01
Source-checked
89Qwen3.5 122B A10B (Reasoning)Alibaba2026-02-2470.33%
2026-09-01
Source-checked
90Grok Build 0.1 0616xAI2026-06-1670%
2026-09-01
Source-checked
91KAT Coder Pro V2Other2026-03-2770%
2026-09-01
Source-checked
92MiMo-V2-Omni-0327Xiaomi2026-03-2769.67%
2026-09-01
Source-checked
93Step 3.7 FlashStepFun2026-06-0169.67%
2026-09-01
Source-checked
94Inkling SmallOther2026-07-3069.33%
2026-09-01
Source-checked
95DeepSeek V4 Flash (Reasoning, High Effort)DeepSeek2026-04-2469%
2026-09-01
Source-checked
96GPT-5.1 Codex (high)OpenAI2025-11-1369%
2026-09-01
Source-checked
97GPT-5.2 (medium)OpenAI2025-12-1169%
2026-09-01
Source-checked
98Qwen 3.7 PlusAlibaba2026-06-0269%
2026-09-01
Source-checked
99MiMo-V2-Flash (Feb 2026)Xiaomi2025-12-1668.67%
2026-09-01
Source-checked
100Claude 4.5 Sonnet (Reasoning)Anthropic2025-09-2968.33%
2026-09-01
Source-checked
101Gemma 4 31B ITGoogle2026-03-0168.33%
2026-09-01
Source-checked
102Grok 4.3 (medium)xAI2026-04-3068.33%
2026-09-01
Source-checked
103MiMo-V2.5Xiaomi2026-04-2268.33%
2026-09-01
Source-checked
104Solar Open2 250BOther2026-08-1268.33%
2026-09-01
Source-checked
105GLM-4.7 (Reasoning)Zhipu2025-12-2268%
2026-09-01
Source-checked
106GLM-5.1Zhipu2026-04-0768%
2026-09-01
Source-checked
107MiMo-V2-Flash (Reasoning)Xiaomi2025-12-1668%
2026-09-01
Source-checked
108Grok 4.3 (low)xAI2026-04-3067.67%
2026-09-01
Source-checked
109Claude Opus 4.5Anthropic2025-11-2467.33%
2026-09-01
Source-checked
110Ring-2.6-1TOther2026-05-0867.33%
2026-09-01
Source-checked
111DeepSeek V4 Pro (Reasoning, High Effort)DeepSeek2026-04-2467%
2026-09-01
Source-checked
112GPT-5.5 Instant (June 2026)OpenAI2026-06-2567%
2026-09-01
Source-checked
113Grok 4xAI2025-07-1067%
2026-09-01
Source-checked
114Ling 3.0 FlashOther2026-08-0467%
2026-09-01
Source-checked
115GLM-5-TurboZhipu2026-03-1566.67%
2026-09-01
Source-checked
116Qwen3.6 35B A3B (Reasoning)Alibaba2026-04-1666.67%
2026-09-01
Source-checked
117MiniMax-M2.1MiniMax2025-12-2366.33%
2026-09-01
Source-checked
118A.X-K2Other2026-08-1266%
2026-09-01
Source-checked
119Grok 4.3xAI2026-04-3066%
2026-09-01
Source-checked
120Kimi K2.6 (Non-reasoning)Other2026-04-2066%
2026-09-01
Source-checked
121GLM 5V Turbo (Reasoning)Zhipu2026-04-0165.67%
2026-09-01
Source-checked
122MiMo-V2-ProXiaomi2026-03-1865.67%
2026-09-01
Source-checked
123Mistral Medium 3.5Mistral2026-04-2965.33%
2026-09-01
Source-checked
124Claude Sonnet 4.6 (Non-reasoning, Low Effort)Anthropic2026-02-1764.67%
2026-09-01
Source-checked
125Qwen3.6 27B (Non-reasoning)Alibaba2026-04-2264.33%
2026-09-01
Source-checked
126GPT-5.5 Instant (May 2026)OpenAI2026-05-0564%
2026-09-01
Source-checked
127GPT-5 (low)OpenAI2025-08-0763.67%
2026-09-01
Source-checked
128JT-35B-FlashOther2026-05-1463.67%
2026-09-01
Source-checked
129OpenAI o1OpenAI2024-09-1263.33%
2026-09-01
Source-checked
130LongCat-2.0Meituan2026-06-2962.67%
2026-09-01
Source-checked
131Claude Opus 4.6Anthropic2026-02-0562.33%
2026-09-01
Source-checked
132Claude Sonnet 4.6Anthropic2026-02-1762.33%
2026-09-01
Source-checked
133Grok 4.20 ReasoningxAI2026-03-0562.33%
2026-09-01
Source-checked
134G9v3-39A5BOther2026-08-0362%
2026-09-01
Source-checked
135Qwen3.5 397B A17B (Non-reasoning)Alibaba2026-02-1662%
2026-09-01
Source-checked
136Gemma 4 12B (Reasoning)Google2026-06-0761.67%
2026-09-01
Source-checked
137Grok 4.20 0309 (Reasoning)xAI2026-03-1060.67%
2026-09-01
Source-checked
138Hy3-preview (Reasoning)Other2026-04-2360.33%
2026-09-01
Source-checked
139Qwen3.6 35B A3B (Non-reasoning)Alibaba2026-04-1660%
2026-09-01
Source-checked
140Ling 3.0 TinyOther2026-08-0658.67%
2026-09-01
Source-checked
141Gemini 3.5 Flash (minimal)Google2026-05-1958.33%
2026-09-01
Source-checked
1Claude Opus 4.8Anthropic2026-05-2858.3%
2026-06-17
Source-checked
143GPT-5.5 (Non-reasoning)OpenAI2026-04-2358%
2026-09-01
Source-checked
144K-EXAONE 2.0 0803Other2026-08-1257.67%
2026-09-01
Source-checked
145DeepSeek R1DeepSeek2025-01-2056.67%
2026-09-01
Source-checked
146Nemotron 3.5 LightningNVIDIA2026-08-1155.33%
2026-09-01
Source-checked
147Gemini 3 Flash PreviewGoogle2025-12-1753%
2026-09-01
Source-checked
2GPT 5.5OpenAI2026-04-2352.1%
2026-06-17
Source-checked
149EXAONE 4.5 33BOther2026-04-0951.67%
2026-09-01
Source-checked
150Llama 4 MaverickMeta2026-04-0550%
2026-09-01
Source-checked
151DeepSeek V4 Pro (Non-reasoning)DeepSeek2026-04-2449.67%
2026-09-01
Source-checked
152Command ACohere2026-03-1548.67%
2026-09-01
Source-checked
153Granite 4.2 30BOther2026-08-2546.67%
2026-09-01
Source-checked
154GLM-5.2 (Non-reasoning)Zhipu2026-06-1645%
2026-09-01
Source-checked
155GLM-5 (Non-reasoning)Zhipu2026-02-1143.33%
2026-09-01
Source-checked
156Granite 4.2 8BOther2026-08-2543.33%
2026-09-01
Source-checked
157G9v3-3BOther2026-07-2340.67%
2026-09-01
Source-checked
158Nemotron 3 Nano Omni 30B A3B ReasoningNVIDIA2026-04-2940.67%
2026-09-01
Source-checked
159MiMo-V2.5-Pro (Non-reasoning)Xiaomi2026-04-2239.33%
2026-09-01
Source-checked
160Hy3-preview (Non-reasoning)Other2026-04-2338.67%
2026-09-01
Source-checked
161Ling-2.6-1TOther2026-04-2338%
2026-09-01
Source-checked
162DeepSeek V4 Flash (Non-reasoning)DeepSeek2026-04-2437.33%
2026-09-01
Source-checked
163Claude 4 Opus (Reasoning)Anthropic2025-05-2236.33%
2026-09-01
Source-checked
164North Mini CodeCohere2026-06-0936%
2026-09-01
Source-checked
165HyperNova 60B 2605Multiverse2026-05-0635.33%
2026-09-01
Source-checked
166Mistral Large 3Mistral2025-12-0234.67%
2026-09-01
Source-checked
167Gemma 4 12B (Non-reasoning)Google2026-06-1031.33%
2026-09-01
Source-checked
168Llama 4 ScoutMeta2026-04-0530.33%
2026-09-01
Source-checked
169Grok 4.3 (Non-reasoning)xAI2026-04-3028%
2026-09-01
Source-checked
170Ling 2.6 FlashOther2026-04-2128%
2026-09-01
Source-checked
171Granite 4.2 3BOther2026-08-2524.33%
2026-09-01
Source-checked
172Grok 4.20 0309 v2 (Non-reasoning)xAI2026-04-0720.67%
2026-09-01
Source-checked
173Granite 4.1 30BOther2026-04-2920.33%
2026-09-01
Source-checked
174DiffusionGemma 26B A4BGoogle2026-06-1018.33%
2026-09-01
Source-checked
175JT-MINIOther2026-04-1512.67%
2026-09-01
Source-checked
176MiniCPM-V 4.6 1.3BOpenBMB2026-05-116.67%
2026-09-01
Source-checked
177MiniCPM5-1B (Non-reasoning)OpenBMB2026-05-254.67%
2026-09-01
Source-checked
178MiniCPM5-1B (Reasoning)OpenBMB2026-06-044.67%
2026-09-01
Source-checked
179LFM2.5-8B-A1BLiquid2026-06-080%
2026-09-01
Source-checked
Method

What it covers

This benchmark sits in the long context family. It uses 100 open-answer questions with long context passages; equality-checker LLM grading, pass@1. The score should travel with its task format, scoring method, source date, and benchmark version.

Score ceiling

Where it breaks down

High scores on in-prompt long context do not prove reliable retrieval from external corpora or multimodal documents.

Receipts

Sources and further reading

Questions about this benchmark

Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.

What does AA-LCR measure?

Long-context synthesis and reasoning: 100 open-answer questions requiring models to integrate evidence across long inputs. 6% weight in AA Intelligence Index v4.1.

What does a high AA-LCR score not prove?

A strong AA-LCR result says nothing about:

  • Retrieval from external databases — the context is in-prompt only.
  • Multimodal long documents — AA-LCR is text-only in the current AA harness.
  • Infinite context — the eval uses a fixed long-context window, not arbitrary document length.

How is AA-LCR scored?

pass@1 accuracy with 3 repeats. Task format: 100 open-answer questions with long context passages; equality-checker LLM grading, pass@1.

Is AA-LCR saturated?

AA-LCR is currently marked Active in the atlas.

Can AA-LCR results be contaminated by training data?

Contamination risk for AA-LCR is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.