Benchmarks / Long context

Artificial Analysis Long Context Reasoning

AA-LCR

Long-context synthesis and reasoning: 100 open-answer questions requiring models to integrate evidence across long inputs. 6% weight in AA Intelligence Index v4.1.

Artificial AnalysisLong contextSolidActiveLow contamination riskSince 2025
What this does not measure
  • Retrieval from external databases — the context is in-prompt only.
  • Multimodal long documents — AA-LCR is text-only in the current AA harness.
  • Infinite context — the eval uses a fixed long-context window, not arbitrary document length.
Analysis

Why this benchmark is useful

Long-context research breaks when models lose thread across distant evidence. AA-LCR is one of the few public suites that grades synthesis across long inputs instead of retrieval trivia alone.

Scope

Coverage map

Task family
Long context
Format
100 open-answer questions with long context passages; equality-checker LLM grading, pass@1.
Scoring
pass@1 accuracy with 3 repeats.
Maintainer
Artificial Analysis
Reading guide

How to read the scores

Read pass@1 with three repeats as a conservative headline. Compare only rows from the same AA harness version and long-context window.

Blind spots

What it does not cover

  • Retrieval from external databases — the context is in-prompt only.
  • Multimodal long documents — AA-LCR is text-only in the current AA harness.
  • Infinite context — the eval uses a fixed long-context window, not arbitrary document length.
Scores

Evidence ledger

64 rows
6464 rows
75%best score
64source-checked
1sources
2026-07-21to 2026-06-17
64

api · api

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. GPT-5.175% ·
  2. Kimi K374.67% ·
  3. GPT 5.474% ·
  4. GPT-5.3 Codex (xhigh)74% ·
  5. GPT-5.6 Luna74% ·

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1GPT-5.1OpenAI2025-11-1375%
2026-07-21
Source-checked
2Kimi K3Moonshot2026-07-1674.67%
2026-07-21
Source-checked
3GPT 5.4OpenAI2026-03-0574%
2026-07-21
Source-checked
4GPT-5.3 Codex (xhigh)OpenAI2026-02-0574%
2026-07-21
Source-checked
5GPT-5.6 LunaOpenAI2026-07-0974%
2026-07-21
Source-checked
6GPT-5.6 TerraOpenAI2026-07-0974%
2026-07-21
Source-checked
7MiniMax M3MiniMax2026-05-3174%
2026-07-21
Source-checked
8GPT-5.6 SolOpenAI2026-07-0973.67%
2026-07-21
Source-checked
9GPT-5.5 (high)OpenAI2026-04-2373.33%
2026-07-21
Source-checked
10MiMo-V2.5-ProXiaomi2026-04-2273.33%
2026-07-21
Source-checked
11Gemini 3.1 Pro PreviewGoogle2026-02-1972.67%
2026-07-21
Source-checked
12GPT-5.2OpenAI2025-12-1172.67%
2026-07-21
Source-checked
13GPT-5.5 (medium)OpenAI2026-04-2372.33%
2026-07-21
Source-checked
14GPT-5.5 (low)OpenAI2026-04-2372%
2026-07-21
Source-checked
15GLM-5.2Zhipu2026-06-1371.33%
2026-07-21
Source-checked
16Gemini 3.5 Flash (medium)Google2026-05-1971%
2026-07-21
Source-checked
17Claude Opus 4.6 (Adaptive Reasoning, Max Effort)Anthropic2026-02-0570.67%
2026-07-21
Source-checked
18Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)Anthropic2026-02-1770.67%
2026-07-21
Source-checked
19Claude Sonnet 5Anthropic2026-06-3070.67%
2026-07-21
Source-checked
20Gemini 3 ProGoogle2025-11-1870.67%
2026-07-21
Source-checked
21Claude Opus 4.7Anthropic2026-04-1670.33%
2026-07-21
Source-checked
22Claude Fable 5Anthropic2026-06-0970%
2026-07-21
Source-checked
23Kimi K2.6Moonshot2026-04-2069.67%
2026-07-21
Source-checked
24Muse SparkOther2026-01-0169.67%
2026-07-21
Source-checked
25Gemini 3.5 FlashGoogle2026-05-1969.33%
2026-07-21
Source-checked
26GPT-5.4 MiniOpenAI2026-03-1769.33%
2026-07-21
Source-checked
27OpenAI o3OpenAI2025-04-1669.33%
2026-07-21
Source-checked
28Qwen3.7 MaxAlibaba2026-05-1969%
2026-07-21
Source-checked
29MiniMax M2.7MiniMax2026-04-1568.67%
2026-07-21
Source-checked
30Grok 4.5xAI2026-07-0867.67%
2026-07-21
Source-checked
31Claude Opus 4.7 (Non-reasoning, High Effort)Anthropic2026-04-1667%
2026-07-21
Source-checked
32Nemotron 3 Ultra 550B A55B (Reasoning)NVIDIA2026-06-0467%
2026-07-21
Source-checked
33DeepSeek V4 Pro (Reasoning, Max Effort)DeepSeek2026-04-2466.33%
2026-07-21
Source-checked
34Kimi K2.7 CodeMoonshot2026-06-1266.33%
2026-07-21
Source-checked
35GPT-5.4 NanoOpenAI2026-03-1766%
2026-07-21
Source-checked
36MiniMax M2.5MiniMax2026-04-0166%
2026-07-21
Source-checked
37Claude Opus 4.5Anthropic2025-11-2465.33%
2026-07-21
Source-checked
38Gemini 3.1 Flash LiteGoogle2026-05-0765.33%
2026-07-21
Source-checked
39Kimi K2.5Moonshot2026-01-2765.33%
2026-07-21
Source-checked
40Qwen 3.7 PlusAlibaba2026-06-0265%
2026-07-21
Source-checked
41Grok 4.3xAI2026-04-3064.33%
2026-07-21
Source-checked
42Step 3.7 FlashStepFun2026-06-0163.67%
2026-07-21
Source-checked
43Muse Spark 1.1Meta2026-07-0963.33%
2026-07-21
Source-checked
44GLM-5.1Zhipu2026-04-0762.33%
2026-07-21
Source-checked
45Gemma 4 31B ITGoogle2026-03-0162%
2026-07-21
Source-checked
46OpenAI o1OpenAI2024-09-1259.33%
2026-07-21
Source-checked
47Claude Opus 4.6Anthropic2026-02-0558.33%
2026-07-21
Source-checked
1Claude Opus 4.8Anthropic2026-05-2858.3%
2026-06-17
Source-checked
49Grok 4.20 ReasoningxAI2026-03-0558%
2026-07-21
Source-checked
50LongCat-2.0Meituan2026-06-2958%
2026-07-21
Source-checked
51Claude Sonnet 4.6Anthropic2026-02-1757.67%
2026-07-21
Source-checked
52Gemma 4 12B (Reasoning)Google2026-06-0755.33%
2026-07-21
Source-checked
53DeepSeek R1DeepSeek2025-01-2054.67%
2026-07-21
Source-checked
2GPT 5.5OpenAI2026-04-2352.1%
2026-06-17
Source-checked
55Gemini 3 Flash PreviewGoogle2025-12-1748%
2026-07-21
Source-checked
56Llama 4 MaverickMeta2026-04-0546%
2026-07-21
Source-checked
57Mistral Large 3Mistral2025-12-0234.67%
2026-07-21
Source-checked
58North Mini CodeCohere2026-06-0932.33%
2026-07-21
Source-checked
59HyperNova 60B 2605Multiverse2026-05-0631.67%
2026-07-21
Source-checked
60Gemma 4 12B (Non-reasoning)Google2026-06-1030.67%
2026-07-21
Source-checked
61Llama 4 ScoutMeta2026-04-0525.8%
2026-07-21
Source-checked
62Command ACohere2026-03-1518%
2026-07-21
Source-checked
63MiniCPM5-1B (Reasoning)OpenBMB2026-06-043.67%
2026-07-21
Source-checked
64LFM2.5-8B-A1BLiquid2026-06-080%
2026-07-21
Source-checked
Method

What it covers

This benchmark sits in the long context family. It uses 100 open-answer questions with long context passages; equality-checker LLM grading, pass@1. The score should travel with its task format, scoring method, source date, and benchmark version.

Score ceiling

Where it breaks down

High scores on in-prompt long context do not prove reliable retrieval from external corpora or multimodal documents.

Receipts

Sources and further reading

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.