Why this benchmark is useful
Long-context research breaks when models lose thread across distant evidence. AA-LCR is one of the few public suites that grades synthesis across long inputs instead of retrieval trivia alone.
Benchmarks / Long context
AA-LCR
Long-context synthesis and reasoning: 100 open-answer questions requiring models to integrate evidence across long inputs. 6% weight in AA Intelligence Index v4.1.
Long-context research breaks when models lose thread across distant evidence. AA-LCR is one of the few public suites that grades synthesis across long inputs instead of retrieval trivia alone.
Read pass@1 with three repeats as a conservative headline. Compare only rows from the same AA harness version and long-context window.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | GPT-5.1OpenAI | 2025-11-13 | 75% | 2026-07-21 | Source-checked |
| 2 | Kimi K3Moonshot | 2026-07-16 | 74.67% | 2026-07-21 | Source-checked |
| 3 | GPT 5.4OpenAI | 2026-03-05 | 74% | 2026-07-21 | Source-checked |
| 4 | GPT-5.3 Codex (xhigh)OpenAI | 2026-02-05 | 74% | 2026-07-21 | Source-checked |
| 5 | GPT-5.6 LunaOpenAI | 2026-07-09 | 74% | 2026-07-21 | Source-checked |
| 6 | GPT-5.6 TerraOpenAI | 2026-07-09 | 74% | 2026-07-21 | Source-checked |
| 7 | MiniMax M3MiniMax | 2026-05-31 | 74% | 2026-07-21 | Source-checked |
| 8 | GPT-5.6 SolOpenAI | 2026-07-09 | 73.67% | 2026-07-21 | Source-checked |
| 9 | GPT-5.5 (high)OpenAI | 2026-04-23 | 73.33% | 2026-07-21 | Source-checked |
| 10 | MiMo-V2.5-ProXiaomi | 2026-04-22 | 73.33% | 2026-07-21 | Source-checked |
| 11 | Gemini 3.1 Pro PreviewGoogle | 2026-02-19 | 72.67% | 2026-07-21 | Source-checked |
| 12 | GPT-5.2OpenAI | 2025-12-11 | 72.67% | 2026-07-21 | Source-checked |
| 13 | GPT-5.5 (medium)OpenAI | 2026-04-23 | 72.33% | 2026-07-21 | Source-checked |
| 14 | GPT-5.5 (low)OpenAI | 2026-04-23 | 72% | 2026-07-21 | Source-checked |
| 15 | GLM-5.2Zhipu | 2026-06-13 | 71.33% | 2026-07-21 | Source-checked |
| 16 | Gemini 3.5 Flash (medium)Google | 2026-05-19 | 71% | 2026-07-21 | Source-checked |
| 17 | Claude Opus 4.6 (Adaptive Reasoning, Max Effort)Anthropic | 2026-02-05 | 70.67% | 2026-07-21 | Source-checked |
| 18 | Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)Anthropic | 2026-02-17 | 70.67% | 2026-07-21 | Source-checked |
| 19 | Claude Sonnet 5Anthropic | 2026-06-30 | 70.67% | 2026-07-21 | Source-checked |
| 20 | Gemini 3 ProGoogle | 2025-11-18 | 70.67% | 2026-07-21 | Source-checked |
| 21 | Claude Opus 4.7Anthropic | 2026-04-16 | 70.33% | 2026-07-21 | Source-checked |
| 22 | Claude Fable 5Anthropic | 2026-06-09 | 70% | 2026-07-21 | Source-checked |
| 23 | Kimi K2.6Moonshot | 2026-04-20 | 69.67% | 2026-07-21 | Source-checked |
| 24 | Muse SparkOther | 2026-01-01 | 69.67% | 2026-07-21 | Source-checked |
| 25 | Gemini 3.5 FlashGoogle | 2026-05-19 | 69.33% | 2026-07-21 | Source-checked |
| 26 | GPT-5.4 MiniOpenAI | 2026-03-17 | 69.33% | 2026-07-21 | Source-checked |
| 27 | OpenAI o3OpenAI | 2025-04-16 | 69.33% | 2026-07-21 | Source-checked |
| 28 | Qwen3.7 MaxAlibaba | 2026-05-19 | 69% | 2026-07-21 | Source-checked |
| 29 | MiniMax M2.7MiniMax | 2026-04-15 | 68.67% | 2026-07-21 | Source-checked |
| 30 | Grok 4.5xAI | 2026-07-08 | 67.67% | 2026-07-21 | Source-checked |
| 31 | Claude Opus 4.7 (Non-reasoning, High Effort)Anthropic | 2026-04-16 | 67% | 2026-07-21 | Source-checked |
| 32 | Nemotron 3 Ultra 550B A55B (Reasoning)NVIDIA | 2026-06-04 | 67% | 2026-07-21 | Source-checked |
| 33 | DeepSeek V4 Pro (Reasoning, Max Effort)DeepSeek | 2026-04-24 | 66.33% | 2026-07-21 | Source-checked |
| 34 | Kimi K2.7 CodeMoonshot | 2026-06-12 | 66.33% | 2026-07-21 | Source-checked |
| 35 | GPT-5.4 NanoOpenAI | 2026-03-17 | 66% | 2026-07-21 | Source-checked |
| 36 | MiniMax M2.5MiniMax | 2026-04-01 | 66% | 2026-07-21 | Source-checked |
| 37 | Claude Opus 4.5Anthropic | 2025-11-24 | 65.33% | 2026-07-21 | Source-checked |
| 38 | Gemini 3.1 Flash LiteGoogle | 2026-05-07 | 65.33% | 2026-07-21 | Source-checked |
| 39 | Kimi K2.5Moonshot | 2026-01-27 | 65.33% | 2026-07-21 | Source-checked |
| 40 | Qwen 3.7 PlusAlibaba | 2026-06-02 | 65% | 2026-07-21 | Source-checked |
| 41 | Grok 4.3xAI | 2026-04-30 | 64.33% | 2026-07-21 | Source-checked |
| 42 | Step 3.7 FlashStepFun | 2026-06-01 | 63.67% | 2026-07-21 | Source-checked |
| 43 | Muse Spark 1.1Meta | 2026-07-09 | 63.33% | 2026-07-21 | Source-checked |
| 44 | GLM-5.1Zhipu | 2026-04-07 | 62.33% | 2026-07-21 | Source-checked |
| 45 | Gemma 4 31B ITGoogle | 2026-03-01 | 62% | 2026-07-21 | Source-checked |
| 46 | OpenAI o1OpenAI | 2024-09-12 | 59.33% | 2026-07-21 | Source-checked |
| 47 | Claude Opus 4.6Anthropic | 2026-02-05 | 58.33% | 2026-07-21 | Source-checked |
| 1 | Claude Opus 4.8Anthropic | 2026-05-28 | 58.3% | 2026-06-17 | Source-checked |
| 49 | Grok 4.20 ReasoningxAI | 2026-03-05 | 58% | 2026-07-21 | Source-checked |
| 50 | LongCat-2.0Meituan | 2026-06-29 | 58% | 2026-07-21 | Source-checked |
| 51 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 57.67% | 2026-07-21 | Source-checked |
| 52 | Gemma 4 12B (Reasoning)Google | 2026-06-07 | 55.33% | 2026-07-21 | Source-checked |
| 53 | DeepSeek R1DeepSeek | 2025-01-20 | 54.67% | 2026-07-21 | Source-checked |
| 2 | GPT 5.5OpenAI | 2026-04-23 | 52.1% | 2026-06-17 | Source-checked |
| 55 | Gemini 3 Flash PreviewGoogle | 2025-12-17 | 48% | 2026-07-21 | Source-checked |
| 56 | Llama 4 MaverickMeta | 2026-04-05 | 46% | 2026-07-21 | Source-checked |
| 57 | Mistral Large 3Mistral | 2025-12-02 | 34.67% | 2026-07-21 | Source-checked |
| 58 | North Mini CodeCohere | 2026-06-09 | 32.33% | 2026-07-21 | Source-checked |
| 59 | HyperNova 60B 2605Multiverse | 2026-05-06 | 31.67% | 2026-07-21 | Source-checked |
| 60 | Gemma 4 12B (Non-reasoning)Google | 2026-06-10 | 30.67% | 2026-07-21 | Source-checked |
| 61 | Llama 4 ScoutMeta | 2026-04-05 | 25.8% | 2026-07-21 | Source-checked |
| 62 | Command ACohere | 2026-03-15 | 18% | 2026-07-21 | Source-checked |
| 63 | MiniCPM5-1B (Reasoning)OpenBMB | 2026-06-04 | 3.67% | 2026-07-21 | Source-checked |
| 64 | LFM2.5-8B-A1BLiquid | 2026-06-08 | 0% | 2026-07-21 | Source-checked |
This benchmark sits in the long context family. It uses 100 open-answer questions with long context passages; equality-checker LLM grading, pass@1. The score should travel with its task format, scoring method, source date, and benchmark version.
High scores on in-prompt long context do not prove reliable retrieval from external corpora or multimodal documents.
The Pack · Editorial newsletter
One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.