Why this benchmark is useful
Long-context research breaks when models lose thread across distant evidence. AA-LCR is one of the few public suites that grades synthesis across long inputs instead of retrieval trivia alone.
AA-LCR
Long-context synthesis and reasoning: 100 open-answer questions requiring models to integrate evidence across long inputs. 6% weight in AA Intelligence Index v4.1.
Long-context research breaks when models lose thread across distant evidence. AA-LCR is one of the few public suites that grades synthesis across long inputs instead of retrieval trivia alone.
Read pass@1 with three repeats as a conservative headline. Compare only rows from the same AA harness version and long-context window.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Muse Spark 1.2Meta | 2026-08-05 | 83.33% | 2026-09-01 | Source-checked |
| 2 | Kimi K3Moonshot | 2026-07-16 | 82.67% | 2026-09-01 | Source-checked |
| 3 | Muse Spark 1.1Meta | 2026-07-09 | 81.33% | 2026-09-01 | Source-checked |
| 4 | Gemini 3.5 FlashGoogle | 2026-05-19 | 81% | 2026-09-01 | Source-checked |
| 5 | MiniMax M3MiniMax | 2026-05-31 | 80.33% | 2026-09-01 | Source-checked |
| 6 | Gemini 3.7 FlashGoogle | 2026-08-13 | 80% | 2026-09-01 | Source-checked |
| 7 | Muse GlimmerMeta | 2026-08-10 | 80% | 2026-09-01 | Source-checked |
| 8 | Gemini 3.5 Flash (medium)Google | 2026-05-19 | 79.67% | 2026-09-01 | Source-checked |
| 9 | GPT-5.6 TerraOpenAI | 2026-07-09 | 79.67% | 2026-09-01 | Source-checked |
| 10 | GPT-5.2OpenAI | 2025-12-11 | 79.33% | 2026-09-01 | Source-checked |
| 11 | GPT-5.2 Codex (xhigh)OpenAI | 2025-12-11 | 79.33% | 2026-09-01 | Source-checked |
| 12 | Gemini 3.1 Pro PreviewGoogle | 2026-02-19 | 79% | 2026-09-01 | Source-checked |
| 13 | Gemini 3.6 Flash (high)Google | 2026-07-21 | 79% | 2026-09-01 | Source-checked |
| 14 | GPT-5.5 (high)OpenAI | 2026-04-23 | 79% | 2026-09-01 | Source-checked |
| 15 | Claude Opus 5 (Adaptive Reasoning, Medium Effort)Anthropic | 2026-07-24 | 78.67% | 2026-09-01 | Source-checked |
| 16 | GPT-5.3 Codex (xhigh)OpenAI | 2026-02-05 | 78.33% | 2026-09-01 | Source-checked |
| 17 | GPT-5.6 LunaOpenAI | 2026-07-09 | 78.33% | 2026-09-01 | Source-checked |
| 18 | Agnes 2.5 Pro BetaOther | 2026-08-26 | 78% | 2026-09-01 | Source-checked |
| 19 | DeepSeek V4 Flash Vision (Reasoning, Max Effort)DeepSeek | 2026-08-21 | 78% | 2026-09-01 | Source-checked |
| 20 | GLM-5.3-FlashZhipu | 2026-08-26 | 78% | 2026-09-01 | Source-checked |
| 21 | GPT 5.4OpenAI | 2026-03-05 | 77.67% | 2026-09-01 | Source-checked |
| 22 | GPT-5.5 (low)OpenAI | 2026-04-23 | 77.67% | 2026-09-01 | Source-checked |
| 23 | GPT-5.6 SolOpenAI | 2026-07-09 | 77.67% | 2026-09-01 | Source-checked |
| 24 | MiMo-V2.5-ProXiaomi | 2026-04-22 | 77.67% | 2026-09-01 | Source-checked |
| 25 | Qwen3.8 27BAlibaba | 2026-08-14 | 77.33% | 2026-09-01 | Source-checked |
| 26 | Claude Opus 5 (Adaptive Reasoning, Low Effort)Anthropic | 2026-07-24 | 77% | 2026-09-01 | Source-checked |
| 27 | Claude Sonnet 5Anthropic | 2026-06-30 | 77% | 2026-09-01 | Source-checked |
| 28 | Kimi K3 (low)Moonshot | 2026-07-16 | 77% | 2026-09-01 | Source-checked |
| 29 | Muse SparkOther | 2026-01-01 | 77% | 2026-09-01 | Source-checked |
| 30 | Qwen3.8-Flash-NextAlibaba | 2026-08-26 | 77% | 2026-09-01 | Source-checked |
| 31 | Claude Fable 5Anthropic | 2026-06-09 | 76.67% | 2026-09-01 | Source-checked |
| 32 | GLM-5.2Zhipu | 2026-06-13 | 76.67% | 2026-09-01 | Source-checked |
| 33 | GPT-5.1OpenAI | 2025-11-13 | 76.67% | 2026-09-01 | Source-checked |
| 34 | GPT-5.5 (medium)OpenAI | 2026-04-23 | 76.67% | 2026-09-01 | Source-checked |
| 35 | Kimi K2.6Moonshot | 2026-04-20 | 76.67% | 2026-09-01 | Source-checked |
| 36 | Claude Opus 5 (Adaptive Reasoning, High Effort)Anthropic | 2026-07-24 | 76.33% | 2026-09-01 | Source-checked |
| 37 | Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)Anthropic | 2026-07-24 | 76.33% | 2026-09-01 | Source-checked |
| 38 | GLM-5.3Zhipu | 2026-08-14 | 76.33% | 2026-09-01 | Source-checked |
| 39 | GPT-5 (high)OpenAI | 2025-08-07 | 76.33% | 2026-09-01 | Source-checked |
| 40 | GPT-5 (medium)OpenAI | 2025-08-07 | 76.33% | 2026-09-01 | Source-checked |
| 41 | Nex-N2-ProOther | 2026-06-02 | 76.33% | 2026-09-01 | Source-checked |
| 42 | Claude Opus 4.5 (Reasoning)Anthropic | 2025-11-24 | 76% | 2026-09-01 | Source-checked |
| 43 | Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic | 2026-07-24 | 75.67% | 2026-09-01 | Source-checked |
| 44 | Claude Opus 4.7Anthropic | 2026-04-16 | 75.33% | 2026-09-01 | Source-checked |
| 45 | DeepSeek V4 Pro 0813 (Reasoning, Max Effort)DeepSeek | 2026-08-13 | 75.33% | 2026-09-01 | Source-checked |
| 46 | MiniMax M2.7MiniMax | 2026-04-15 | 75.33% | 2026-09-01 | Source-checked |
| 47 | Qwen3.8 2.4T A95BAlibaba | 2026-08-12 | 75.33% | 2026-09-01 | Source-checked |
| 48 | Grok 4.6xAI | 2026-08-12 | 75% | 2026-09-01 | Source-checked |
| 49 | Kimi K2.7 CodeMoonshot | 2026-06-12 | 75% | 2026-09-01 | Source-checked |
| 50 | Apodex 1.1Other | 2026-08-30 | 74.67% | 2026-09-01 | Source-checked |
| 51 | Gemini 3.5 Flash-LiteGoogle | 2026-07-21 | 74.67% | 2026-09-01 | Source-checked |
| 52 | Hy3Other | 2026-07-06 | 74.67% | 2026-09-01 | Source-checked |
| 53 | Qwen3.7 MaxAlibaba | 2026-05-19 | 74.67% | 2026-09-01 | Source-checked |
| 54 | Claude Opus 4.6 (Adaptive Reasoning, Max Effort)Anthropic | 2026-02-05 | 74.33% | 2026-09-01 | Source-checked |
| 55 | DeepSeek V4 Flash 0731 (Reasoning, Max Effort)DeepSeek | 2026-07-31 | 74.33% | 2026-09-01 | Source-checked |
| 56 | Qwen3.8 MaxAlibaba | 2026-08-03 | 74.33% | 2026-09-01 | Source-checked |
| 57 | Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)Anthropic | 2026-02-17 | 74% | 2026-09-01 | Source-checked |
| 58 | Grok 4.5xAI | 2026-07-08 | 74% | 2026-09-01 | Source-checked |
| 59 | GPT-5.4 (low)OpenAI | 2026-03-05 | 73.67% | 2026-09-01 | Source-checked |
| 60 | Claude 4.1 Opus (Reasoning)Anthropic | 2025-08-05 | 73.33% | 2026-09-01 | Source-checked |
| 61 | Inkling (xhigh)Other | 2026-07-15 | 73.33% | 2026-09-01 | Source-checked |
| 62 | OpenAI o3OpenAI | 2025-04-16 | 73.33% | 2026-09-01 | Source-checked |
| 63 | Qwen3.6 27B (Reasoning)Alibaba | 2026-04-22 | 73.33% | 2026-09-01 | Source-checked |
| 64 | Agnes 2.5 Pro AlphaOther | 2026-07-24 | 73% | 2026-09-01 | Source-checked |
| 65 | Gemini 3 Flash Preview (Reasoning)Google | 2025-12-17 | 73% | 2026-09-01 | Source-checked |
| 66 | Gemini 3 ProGoogle | 2025-11-18 | 73% | 2026-09-01 | Source-checked |
| 67 | GPT-5.4 MiniOpenAI | 2026-03-17 | 73% | 2026-09-01 | Source-checked |
| 68 | Kimi K2.5Moonshot | 2026-01-27 | 73% | 2026-09-01 | Source-checked |
| 69 | Qwen3.5 397B A17B (Reasoning)Alibaba | 2026-02-16 | 72.67% | 2026-09-01 | Source-checked |
| 70 | Claude Opus 4.7 (Non-reasoning, High Effort)Anthropic | 2026-04-16 | 72.33% | 2026-09-01 | Source-checked |
| 71 | Motif 3Other | 2026-08-12 | 72.33% | 2026-09-01 | Source-checked |
| 72 | Qwen3.5 27B (Reasoning)Alibaba | 2026-02-24 | 72.33% | 2026-09-01 | Source-checked |
| 73 | Qwen3.6 PlusAlibaba | 2026-04-02 | 72.33% | 2026-09-01 | Source-checked |
| 74 | GPT-5.4 NanoOpenAI | 2026-03-17 | 72% | 2026-09-01 | Source-checked |
| 75 | MiniMax M2.5MiniMax | 2026-04-01 | 72% | 2026-09-01 | Source-checked |
| 76 | Qwen3.6 Max PreviewAlibaba | 2026-04-20 | 72% | 2026-09-01 | Source-checked |
| 77 | MiMo-V2-OmniXiaomi | 2026-03-19 | 71.67% | 2026-09-01 | Source-checked |
| 78 | Gemini 3.1 Flash LiteGoogle | 2026-05-07 | 71.33% | 2026-09-01 | Source-checked |
| 79 | GPT-5 Codex (high)OpenAI | 2025-09-23 | 71% | 2026-09-01 | Source-checked |
| 80 | Nemotron 3 Ultra 550B A55B (Reasoning)NVIDIA | 2026-06-04 | 71% | 2026-09-01 | Source-checked |
| 81 | DeepSeek V3.2 (Reasoning)DeepSeek | 2025-12-01 | 70.67% | 2026-09-01 | Source-checked |
| 82 | GLM-5 (Reasoning)Zhipu | 2026-02-11 | 70.67% | 2026-09-01 | Source-checked |
| 83 | Motif 3 (Beta)Other | 2026-07-21 | 70.67% | 2026-09-01 | Source-checked |
| 84 | Solar Pro 4Other | 2026-08-06 | 70.67% | 2026-09-01 | Source-checked |
| 85 | Gemini 3 Pro Preview (low)Google | 2025-11-18 | 70.33% | 2026-09-01 | Source-checked |
| 86 | JT-4.1 Flash 236B A21BOther | 2026-07-09 | 70.33% | 2026-09-01 | Source-checked |
| 87 | Kimi K2 ThinkingMoonshot | 2025-11-06 | 70.33% | 2026-09-01 | Source-checked |
| 88 | Qwen3 Max ThinkingAlibaba | 2026-01-26 | 70.33% | 2026-09-01 | Source-checked |
| 89 | Qwen3.5 122B A10B (Reasoning)Alibaba | 2026-02-24 | 70.33% | 2026-09-01 | Source-checked |
| 90 | Grok Build 0.1 0616xAI | 2026-06-16 | 70% | 2026-09-01 | Source-checked |
| 91 | KAT Coder Pro V2Other | 2026-03-27 | 70% | 2026-09-01 | Source-checked |
| 92 | MiMo-V2-Omni-0327Xiaomi | 2026-03-27 | 69.67% | 2026-09-01 | Source-checked |
| 93 | Step 3.7 FlashStepFun | 2026-06-01 | 69.67% | 2026-09-01 | Source-checked |
| 94 | Inkling SmallOther | 2026-07-30 | 69.33% | 2026-09-01 | Source-checked |
| 95 | DeepSeek V4 Flash (Reasoning, High Effort)DeepSeek | 2026-04-24 | 69% | 2026-09-01 | Source-checked |
| 96 | GPT-5.1 Codex (high)OpenAI | 2025-11-13 | 69% | 2026-09-01 | Source-checked |
| 97 | GPT-5.2 (medium)OpenAI | 2025-12-11 | 69% | 2026-09-01 | Source-checked |
| 98 | Qwen 3.7 PlusAlibaba | 2026-06-02 | 69% | 2026-09-01 | Source-checked |
| 99 | MiMo-V2-Flash (Feb 2026)Xiaomi | 2025-12-16 | 68.67% | 2026-09-01 | Source-checked |
| 100 | Claude 4.5 Sonnet (Reasoning)Anthropic | 2025-09-29 | 68.33% | 2026-09-01 | Source-checked |
| 101 | Gemma 4 31B ITGoogle | 2026-03-01 | 68.33% | 2026-09-01 | Source-checked |
| 102 | Grok 4.3 (medium)xAI | 2026-04-30 | 68.33% | 2026-09-01 | Source-checked |
| 103 | MiMo-V2.5Xiaomi | 2026-04-22 | 68.33% | 2026-09-01 | Source-checked |
| 104 | Solar Open2 250BOther | 2026-08-12 | 68.33% | 2026-09-01 | Source-checked |
| 105 | GLM-4.7 (Reasoning)Zhipu | 2025-12-22 | 68% | 2026-09-01 | Source-checked |
| 106 | GLM-5.1Zhipu | 2026-04-07 | 68% | 2026-09-01 | Source-checked |
| 107 | MiMo-V2-Flash (Reasoning)Xiaomi | 2025-12-16 | 68% | 2026-09-01 | Source-checked |
| 108 | Grok 4.3 (low)xAI | 2026-04-30 | 67.67% | 2026-09-01 | Source-checked |
| 109 | Claude Opus 4.5Anthropic | 2025-11-24 | 67.33% | 2026-09-01 | Source-checked |
| 110 | Ring-2.6-1TOther | 2026-05-08 | 67.33% | 2026-09-01 | Source-checked |
| 111 | DeepSeek V4 Pro (Reasoning, High Effort)DeepSeek | 2026-04-24 | 67% | 2026-09-01 | Source-checked |
| 112 | GPT-5.5 Instant (June 2026)OpenAI | 2026-06-25 | 67% | 2026-09-01 | Source-checked |
| 113 | Grok 4xAI | 2025-07-10 | 67% | 2026-09-01 | Source-checked |
| 114 | Ling 3.0 FlashOther | 2026-08-04 | 67% | 2026-09-01 | Source-checked |
| 115 | GLM-5-TurboZhipu | 2026-03-15 | 66.67% | 2026-09-01 | Source-checked |
| 116 | Qwen3.6 35B A3B (Reasoning)Alibaba | 2026-04-16 | 66.67% | 2026-09-01 | Source-checked |
| 117 | MiniMax-M2.1MiniMax | 2025-12-23 | 66.33% | 2026-09-01 | Source-checked |
| 118 | A.X-K2Other | 2026-08-12 | 66% | 2026-09-01 | Source-checked |
| 119 | Grok 4.3xAI | 2026-04-30 | 66% | 2026-09-01 | Source-checked |
| 120 | Kimi K2.6 (Non-reasoning)Other | 2026-04-20 | 66% | 2026-09-01 | Source-checked |
| 121 | GLM 5V Turbo (Reasoning)Zhipu | 2026-04-01 | 65.67% | 2026-09-01 | Source-checked |
| 122 | MiMo-V2-ProXiaomi | 2026-03-18 | 65.67% | 2026-09-01 | Source-checked |
| 123 | Mistral Medium 3.5Mistral | 2026-04-29 | 65.33% | 2026-09-01 | Source-checked |
| 124 | Claude Sonnet 4.6 (Non-reasoning, Low Effort)Anthropic | 2026-02-17 | 64.67% | 2026-09-01 | Source-checked |
| 125 | Qwen3.6 27B (Non-reasoning)Alibaba | 2026-04-22 | 64.33% | 2026-09-01 | Source-checked |
| 126 | GPT-5.5 Instant (May 2026)OpenAI | 2026-05-05 | 64% | 2026-09-01 | Source-checked |
| 127 | GPT-5 (low)OpenAI | 2025-08-07 | 63.67% | 2026-09-01 | Source-checked |
| 128 | JT-35B-FlashOther | 2026-05-14 | 63.67% | 2026-09-01 | Source-checked |
| 129 | OpenAI o1OpenAI | 2024-09-12 | 63.33% | 2026-09-01 | Source-checked |
| 130 | LongCat-2.0Meituan | 2026-06-29 | 62.67% | 2026-09-01 | Source-checked |
| 131 | Claude Opus 4.6Anthropic | 2026-02-05 | 62.33% | 2026-09-01 | Source-checked |
| 132 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 62.33% | 2026-09-01 | Source-checked |
| 133 | Grok 4.20 ReasoningxAI | 2026-03-05 | 62.33% | 2026-09-01 | Source-checked |
| 134 | G9v3-39A5BOther | 2026-08-03 | 62% | 2026-09-01 | Source-checked |
| 135 | Qwen3.5 397B A17B (Non-reasoning)Alibaba | 2026-02-16 | 62% | 2026-09-01 | Source-checked |
| 136 | Gemma 4 12B (Reasoning)Google | 2026-06-07 | 61.67% | 2026-09-01 | Source-checked |
| 137 | Grok 4.20 0309 (Reasoning)xAI | 2026-03-10 | 60.67% | 2026-09-01 | Source-checked |
| 138 | Hy3-preview (Reasoning)Other | 2026-04-23 | 60.33% | 2026-09-01 | Source-checked |
| 139 | Qwen3.6 35B A3B (Non-reasoning)Alibaba | 2026-04-16 | 60% | 2026-09-01 | Source-checked |
| 140 | Ling 3.0 TinyOther | 2026-08-06 | 58.67% | 2026-09-01 | Source-checked |
| 141 | Gemini 3.5 Flash (minimal)Google | 2026-05-19 | 58.33% | 2026-09-01 | Source-checked |
| 1 | Claude Opus 4.8Anthropic | 2026-05-28 | 58.3% | 2026-06-17 | Source-checked |
| 143 | GPT-5.5 (Non-reasoning)OpenAI | 2026-04-23 | 58% | 2026-09-01 | Source-checked |
| 144 | K-EXAONE 2.0 0803Other | 2026-08-12 | 57.67% | 2026-09-01 | Source-checked |
| 145 | DeepSeek R1DeepSeek | 2025-01-20 | 56.67% | 2026-09-01 | Source-checked |
| 146 | Nemotron 3.5 LightningNVIDIA | 2026-08-11 | 55.33% | 2026-09-01 | Source-checked |
| 147 | Gemini 3 Flash PreviewGoogle | 2025-12-17 | 53% | 2026-09-01 | Source-checked |
| 2 | GPT 5.5OpenAI | 2026-04-23 | 52.1% | 2026-06-17 | Source-checked |
| 149 | EXAONE 4.5 33BOther | 2026-04-09 | 51.67% | 2026-09-01 | Source-checked |
| 150 | Llama 4 MaverickMeta | 2026-04-05 | 50% | 2026-09-01 | Source-checked |
| 151 | DeepSeek V4 Pro (Non-reasoning)DeepSeek | 2026-04-24 | 49.67% | 2026-09-01 | Source-checked |
| 152 | Command ACohere | 2026-03-15 | 48.67% | 2026-09-01 | Source-checked |
| 153 | Granite 4.2 30BOther | 2026-08-25 | 46.67% | 2026-09-01 | Source-checked |
| 154 | GLM-5.2 (Non-reasoning)Zhipu | 2026-06-16 | 45% | 2026-09-01 | Source-checked |
| 155 | GLM-5 (Non-reasoning)Zhipu | 2026-02-11 | 43.33% | 2026-09-01 | Source-checked |
| 156 | Granite 4.2 8BOther | 2026-08-25 | 43.33% | 2026-09-01 | Source-checked |
| 157 | G9v3-3BOther | 2026-07-23 | 40.67% | 2026-09-01 | Source-checked |
| 158 | Nemotron 3 Nano Omni 30B A3B ReasoningNVIDIA | 2026-04-29 | 40.67% | 2026-09-01 | Source-checked |
| 159 | MiMo-V2.5-Pro (Non-reasoning)Xiaomi | 2026-04-22 | 39.33% | 2026-09-01 | Source-checked |
| 160 | Hy3-preview (Non-reasoning)Other | 2026-04-23 | 38.67% | 2026-09-01 | Source-checked |
| 161 | Ling-2.6-1TOther | 2026-04-23 | 38% | 2026-09-01 | Source-checked |
| 162 | DeepSeek V4 Flash (Non-reasoning)DeepSeek | 2026-04-24 | 37.33% | 2026-09-01 | Source-checked |
| 163 | Claude 4 Opus (Reasoning)Anthropic | 2025-05-22 | 36.33% | 2026-09-01 | Source-checked |
| 164 | North Mini CodeCohere | 2026-06-09 | 36% | 2026-09-01 | Source-checked |
| 165 | HyperNova 60B 2605Multiverse | 2026-05-06 | 35.33% | 2026-09-01 | Source-checked |
| 166 | Mistral Large 3Mistral | 2025-12-02 | 34.67% | 2026-09-01 | Source-checked |
| 167 | Gemma 4 12B (Non-reasoning)Google | 2026-06-10 | 31.33% | 2026-09-01 | Source-checked |
| 168 | Llama 4 ScoutMeta | 2026-04-05 | 30.33% | 2026-09-01 | Source-checked |
| 169 | Grok 4.3 (Non-reasoning)xAI | 2026-04-30 | 28% | 2026-09-01 | Source-checked |
| 170 | Ling 2.6 FlashOther | 2026-04-21 | 28% | 2026-09-01 | Source-checked |
| 171 | Granite 4.2 3BOther | 2026-08-25 | 24.33% | 2026-09-01 | Source-checked |
| 172 | Grok 4.20 0309 v2 (Non-reasoning)xAI | 2026-04-07 | 20.67% | 2026-09-01 | Source-checked |
| 173 | Granite 4.1 30BOther | 2026-04-29 | 20.33% | 2026-09-01 | Source-checked |
| 174 | DiffusionGemma 26B A4BGoogle | 2026-06-10 | 18.33% | 2026-09-01 | Source-checked |
| 175 | JT-MINIOther | 2026-04-15 | 12.67% | 2026-09-01 | Source-checked |
| 176 | MiniCPM-V 4.6 1.3BOpenBMB | 2026-05-11 | 6.67% | 2026-09-01 | Source-checked |
| 177 | MiniCPM5-1B (Non-reasoning)OpenBMB | 2026-05-25 | 4.67% | 2026-09-01 | Source-checked |
| 178 | MiniCPM5-1B (Reasoning)OpenBMB | 2026-06-04 | 4.67% | 2026-09-01 | Source-checked |
| 179 | LFM2.5-8B-A1BLiquid | 2026-06-08 | 0% | 2026-09-01 | Source-checked |
This benchmark sits in the long context family. It uses 100 open-answer questions with long context passages; equality-checker LLM grading, pass@1. The score should travel with its task format, scoring method, source date, and benchmark version.
High scores on in-prompt long context do not prove reliable retrieval from external corpora or multimodal documents.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
Long-context synthesis and reasoning: 100 open-answer questions requiring models to integrate evidence across long inputs. 6% weight in AA Intelligence Index v4.1.
A strong AA-LCR result says nothing about:
pass@1 accuracy with 3 repeats. Task format: 100 open-answer questions with long context passages; equality-checker LLM grading, pass@1.
AA-LCR is currently marked Active in the atlas.
Contamination risk for AA-LCR is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.