Why this benchmark is useful
Long context is now a product claim on almost every frontier model. LongBench v2 helps separate window size from actual long-context reasoning.
Benchmarks / Long context
LongBench 2
Whether long-context models can reason over long, realistic inputs instead of merely accepting a large token window.
Long context is now a product claim on almost every frontier model. LongBench v2 helps separate window size from actual long-context reasoning.
Read scores next to the context length, prompt setting, and whether chain-of-thought was allowed. The same model can move when the reasoning setup changes.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Claude Opus 4.8AnthropicNo-CoT setting. CoT evaluation shows different results — read both. | 2026-05-28 | 58.4% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 2 | Llama 4 ScoutMeta10M context variant — strong on retrieval-heavy document tasks. | 2026-04-05 | 57.2% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 3 | GPT 5.5OpenAI | 2026-04-23 | 56.1% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 4 | Gemini 3.1 ProGoogle | 2026-02-19 | 55.4% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 5 | Gemini 2.0 UltraGoogle | 2024-12-11 | 54.8% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 6 | GPT 5.4OpenAI | 2026-03-05 | 54.2% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 7 | Gemini 3.5 FlashGoogle | 2026-05-19 | 53.8% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 8 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 52.6% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 9 | Gemini 3 Flash PreviewGoogle | 2025-12-17 | 51.4% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 10 | o3OpenAIReasoning models show mixed results on retrieval-heavy long-context tasks. | 2025-04-16 | 51.2% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 11 | Mistral Large 3Mistral | 2025-12-02 | 50.8% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 12 | Llama 4 MaverickMeta | 2026-04-05 | 49.8% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 13 | Kimi K2.6Moonshot | 2026-04-20 | 49.1% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 14 | DeepSeek V4DeepSeek | 2026-04-24 | 48.6% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 15 | Kimi K2.5Moonshot | 2026-01-27 | 48.2% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 16 | Qwen 3.7 MaxAlibaba | 2026-05-20 | 47.8% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 17 | Grok 4.3xAI | 2026-04-30 | 46.2% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 18 | MiniMax M3MiniMax | 2026-05-31 | 45.6% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 19 | Claude Haiku 4.5Anthropic | 2025-10-01 | 44.8% | LongBench v2 leaderboard2026-04-01 | Needs audit |
Primary source is the LongBench v2 leaderboard. VerdictPal should preserve the no-CoT / CoT setting with every row.
A good LongBench result does not prove citation fidelity, source selection, or privacy-safe document handling.
The Pack · Editorial newsletter
One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.