Why this benchmark is useful
Long context is now a product claim on almost every frontier model. LongBench v2 helps separate window size from actual long-context reasoning.
LongBench 2
Whether long-context models can reason over long, realistic inputs instead of merely accepting a large token window.
Long context is now a product claim on almost every frontier model. LongBench v2 helps separate window size from actual long-context reasoning.
Read scores next to the context length, prompt setting, and whether chain-of-thought was allowed. The same model can move when the reasoning setup changes.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Claude Opus 4.8AnthropicNo-CoT setting. CoT evaluation shows different results — read both. | 2026-05-28 | 58.4% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 2 | Llama 4 ScoutMeta10M context variant — strong on retrieval-heavy document tasks. | 2026-04-05 | 57.2% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 3 | GPT 5.5OpenAI | 2026-04-23 | 56.1% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 4 | Gemini 3.1 ProGoogle | 2026-02-19 | 55.4% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 5 | Gemini 2.0 UltraGoogle | 2024-12-11 | 54.8% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 6 | GPT 5.4OpenAI | 2026-03-05 | 54.2% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 7 | Gemini 3.5 FlashGoogle | 2026-05-19 | 53.8% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 8 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 52.6% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 9 | Gemini 3 Flash PreviewGoogle | 2025-12-17 | 51.4% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 10 | o3OpenAIReasoning models show mixed results on retrieval-heavy long-context tasks. | 2025-04-16 | 51.2% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 11 | Mistral Large 3Mistral | 2025-12-02 | 50.8% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 12 | Llama 4 MaverickMeta | 2026-04-05 | 49.8% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 13 | Kimi K2.6Moonshot | 2026-04-20 | 49.1% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 14 | DeepSeek V4DeepSeek | 2026-04-24 | 48.6% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 15 | Kimi K2.5Moonshot | 2026-01-27 | 48.2% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 16 | Qwen 3.7 MaxAlibaba | 2026-05-20 | 47.8% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 17 | Grok 4.3xAI | 2026-04-30 | 46.2% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 18 | MiniMax M3MiniMax | 2026-05-31 | 45.6% | LongBench v2 leaderboard2026-04-01 | Needs audit |
| 19 | Claude Haiku 4.5Anthropic | 2025-10-01 | 44.8% | LongBench v2 leaderboard2026-04-01 | Needs audit |
Primary source is the LongBench v2 leaderboard. VerdictPal should preserve the no-CoT / CoT setting with every row.
A good LongBench result does not prove citation fidelity, source selection, or privacy-safe document handling.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
Whether long-context models can reason over long, realistic inputs instead of merely accepting a large token window.
A strong LongBench v2 result says nothing about:
Accuracy-style aggregate by evaluation setting. Compare only within the same setting. Task format: Long-document and long-context reasoning tasks with leaderboard settings for no-CoT and CoT evaluation.
LongBench v2 is currently marked Active in the atlas.
Contamination risk for LongBench v2 is graded Medium contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.