Benchmarks / Long context

LongBench v2

LongBench 2

Whether long-context models can reason over long, realistic inputs instead of merely accepting a large token window.

Long contextSolidActiveMedium contamination riskSince 2025
What this does not measure
  • Retrieval quality — the benchmark tests model use of provided context, not whether your search system found the right documents.
  • Every long document task — legal review, literature review, codebase navigation, and meeting analysis fail in different ways.
  • A guaranteed usable context window — advertised token limits can still degrade before the limit on real prompts.
Analysis

Why this benchmark is useful

Long context is now a product claim on almost every frontier model. LongBench v2 helps separate window size from actual long-context reasoning.

Scope

Coverage map

Task family
Long context
Format
Long-document and long-context reasoning tasks with leaderboard settings for no-CoT and CoT evaluation.
Scoring
Accuracy-style aggregate by evaluation setting. Compare only within the same setting.
Maintainer
LongBench team
Reading guide

How to read the scores

Read scores next to the context length, prompt setting, and whether chain-of-thought was allowed. The same model can move when the reasoning setup changes.

Blind spots

What it does not cover

  • Retrieval quality — the benchmark tests model use of provided context, not whether your search system found the right documents.
  • Every long document task — legal review, literature review, codebase navigation, and meeting analysis fail in different ways.
  • A guaranteed usable context window — advertised token limits can still degrade before the limit on real prompts.
Scores

Evidence ledger

19 rows
1919 rows
58.4%best score
0source-checked
1sources
2026-04-01source date
LongBench19

official-page · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Claude Opus 4.858.4% · LongBench v2 leaderboard
  2. Llama 4 Scout57.2% · LongBench v2 leaderboard
  3. GPT 5.556.1% · LongBench v2 leaderboard
  4. Gemini 3.1 Pro55.4% · LongBench v2 leaderboard
  5. Gemini 2.0 Ultra54.8% · LongBench v2 leaderboard

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Claude Opus 4.8AnthropicNo-CoT setting. CoT evaluation shows different results — read both.2026-05-2858.4%
LongBench v2 leaderboard2026-04-01
Needs audit
2Llama 4 ScoutMeta10M context variant — strong on retrieval-heavy document tasks.2026-04-0557.2%
LongBench v2 leaderboard2026-04-01
Needs audit
3GPT 5.5OpenAI2026-04-2356.1%
LongBench v2 leaderboard2026-04-01
Needs audit
4Gemini 3.1 ProGoogle2026-02-1955.4%
LongBench v2 leaderboard2026-04-01
Needs audit
5Gemini 2.0 UltraGoogle2024-12-1154.8%
LongBench v2 leaderboard2026-04-01
Needs audit
6GPT 5.4OpenAI2026-03-0554.2%
LongBench v2 leaderboard2026-04-01
Needs audit
7Gemini 3.5 FlashGoogle2026-05-1953.8%
LongBench v2 leaderboard2026-04-01
Needs audit
8Claude Sonnet 4.6Anthropic2026-02-1752.6%
LongBench v2 leaderboard2026-04-01
Needs audit
9Gemini 3 Flash PreviewGoogle2025-12-1751.4%
LongBench v2 leaderboard2026-04-01
Needs audit
10o3OpenAIReasoning models show mixed results on retrieval-heavy long-context tasks.2025-04-1651.2%
LongBench v2 leaderboard2026-04-01
Needs audit
11Mistral Large 3Mistral2025-12-0250.8%
LongBench v2 leaderboard2026-04-01
Needs audit
12Llama 4 MaverickMeta2026-04-0549.8%
LongBench v2 leaderboard2026-04-01
Needs audit
13Kimi K2.6Moonshot2026-04-2049.1%
LongBench v2 leaderboard2026-04-01
Needs audit
14DeepSeek V4DeepSeek2026-04-2448.6%
LongBench v2 leaderboard2026-04-01
Needs audit
15Kimi K2.5Moonshot2026-01-2748.2%
LongBench v2 leaderboard2026-04-01
Needs audit
16Qwen 3.7 MaxAlibaba2026-05-2047.8%
LongBench v2 leaderboard2026-04-01
Needs audit
17Grok 4.3xAI2026-04-3046.2%
LongBench v2 leaderboard2026-04-01
Needs audit
18MiniMax M3MiniMax2026-05-3145.6%
LongBench v2 leaderboard2026-04-01
Needs audit
19Claude Haiku 4.5Anthropic2025-10-0144.8%
LongBench v2 leaderboard2026-04-01
Needs audit
Method

What it covers

Primary source is the LongBench v2 leaderboard. VerdictPal should preserve the no-CoT / CoT setting with every row.

Score ceiling

Where it breaks down

A good LongBench result does not prove citation fidelity, source selection, or privacy-safe document handling.

Tools that report it

Receipts

Sources and further reading

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.