Benchmarks / Long context

LongBench v2

LongBench 2

Whether long-context models can reason over long, realistic inputs instead of merely accepting a large token window.

What this does not measure
  • Retrieval quality — the benchmark tests model use of provided context, not whether your search system found the right documents.
  • Every long document task — legal review, literature review, codebase navigation, and meeting analysis fail in different ways.
  • A guaranteed usable context window — advertised token limits can still degrade before the limit on real prompts.
Analysis

Why this benchmark is useful

Long context is now a product claim on almost every frontier model. LongBench v2 helps separate window size from actual long-context reasoning.

Scope

Coverage map

Task family
Long context
Format
Long-document and long-context reasoning tasks with leaderboard settings for no-CoT and CoT evaluation.
Scoring
Accuracy-style aggregate by evaluation setting. Compare only within the same setting.
Maintainer
LongBench team
Reading guide

How to read the scores

Read scores next to the context length, prompt setting, and whether chain-of-thought was allowed. The same model can move when the reasoning setup changes.

Blind spots

What it does not cover

  • Retrieval quality — the benchmark tests model use of provided context, not whether your search system found the right documents.
  • Every long document task — legal review, literature review, codebase navigation, and meeting analysis fail in different ways.
  • A guaranteed usable context window — advertised token limits can still degrade before the limit on real prompts.
Scores

Evidence ledger

19 rows
1919 rows
58.4%best score
0source-checked
1sources
2026-04-01source date
LongBench19

official-page · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Claude Opus 4.858.4% · LongBench v2 leaderboard
  2. Llama 4 Scout57.2% · LongBench v2 leaderboard
  3. GPT 5.556.1% · LongBench v2 leaderboard
  4. Gemini 3.1 Pro55.4% · LongBench v2 leaderboard
  5. Gemini 2.0 Ultra54.8% · LongBench v2 leaderboard

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Claude Opus 4.8AnthropicNo-CoT setting. CoT evaluation shows different results — read both.2026-05-2858.4%
LongBench v2 leaderboard2026-04-01
Needs audit
2Llama 4 ScoutMeta10M context variant — strong on retrieval-heavy document tasks.2026-04-0557.2%
LongBench v2 leaderboard2026-04-01
Needs audit
3GPT 5.5OpenAI2026-04-2356.1%
LongBench v2 leaderboard2026-04-01
Needs audit
4Gemini 3.1 ProGoogle2026-02-1955.4%
LongBench v2 leaderboard2026-04-01
Needs audit
5Gemini 2.0 UltraGoogle2024-12-1154.8%
LongBench v2 leaderboard2026-04-01
Needs audit
6GPT 5.4OpenAI2026-03-0554.2%
LongBench v2 leaderboard2026-04-01
Needs audit
7Gemini 3.5 FlashGoogle2026-05-1953.8%
LongBench v2 leaderboard2026-04-01
Needs audit
8Claude Sonnet 4.6Anthropic2026-02-1752.6%
LongBench v2 leaderboard2026-04-01
Needs audit
9Gemini 3 Flash PreviewGoogle2025-12-1751.4%
LongBench v2 leaderboard2026-04-01
Needs audit
10o3OpenAIReasoning models show mixed results on retrieval-heavy long-context tasks.2025-04-1651.2%
LongBench v2 leaderboard2026-04-01
Needs audit
11Mistral Large 3Mistral2025-12-0250.8%
LongBench v2 leaderboard2026-04-01
Needs audit
12Llama 4 MaverickMeta2026-04-0549.8%
LongBench v2 leaderboard2026-04-01
Needs audit
13Kimi K2.6Moonshot2026-04-2049.1%
LongBench v2 leaderboard2026-04-01
Needs audit
14DeepSeek V4DeepSeek2026-04-2448.6%
LongBench v2 leaderboard2026-04-01
Needs audit
15Kimi K2.5Moonshot2026-01-2748.2%
LongBench v2 leaderboard2026-04-01
Needs audit
16Qwen 3.7 MaxAlibaba2026-05-2047.8%
LongBench v2 leaderboard2026-04-01
Needs audit
17Grok 4.3xAI2026-04-3046.2%
LongBench v2 leaderboard2026-04-01
Needs audit
18MiniMax M3MiniMax2026-05-3145.6%
LongBench v2 leaderboard2026-04-01
Needs audit
19Claude Haiku 4.5Anthropic2025-10-0144.8%
LongBench v2 leaderboard2026-04-01
Needs audit
Method

What it covers

Primary source is the LongBench v2 leaderboard. VerdictPal should preserve the no-CoT / CoT setting with every row.

Score ceiling

Where it breaks down

A good LongBench result does not prove citation fidelity, source selection, or privacy-safe document handling.

Tools that report it

Receipts

Sources and further reading

Questions about this benchmark

Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.

What does LongBench v2 measure?

Whether long-context models can reason over long, realistic inputs instead of merely accepting a large token window.

What does a high LongBench v2 score not prove?

A strong LongBench v2 result says nothing about:

  • Retrieval quality — the benchmark tests model use of provided context, not whether your search system found the right documents.
  • Every long document task — legal review, literature review, codebase navigation, and meeting analysis fail in different ways.
  • A guaranteed usable context window — advertised token limits can still degrade before the limit on real prompts.

How is LongBench v2 scored?

Accuracy-style aggregate by evaluation setting. Compare only within the same setting. Task format: Long-document and long-context reasoning tasks with leaderboard settings for no-CoT and CoT evaluation.

Is LongBench v2 saturated?

LongBench v2 is currently marked Active in the atlas.

Can LongBench v2 results be contaminated by training data?

Contamination risk for LongBench v2 is graded Medium contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.