Benchmarks / Coding

DeepSWE

Long-horizon software engineering on 113 contamination-free tasks across 91 repos and five languages — patches must pass hand-written behavioral verifiers, not string diffs.

CodingFlagshipActiveLow contamination riskSince 2026
What this does not measure
  • Python-only issue fixing like SWE-bench Verified — DeepSWE spans Go, Rust, TypeScript, and more.
  • Cheap runs — frontier agents burn six-figure output tokens per task on the public harness.
  • Vendor marketing slides — scores come from a fixed mini-swe-agent harness so models are comparable.
Analysis

Why this benchmark is useful

Editorial brief pendingWe publish the methodology and ledger first; benchmark-specific analysis ships after desk review.

Scope

Coverage map

Task family
Coding
Format
113 tasks from scratch (not mined PRs); pass@1 via behavioral verifiers; mini-swe-agent harness.
Scoring
pass@1 (% resolved).
Maintainer
Datacurve
Reading guide

How to read the scores

Reading guide pending. Use task format, scoring method, and source dates in the ledger until the desk brief ships.

Blind spots

What it does not cover

  • Python-only issue fixing like SWE-bench Verified — DeepSWE spans Go, Rust, TypeScript, and more.
  • Cheap runs — frontier agents burn six-figure output tokens per task on the public harness.
  • Vendor marketing slides — scores come from a fixed mini-swe-agent harness so models are comparable.
Scores

Evidence ledger

6 rows
66 rows
70%best score
6source-checked
1sources
2026-06-20source date
DeepSWE6

official-page · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Claude Fable 5 (max)70% · DeepSWE leaderboard v1.1
  2. GPT 5.5 (xhigh)67% · DeepSWE leaderboard v1.1
  3. Claude Opus 4.8 (max)59% · DeepSWE leaderboard v1.1
  4. GPT 5.4 (xhigh)52% · DeepSWE leaderboard v1.1
  5. GLM 5.2 (max)44% · DeepSWE leaderboard v1.1

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Claude Fable 5 (max)Anthropic2026-06-0970%
DeepSWE leaderboard v1.12026-06-20
Source-checked
2GPT 5.5 (xhigh)OpenAI2026-04-2367%
DeepSWE leaderboard v1.12026-06-20
Source-checked
3Claude Opus 4.8 (max)Anthropic2026-05-2859%
DeepSWE leaderboard v1.12026-06-20
Source-checked
4GPT 5.4 (xhigh)OpenAI2026-03-1552%
DeepSWE leaderboard v1.12026-06-20
Source-checked
5GLM 5.2 (max)Zhipu2026-05-1544%
DeepSWE leaderboard v1.12026-06-20
Source-checked
6Gemini 3.5 Flash (medium)Google2026-04-1037%
DeepSWE leaderboard v1.12026-06-20
Source-checked
Method

What it covers

Data-quality note pending. Every ledger row still carries source URL, source date, and ingest timestamp.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

Tools that report it

Receipts

Sources and further reading

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.