Benchmarks / Coding

DeepSWE

Long-horizon software engineering on 113 contamination-free tasks across 91 repos and five languages — patches must pass hand-written behavioral verifiers, not string diffs.

What this does not measure
  • Python-only issue fixing like SWE-bench Verified — DeepSWE spans Go, Rust, TypeScript, and more.
  • Cheap runs — frontier agents burn six-figure output tokens per task on the public harness.
  • Vendor marketing slides — scores come from a fixed mini-swe-agent harness so models are comparable.
Analysis

Why this benchmark is useful

When you need long-horizon, multi-language software work with behavioral verifiers — harder contamination posture than classic SWE-bench Verified.

Scope

Coverage map

Task family
Coding
Format
113 tasks from scratch (not mined PRs); pass@1 via behavioral verifiers; mini-swe-agent harness.
Scoring
pass@1 (% resolved).
Maintainer
Datacurve
Reading guide

How to read the scores

pass@1 (% resolved) on a fixed mini-swe-agent harness. Compare harness-matched runs; token burn per task is high, so cost is part of the story.

Blind spots

What it does not cover

  • Python-only issue fixing like SWE-bench Verified — DeepSWE spans Go, Rust, TypeScript, and more.
  • Cheap runs — frontier agents burn six-figure output tokens per task on the public harness.
  • Vendor marketing slides — scores come from a fixed mini-swe-agent harness so models are comparable.
Scores

Evidence ledger

8 rows
88 rows
74%best score
8source-checked
1sources
2026-08-04to 2026-06-20
DeepSWE8

official-page · manual

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Claude Opus 5 (max)74% · DeepSWE leaderboard
  2. GPT-5.6 Sol (max)73% · DeepSWE leaderboard
  3. Claude Fable 5 (max)70% · DeepSWE leaderboard
  4. GPT 5.5 (xhigh)67% · DeepSWE leaderboard v1.1
  5. Claude Opus 4.8 (max)59% · DeepSWE leaderboard v1.1

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Claude Opus 5 (max)Anthropic2026-07-2474%
DeepSWE leaderboard2026-08-04
Source-checked
2GPT-5.6 Sol (max)OpenAI2026-07-0973%
DeepSWE leaderboard2026-08-04
Source-checked
2GPT 5.5 (xhigh)OpenAI2026-04-2367%
DeepSWE leaderboard v1.12026-06-20
Source-checked
3Claude Fable 5 (max)Anthropic2026-06-0970%
DeepSWE leaderboard2026-08-04
Source-checked
3Claude Opus 4.8 (max)Anthropic2026-05-2859%
DeepSWE leaderboard v1.12026-06-20
Source-checked
4GPT 5.4 (xhigh)OpenAI2026-03-1552%
DeepSWE leaderboard v1.12026-06-20
Source-checked
5GLM 5.2 (max)Zhipu2026-05-1544%
DeepSWE leaderboard v1.12026-06-20
Source-checked
6Gemini 3.5 Flash (medium)Google2026-04-1037%
DeepSWE leaderboard v1.12026-06-20
Source-checked
Method

What it covers

This benchmark sits in the coding family. It uses 113 tasks from scratch (not mined PRs); pass@1 via behavioral verifiers; mini-swe-agent harness. The score should travel with its task format, scoring method, source date, and benchmark version.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

Tools that report it

Receipts

Sources and further reading

Questions about this benchmark

Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.

What does DeepSWE measure?

Long-horizon software engineering on 113 contamination-free tasks across 91 repos and five languages — patches must pass hand-written behavioral verifiers, not string diffs.

What does a high DeepSWE score not prove?

A strong DeepSWE result says nothing about:

  • Python-only issue fixing like SWE-bench Verified — DeepSWE spans Go, Rust, TypeScript, and more.
  • Cheap runs — frontier agents burn six-figure output tokens per task on the public harness.
  • Vendor marketing slides — scores come from a fixed mini-swe-agent harness so models are comparable.

How is DeepSWE scored?

pass@1 (% resolved). Task format: 113 tasks from scratch (not mined PRs); pass@1 via behavioral verifiers; mini-swe-agent harness.

Is DeepSWE saturated?

DeepSWE is currently marked Active in the atlas.

Can DeepSWE results be contaminated by training data?

Contamination risk for DeepSWE is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.