Why this benchmark is useful
When you need long-horizon, multi-language software work with behavioral verifiers — harder contamination posture than classic SWE-bench Verified.
Long-horizon software engineering on 113 contamination-free tasks across 91 repos and five languages — patches must pass hand-written behavioral verifiers, not string diffs.
When you need long-horizon, multi-language software work with behavioral verifiers — harder contamination posture than classic SWE-bench Verified.
pass@1 (% resolved) on a fixed mini-swe-agent harness. Compare harness-matched runs; token burn per task is high, so cost is part of the story.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Claude Opus 5 (max)Anthropic | 2026-07-24 | 74% | DeepSWE leaderboard2026-08-04 | Source-checked |
| 2 | GPT-5.6 Sol (max)OpenAI | 2026-07-09 | 73% | DeepSWE leaderboard2026-08-04 | Source-checked |
| 2 | GPT 5.5 (xhigh)OpenAI | 2026-04-23 | 67% | DeepSWE leaderboard v1.12026-06-20 | Source-checked |
| 3 | Claude Fable 5 (max)Anthropic | 2026-06-09 | 70% | DeepSWE leaderboard2026-08-04 | Source-checked |
| 3 | Claude Opus 4.8 (max)Anthropic | 2026-05-28 | 59% | DeepSWE leaderboard v1.12026-06-20 | Source-checked |
| 4 | GPT 5.4 (xhigh)OpenAI | 2026-03-15 | 52% | DeepSWE leaderboard v1.12026-06-20 | Source-checked |
| 5 | GLM 5.2 (max)Zhipu | 2026-05-15 | 44% | DeepSWE leaderboard v1.12026-06-20 | Source-checked |
| 6 | Gemini 3.5 Flash (medium)Google | 2026-04-10 | 37% | DeepSWE leaderboard v1.12026-06-20 | Source-checked |
This benchmark sits in the coding family. It uses 113 tasks from scratch (not mined PRs); pass@1 via behavioral verifiers; mini-swe-agent harness. The score should travel with its task format, scoring method, source date, and benchmark version.
No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
Long-horizon software engineering on 113 contamination-free tasks across 91 repos and five languages — patches must pass hand-written behavioral verifiers, not string diffs.
A strong DeepSWE result says nothing about:
pass@1 (% resolved). Task format: 113 tasks from scratch (not mined PRs); pass@1 via behavioral verifiers; mini-swe-agent harness.
DeepSWE is currently marked Active in the atlas.
Contamination risk for DeepSWE is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.