Why this benchmark is useful
Editorial brief pendingWe publish the methodology and ledger first; benchmark-specific analysis ships after desk review.
Benchmarks / Coding
Long-horizon software engineering on 113 contamination-free tasks across 91 repos and five languages — patches must pass hand-written behavioral verifiers, not string diffs.
Editorial brief pendingWe publish the methodology and ledger first; benchmark-specific analysis ships after desk review.
Reading guide pending. Use task format, scoring method, and source dates in the ledger until the desk brief ships.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Claude Fable 5 (max)Anthropic | 2026-06-09 | 70% | DeepSWE leaderboard v1.12026-06-20 | Source-checked |
| 2 | GPT 5.5 (xhigh)OpenAI | 2026-04-23 | 67% | DeepSWE leaderboard v1.12026-06-20 | Source-checked |
| 3 | Claude Opus 4.8 (max)Anthropic | 2026-05-28 | 59% | DeepSWE leaderboard v1.12026-06-20 | Source-checked |
| 4 | GPT 5.4 (xhigh)OpenAI | 2026-03-15 | 52% | DeepSWE leaderboard v1.12026-06-20 | Source-checked |
| 5 | GLM 5.2 (max)Zhipu | 2026-05-15 | 44% | DeepSWE leaderboard v1.12026-06-20 | Source-checked |
| 6 | Gemini 3.5 Flash (medium)Google | 2026-04-10 | 37% | DeepSWE leaderboard v1.12026-06-20 | Source-checked |
Data-quality note pending. Every ledger row still carries source URL, source date, and ingest timestamp.
No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.
The Pack · Editorial newsletter
One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.