Why this benchmark is useful
It turns code ability into a patch or pass-rate signal you can inspect. That makes it useful for comparing agent harnesses, not just base models.
Benchmarks / Coding
Whether a model can resolve real GitHub issues from popular Python repos — it must produce a patch that makes the project's hidden test suite pass. 'Verified' is a 500-issue, human-validated subset of the original SWE-bench.
It turns code ability into a patch or pass-rate signal you can inspect. That makes it useful for comparing agent harnesses, not just base models.
Read SWE-bench Verified as a saturated signal with medium contamination risk. Compare models only when the source uses the same harness, prompting setup, sampling policy, and score unit.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Claude Mythos PreviewAnthropicBest-of-3 agent harness; compare to single-attempt runs on the same leaderboard before drawing conclusions. | 2026-06-01 | 93.9% | SWE-bench official leaderboard2026-06-02 | Source-checked |
| 2 | Claude Opus 4.8Anthropic | 2026-05-28 | 88.6% | SWE-bench official leaderboard2026-05-30 | Source-checked |
| 3 | Claude Opus 4.7 (Adaptive)Anthropic | 2026-04-16 | 87.6% | SWE-bench official leaderboard2026-05-30 | Source-checked |
| 4 | GPT 5.5OpenAIDesk-attributed from GPT 5.5 card — reproducible harness score, not a vendor slide. | 2026-04-23 | 82.5% | SWE-bench Verified leaderboard2026-06-01 | Source-checked |
| 4 | Claude Opus 4.5 (SEAL)Anthropic | 2025-11-01 | 80.9% | SWE-bench SEAL leaderboard2026-05-30 | Source-checked |
| 5 | Claude Opus 4.6Anthropic | 2026-02-01 | 80.8% | SWE-bench official leaderboard2026-05-30 | Source-checked |
| 6 | DeepSeek V4 Pro MaxDeepSeek | 2026-04-24 | 80.6% | SWE-bench official leaderboard2026-05-30 | Source-checked |
| 6 | Gemini 3.1 ProGoogle | 2026-02-19 | 80.6% | SWE-bench official leaderboard2026-05-30 | Source-checked |
| 8 | Kimi K2.6Moonshot | 2026-04-20 | 80.2% | SWE-bench official leaderboard2026-05-30 | Source-checked |
| 8 | MiniMax M2.5MiniMax | 2026-04-01 | 80.2% | SWE-bench official leaderboard2026-05-30 | Source-checked |
| 10 | GPT 5.2OpenAI | 2025-12-11 | 80% | SWE-bench official leaderboard2026-05-30 | Source-checked |
| 10 | GPT 5.4OpenAI | 2026-03-05 | 80% | SWE-bench official leaderboard2026-05-30 | Source-checked |
| 12 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 79.6% | SWE-bench official leaderboard2026-05-30 | Source-checked |
| 14 | Qwen 3.7 MaxAlibaba | 2026-05-20 | 79.4% | SWE-bench Verified leaderboard2026-06-09 | Source-checked |
| 13 | DeepSeek V4 Flash MaxDeepSeek | 2026-05-10 | 79% | SWE-bench official leaderboard2026-05-30 | Source-checked |
| 14 | Qwen 3.6 PlusAlibaba | 2026-04-01 | 78.8% | SWE-bench official leaderboard2026-05-30 | Source-checked |
| 17 | MiniMax M3MiniMax | 2026-05-31 | 78.4% | SWE-bench Verified leaderboard2026-06-09 | Source-checked |
| 15 | Gemini 3 FlashGoogle | 2025-12-17 | 78% | SWE-bench official leaderboard2026-05-30 | Source-checked |
| 19 | Grok 4.3xAI | 2026-04-30 | 77.6% | SWE-bench Verified leaderboard2026-06-09 | Source-checked |
| 20 | Gemini 3.5 FlashGoogle | 2026-05-19 | 77.2% | SWE-bench Verified leaderboard2026-06-09 | Source-checked |
| 21 | Mistral Large 3Mistral | 2025-12-02 | 76.2% | SWE-bench Verified leaderboard2026-06-09 | Source-checked |
| 22 | Llama 4 MaverickMeta | 2026-04-05 | 74.8% | SWE-bench Verified leaderboard2026-06-09 | Source-checked |
| 23 | LongCat-2.0MeituanVendor-reported SWE-bench Pro score, cited as narrowly above GPT-5.5 in Meituan's release coverage. Not the canonical SWE-bench Verified row. | 2026-06-29 | 59.5% | Meituan LongCat-2.0 release materials2026-06-29 | Needs audit |
Frontier models now cluster ~78–87% on Verified; contamination and harness variance pin discrimination. Prefer DeepSWE or SWE-Lancer for current coding-agent signal.
No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.
The Pack · Editorial newsletter
One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.