Why this benchmark is useful
Once the flagship real-issue Python coding test — still useful lineage. Frontier scores now cluster tightly; prefer DeepSWE or SWE-Lancer for current agent signal.
Whether a model can resolve real GitHub issues from popular Python repos — it must produce a patch that makes the project's hidden test suite pass. 'Verified' is a 500-issue, human-validated subset of the original SWE-bench.
Once the flagship real-issue Python coding test — still useful lineage. Frontier scores now cluster tightly; prefer DeepSWE or SWE-Lancer for current agent signal.
% of Verified issues resolved (hidden tests pass). Harness, retries, and tools swing scores tens of points — always read the agent setup. ~78–87% clustering means weak frontier separation today.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Claude Mythos PreviewAnthropicBest-of-3 agent harness; compare to single-attempt runs on the same leaderboard before drawing conclusions. | 2026-06-01 | 93.9% | SWE-bench official leaderboard2026-06-02 | Source-checked |
| 2 | Claude Opus 4.8Anthropic | 2026-05-28 | 88.6% | SWE-bench official leaderboard2026-05-30 | Source-checked |
| 3 | Claude Opus 4.7 (Adaptive)Anthropic | 2026-04-16 | 87.6% | SWE-bench official leaderboard2026-05-30 | Source-checked |
| 4 | GPT 5.5OpenAIDesk-attributed from GPT 5.5 card — reproducible harness score, not a vendor slide. | 2026-04-23 | 82.5% | SWE-bench Verified leaderboard2026-06-01 | Source-checked |
| 4 | Claude Opus 4.5 (SEAL)Anthropic | 2025-11-01 | 80.9% | SWE-bench SEAL leaderboard2026-05-30 | Source-checked |
| 5 | Claude Opus 4.6Anthropic | 2026-02-01 | 80.8% | SWE-bench official leaderboard2026-05-30 | Source-checked |
| 6 | DeepSeek V4 Pro MaxDeepSeek | 2026-04-24 | 80.6% | SWE-bench official leaderboard2026-05-30 | Source-checked |
| 6 | Gemini 3.1 ProGoogle | 2026-02-19 | 80.6% | SWE-bench official leaderboard2026-05-30 | Source-checked |
| 8 | Kimi K2.6Moonshot | 2026-04-20 | 80.2% | SWE-bench official leaderboard2026-05-30 | Source-checked |
| 8 | MiniMax M2.5MiniMax | 2026-04-01 | 80.2% | SWE-bench official leaderboard2026-05-30 | Source-checked |
| 10 | GPT 5.2OpenAI | 2025-12-11 | 80% | SWE-bench official leaderboard2026-05-30 | Source-checked |
| 10 | GPT 5.4OpenAI | 2026-03-05 | 80% | SWE-bench official leaderboard2026-05-30 | Source-checked |
| 12 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 79.6% | SWE-bench official leaderboard2026-05-30 | Source-checked |
| 14 | Qwen 3.7 MaxAlibaba | 2026-05-20 | 79.4% | SWE-bench Verified leaderboard2026-06-09 | Source-checked |
| 13 | DeepSeek V4 Flash MaxDeepSeek | 2026-05-10 | 79% | SWE-bench official leaderboard2026-05-30 | Source-checked |
| 14 | Qwen 3.6 PlusAlibaba | 2026-04-01 | 78.8% | SWE-bench official leaderboard2026-05-30 | Source-checked |
| 17 | MiniMax M3MiniMax | 2026-05-31 | 78.4% | SWE-bench Verified leaderboard2026-06-09 | Source-checked |
| 15 | Gemini 3 FlashGoogle | 2025-12-17 | 78% | SWE-bench official leaderboard2026-05-30 | Source-checked |
| 19 | Grok 4.3xAI | 2026-04-30 | 77.6% | SWE-bench Verified leaderboard2026-06-09 | Source-checked |
| 20 | Gemini 3.5 FlashGoogle | 2026-05-19 | 77.2% | SWE-bench Verified leaderboard2026-06-09 | Source-checked |
| 21 | Mistral Large 3Mistral | 2025-12-02 | 76.2% | SWE-bench Verified leaderboard2026-06-09 | Source-checked |
| 22 | Llama 4 MaverickMeta | 2026-04-05 | 74.8% | SWE-bench Verified leaderboard2026-06-09 | Source-checked |
| 23 | LongCat-2.0MeituanVendor-reported SWE-bench Pro score, cited as narrowly above GPT-5.5 in Meituan's release coverage. Not the canonical SWE-bench Verified row. | 2026-06-29 | 59.5% | Meituan LongCat-2.0 release materials2026-06-29 | Needs audit |
Frontier models now cluster ~78–87% on Verified; contamination and harness variance pin discrimination. Prefer DeepSWE or SWE-Lancer for current coding-agent signal.
No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
Whether a model can resolve real GitHub issues from popular Python repos — it must produce a patch that makes the project's hidden test suite pass. 'Verified' is a 500-issue, human-validated subset of the original SWE-bench.
A strong SWE-bench Verified result says nothing about:
% of issues resolved (tests pass). Task format: 500 human-validated GitHub issues; success = the gold test suite passes after the model's patch.
SWE-bench Verified is currently marked Saturated in the atlas.
Contamination risk for SWE-bench Verified is graded Medium contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.