Why this benchmark is useful
When you care whether a model can select, sequence, and recover tool calls over MCP — schema adherence and chain-of-tool reasoning, not chat polish.
How reliably a model selects, sequences, and recovers when invoking tools exposed via the Model Context Protocol — including schema adherence, error handling, and chain-of-tool reasoning.
When you care whether a model can select, sequence, and recover tool calls over MCP — schema adherence and chain-of-tool reasoning, not chat polish.
End-to-end task completion percentage against an MCP server. Harness and server implementation can swing scores; cost of retries is not scored.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Gemini 3.5 FlashGoogle | 2026-05-19 | 83.6% | MCP Atlas leaderboard2026-05-15 | Source-checked |
| 2 | Claude Opus 4.8Anthropic | 2026-05-28 | 79.2% | MCP Atlas leaderboard2026-06-09 | Needs audit |
| 3 | GPT 5.5OpenAI | 2026-04-23 | 77.8% | MCP Atlas leaderboard2026-06-09 | Needs audit |
| 4 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 75.4% | MCP Atlas leaderboard2026-06-09 | Needs audit |
| 5 | Gemini 3.1 ProGoogle | 2026-02-19 | 74.1% | MCP Atlas leaderboard2026-06-09 | Needs audit |
| 6 | DeepSeek V4 Pro MaxDeepSeek | 2026-04-24 | 72.6% | MCP Atlas leaderboard2026-06-09 | Needs audit |
| 7 | Qwen 3.7 MaxAlibaba | 2026-05-20 | 71.8% | MCP Atlas leaderboard2026-06-09 | Needs audit |
| 8 | Mistral Large 3Mistral | 2025-12-02 | 70.4% | MCP Atlas leaderboard2026-06-09 | Needs audit |
| 9 | Kimi K2.6Moonshot | 2026-04-20 | 69.7% | MCP Atlas leaderboard2026-06-09 | Needs audit |
| 10 | MiniMax M3MiniMax | 2026-05-31 | 68.9% | MCP Atlas leaderboard2026-06-09 | Needs audit |
This benchmark sits in the tool use family. It uses Scored tool-call sequences against an MCP server; the suite measures both call correctness and recovery from intermediate errors. The score should travel with its task format, scoring method, source date, and benchmark version.
No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
How reliably a model selects, sequences, and recovers when invoking tools exposed via the Model Context Protocol — including schema adherence, error handling, and chain-of-tool reasoning.
A strong MCP Atlas result says nothing about:
Percentage of tasks completed end-to-end; see the source repo for harness and scoring script details. Task format: Scored tool-call sequences against an MCP server; the suite measures both call correctness and recovery from intermediate errors.
MCP Atlas is currently marked Active in the atlas.
Contamination risk for MCP Atlas is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.