Why this benchmark is useful
Tool calling is the plumbing behind agents, automations, and research assistants. BFCL gives that plumbing a public test instead of treating every function-call demo as evidence.
Benchmarks / Tool use
BFCL
Whether a model can choose and format function calls correctly across single-turn, multi-turn, parallel, and live API-style tool-use tasks.
Tool calling is the plumbing behind agents, automations, and research assistants. BFCL gives that plumbing a public test instead of treating every function-call demo as evidence.
Compare models within the same BFCL version and category. A model can be strong at simple calls and weak at parallel or multi-turn calls.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Claude Sonnet 4.6AnthropicV4 aggregate score; drill into multi-turn and parallel sub-scores before choosing. | 2026-02-17 | 88.7% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 2 | GPT 5.5OpenAI | 2026-04-23 | 87.4% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 3 | Claude Opus 4.8Anthropic | 2026-05-28 | 86.9% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 4 | Kimi K2.6Moonshot | 2026-04-20 | 85.2% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 5 | GPT 5.4OpenAI | 2026-03-05 | 83.8% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 6 | Gemini 3.5 FlashGoogle | 2026-05-19 | 82.3% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 7 | Gemini 3.1 ProGoogle | 2026-02-19 | 81.6% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 8 | Mistral Large 3Mistral | 2025-12-02 | 80.1% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 9 | Gemini 3 Flash PreviewGoogle | 2025-12-17 | 79.4% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 10 | Qwen 3.7 MaxAlibaba | 2026-05-20 | 78.4% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 11 | o3OpenAIStrong in simple single-turn; drops in multi-turn complexity. | 2025-04-16 | 78.1% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 12 | Grok 4.3xAI | 2026-04-30 | 77.2% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 13 | DeepSeek V4 Flash MaxDeepSeek | 2026-04-24 | 76.8% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 14 | MiniMax M3MiniMax | 2026-05-31 | 76.2% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 15 | Llama 4 MaverickMeta | 2026-04-05 | 75.9% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 16 | DeepSeek V4DeepSeek | 2026-04-24 | 74.6% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 17 | Llama 4 ScoutMeta | 2026-04-05 | 74.1% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 18 | DeepSeek R1DeepSeek | 2025-01-20 | 72.4% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 19 | Claude Haiku 4.5Anthropic | 2025-10-01 | 71.8% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 20 | Cohere Command ACohere | 2026-03-15 | 70.6% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
Primary source is the Berkeley Gorilla leaderboard. VerdictPal should store category and BFCL version with every score row.
The leaderboard does not replace a product-level integration test with your schemas, auth model, error handling, and adversarial inputs.
The Pack · Editorial newsletter
One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.