Why this benchmark is useful
Tool calling is the plumbing behind agents, automations, and research assistants. BFCL gives that plumbing a public test instead of treating every function-call demo as evidence.
BFCL
Whether a model can choose and format function calls correctly across single-turn, multi-turn, parallel, and live API-style tool-use tasks.
Tool calling is the plumbing behind agents, automations, and research assistants. BFCL gives that plumbing a public test instead of treating every function-call demo as evidence.
Compare models within the same BFCL version and category. A model can be strong at simple calls and weak at parallel or multi-turn calls.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Claude Sonnet 4.6AnthropicV4 aggregate score; drill into multi-turn and parallel sub-scores before choosing. | 2026-02-17 | 88.7% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 2 | GPT 5.5OpenAI | 2026-04-23 | 87.4% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 3 | Claude Opus 4.8Anthropic | 2026-05-28 | 86.9% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 4 | Kimi K2.6Moonshot | 2026-04-20 | 85.2% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 5 | GPT 5.4OpenAI | 2026-03-05 | 83.8% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 6 | Gemini 3.5 FlashGoogle | 2026-05-19 | 82.3% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 7 | Gemini 3.1 ProGoogle | 2026-02-19 | 81.6% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 8 | Mistral Large 3Mistral | 2025-12-02 | 80.1% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 9 | Gemini 3 Flash PreviewGoogle | 2025-12-17 | 79.4% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 10 | Qwen 3.7 MaxAlibaba | 2026-05-20 | 78.4% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 11 | o3OpenAIStrong in simple single-turn; drops in multi-turn complexity. | 2025-04-16 | 78.1% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 12 | Grok 4.3xAI | 2026-04-30 | 77.2% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 13 | DeepSeek V4 Flash MaxDeepSeek | 2026-04-24 | 76.8% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 14 | MiniMax M3MiniMax | 2026-05-31 | 76.2% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 15 | Llama 4 MaverickMeta | 2026-04-05 | 75.9% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 16 | DeepSeek V4DeepSeek | 2026-04-24 | 74.6% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 17 | Llama 4 ScoutMeta | 2026-04-05 | 74.1% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 18 | DeepSeek R1DeepSeek | 2025-01-20 | 72.4% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 19 | Claude Haiku 4.5Anthropic | 2025-10-01 | 71.8% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
| 20 | Cohere Command ACohere | 2026-03-15 | 70.6% | Berkeley Function Calling Leaderboard V42026-04-15 | Needs audit |
Primary source is the Berkeley Gorilla leaderboard. VerdictPal should store category and BFCL version with every score row.
The leaderboard does not replace a product-level integration test with your schemas, auth model, error handling, and adversarial inputs.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
Whether a model can choose and format function calls correctly across single-turn, multi-turn, parallel, and live API-style tool-use tasks.
A strong BFCL result says nothing about:
Accuracy-style score by BFCL category and version. The exact category matters more than the aggregate. Task format: Function/tool-calling tasks with schema-constrained expected outputs across simple, multi-turn, parallel, and live-style settings.
BFCL is currently marked Active in the atlas.
Contamination risk for BFCL is graded Medium contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.