Why this benchmark is useful
Replaces τ²-Bench Telecom in AA Index v4.1 with harder, more realistic agentic scenarios that better separate frontier models.
Benchmarks / Agentic
tau3-bench-banking
Dual-control agent-user simulation for banking tasks: 97 scenarios with knowledge retrieval, backend database state evaluation, pass@1. 14% weight in AA Intelligence Index v4.1.
Replaces τ²-Bench Telecom in AA Index v4.1 with harder, more realistic agentic scenarios that better separate frontier models.
Read τ³-Bench Banking as a defunct signal with low contamination risk. Compare models only when the source uses the same harness, prompting setup, sampling policy, and score unit.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Gemini 3.5 Flash (medium)Google | 2026-05-19 | 95.61% | 2026-07-21 | Source-checked |
| 2 | MiniMax M2.5MiniMax | 2026-04-01 | 95.32% | 2026-07-21 | Source-checked |
| 3 | Grok 4.20 ReasoningxAI | 2026-03-05 | 92.98% | 2026-07-21 | Source-checked |
| 4 | Claude Opus 4.6 (Adaptive Reasoning, Max Effort)Anthropic | 2026-02-05 | 92.11% | 2026-07-21 | Source-checked |
| 5 | Gemini 3 ProGoogle | 2025-11-18 | 87.13% | 2026-07-21 | Source-checked |
| 6 | Claude Opus 4.5Anthropic | 2025-11-24 | 86.26% | 2026-07-21 | Source-checked |
| 7 | GPT-5.3 Codex (xhigh)OpenAI | 2026-02-05 | 85.96% | 2026-07-21 | Source-checked |
| 8 | Claude Opus 4.6Anthropic | 2026-02-05 | 84.8% | 2026-07-21 | Source-checked |
| 9 | GPT-5.2OpenAI | 2025-12-11 | 84.8% | 2026-07-21 | Source-checked |
| 10 | MiniCPM5-1B (Reasoning)OpenBMB | 2026-06-04 | 80.99% | 2026-07-21 | Source-checked |
| 11 | OpenAI o3OpenAI | 2025-04-16 | 80.7% | 2026-07-21 | Source-checked |
| 12 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 79.53% | 2026-07-21 | Source-checked |
| 13 | Claude Opus 4.7 (Non-reasoning, High Effort)Anthropic | 2026-04-16 | 73.98% | 2026-07-21 | Source-checked |
| 14 | OpenAI o1OpenAI | 2024-09-12 | 62.57% | 2026-07-21 | Source-checked |
| 1 | Claude Fable 5Anthropic | 2026-06-09 | 62.4% | 2026-06-17 | Source-checked |
| 2 | GPT 5.5OpenAI | 2026-04-23 | 54.8% | 2026-06-17 | Source-checked |
| 17 | Gemini 3 Flash PreviewGoogle | 2025-12-17 | 43.27% | 2026-07-21 | Source-checked |
| 18 | DeepSeek R1DeepSeek | 2025-01-20 | 36.55% | 2026-07-21 | Source-checked |
| 19 | Gemma 4 12B (Reasoning)Google | 2026-06-07 | 36.26% | 2026-07-21 | Source-checked |
| 20 | Kimi K3Moonshot | 2026-07-16 | 33.4% | 2026-07-21 | Source-checked |
| 21 | GPT-5.6 SolOpenAI | 2026-07-09 | 32.99% | 2026-07-21 | Source-checked |
| 22 | Grok 4.5xAI | 2026-07-08 | 32.58% | 2026-07-21 | Source-checked |
| 23 | Gemma 4 12B (Non-reasoning)Google | 2026-06-10 | 31.87% | 2026-07-21 | Source-checked |
| 24 | GPT-5.6 TerraOpenAI | 2026-07-09 | 31.75% | 2026-07-21 | Source-checked |
| 25 | Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)Anthropic | 2026-02-17 | 30.52% | 2026-07-21 | Source-checked |
| 26 | GPT 5.4OpenAI | 2026-03-05 | 30.31% | 2026-07-21 | Source-checked |
| 27 | GPT-5.5 (high)OpenAI | 2026-04-23 | 29.48% | 2026-07-21 | Source-checked |
| 28 | Claude Opus 4.7Anthropic | 2026-04-16 | 28.87% | 2026-07-21 | Source-checked |
| 29 | Claude Sonnet 5Anthropic | 2026-06-30 | 28.25% | 2026-07-21 | Source-checked |
| 30 | Claude Opus 4.8Anthropic | 2026-05-28 | 27.63% | 2026-07-21 | Source-checked |
| 31 | GPT-5.6 LunaOpenAI | 2026-07-09 | 27.22% | 2026-07-21 | Source-checked |
| 32 | GLM-5.2Zhipu | 2026-06-13 | 26.8% | 2026-07-21 | Source-checked |
| 33 | DeepSeek V4 Pro (Reasoning, Max Effort)DeepSeek | 2026-04-24 | 25.77% | 2026-07-21 | Source-checked |
| 34 | GPT-5.5 (medium)OpenAI | 2026-04-23 | 25.77% | 2026-07-21 | Source-checked |
| 35 | Gemini 3.5 FlashGoogle | 2026-05-19 | 25.36% | 2026-07-21 | Source-checked |
| 36 | Muse Spark 1.1Meta | 2026-07-09 | 25.15% | 2026-07-21 | Source-checked |
| 37 | GPT-5.4 MiniOpenAI | 2026-03-17 | 21.44% | 2026-07-21 | Source-checked |
| 38 | GPT-5.5 (low)OpenAI | 2026-04-23 | 21.24% | 2026-07-21 | Source-checked |
| 39 | GPT-5.4 NanoOpenAI | 2026-03-17 | 21.03% | 2026-07-21 | Source-checked |
| 40 | Kimi K2.6Moonshot | 2026-04-20 | 20.62% | 2026-07-21 | Source-checked |
| 41 | Muse SparkOther | 2026-01-01 | 19.59% | 2026-07-21 | Source-checked |
| 42 | Kimi K2.7 CodeMoonshot | 2026-06-12 | 18.14% | 2026-07-21 | Source-checked |
| 43 | Qwen 3.7 PlusAlibaba | 2026-06-02 | 17.87% | 2026-07-21 | Source-checked |
| 44 | Gemini 3.1 Pro PreviewGoogle | 2026-02-19 | 16.49% | 2026-07-21 | Source-checked |
| 45 | LFM2.5-8B-A1BLiquid | 2026-06-08 | 16.08% | 2026-07-21 | Source-checked |
| 46 | Command ACohere | 2026-03-15 | 15.2% | 2026-07-21 | Source-checked |
| 47 | Gemma 4 31B ITGoogle | 2026-03-01 | 15.05% | 2026-07-21 | Source-checked |
| 48 | Kimi K2.5Moonshot | 2026-01-27 | 14.23% | 2026-07-21 | Source-checked |
| 49 | GPT-5.1OpenAI | 2025-11-13 | 14.02% | 2026-07-21 | Source-checked |
| 50 | Nemotron 3 Ultra 550B A55B (Reasoning)NVIDIA | 2026-06-04 | 13.81% | 2026-07-21 | Source-checked |
| 51 | MiniMax M3MiniMax | 2026-05-31 | 12.99% | 2026-07-21 | Source-checked |
| 52 | LongCat-2.0Meituan | 2026-06-29 | 12.78% | 2026-07-21 | Source-checked |
| 53 | Grok 4.3xAI | 2026-04-30 | 12.16% | 2026-07-21 | Source-checked |
| 54 | GLM-5.1Zhipu | 2026-04-07 | 11.55% | 2026-07-21 | Source-checked |
| 55 | Step 3.7 FlashStepFun | 2026-06-01 | 11.34% | 2026-07-21 | Source-checked |
| 56 | Qwen3.7 MaxAlibaba | 2026-05-19 | 10.93% | 2026-07-21 | Source-checked |
| 57 | MiniMax M2.7MiniMax | 2026-04-15 | 8.87% | 2026-07-21 | Source-checked |
| 58 | Gemini 3.1 Flash LiteGoogle | 2026-05-07 | 8.66% | 2026-07-21 | Source-checked |
| 59 | MiMo-V2.5-ProXiaomi | 2026-04-22 | 8.66% | 2026-07-21 | Source-checked |
| 60 | North Mini CodeCohere | 2026-06-09 | 6.39% | 2026-07-21 | Source-checked |
| 61 | Mistral Large 3Mistral | 2025-12-02 | 5.77% | 2026-07-21 | Source-checked |
| 62 | HyperNova 60B 2605Multiverse | 2026-05-06 | 5.36% | 2026-07-21 | Source-checked |
| 63 | Llama 4 MaverickMeta | 2026-04-05 | 3.92% | 2026-07-21 | Source-checked |
| 64 | Llama 4 ScoutMeta | 2026-04-05 | 3.3% | 2026-07-21 | Source-checked |
Closed — verified AA τ³-Bench Banking pass@1 scores now cluster at 99%+. Retained for methodology context; use harder agentic suites for frontier separation.
No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.
The Pack · Editorial newsletter
One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.