Why this benchmark is useful
Replaces τ²-Bench Telecom in AA Index v4.1 with harder, more realistic agentic scenarios that better separate frontier models.
tau3-bench-banking
Dual-control agent-user simulation for banking tasks: 97 scenarios with knowledge retrieval, backend database state evaluation, pass@1. 14% weight in AA Intelligence Index v4.1.
Replaces τ²-Bench Telecom in AA Index v4.1 with harder, more realistic agentic scenarios that better separate frontier models.
pass@1 on backend database state across 97 banking scenarios. Artificial Analysis now reports τ³ on a low scale (typical verified scores well below 60) — treat it as a live agentic separator, not a solved exam.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Qwen3.8 MaxAlibaba | 2026-08-03 | 51.34% | 2026-09-01 | Source-checked |
| 2 | Grok 4.6xAI | 2026-08-12 | 50.72% | 2026-09-01 | Source-checked |
| 3 | GLM-5.3Zhipu | 2026-08-14 | 50.31% | 2026-09-01 | Source-checked |
| 4 | Qwen3.8 2.4T A95BAlibaba | 2026-08-12 | 49.07% | 2026-09-01 | Source-checked |
| 5 | Qwen3.8 27BAlibaba | 2026-08-14 | 48.04% | 2026-09-01 | Source-checked |
| 6 | GLM-5.3-FlashZhipu | 2026-08-26 | 47.22% | 2026-09-01 | Source-checked |
| 7 | Kimi K3Moonshot | 2026-07-16 | 45.98% | 2026-09-01 | Source-checked |
| 8 | Qwen3.8-Flash-NextAlibaba | 2026-08-26 | 45.36% | 2026-09-01 | Source-checked |
| 9 | Claude Opus 5 (Adaptive Reasoning, High Effort)Anthropic | 2026-07-24 | 44.74% | 2026-09-01 | Source-checked |
| 10 | GPT-5.6 SolOpenAI | 2026-07-09 | 44.33% | 2026-09-01 | Source-checked |
| 11 | Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)Anthropic | 2026-07-24 | 43.3% | 2026-09-01 | Source-checked |
| 12 | Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic | 2026-07-24 | 42.06% | 2026-09-01 | Source-checked |
| 13 | Grok 4.5xAI | 2026-07-08 | 42.06% | 2026-09-01 | Source-checked |
| 14 | Kimi K3 (low)Moonshot | 2026-07-16 | 41.65% | 2026-09-01 | Source-checked |
| 15 | DeepSeek V4 Flash Vision (Reasoning, Max Effort)DeepSeek | 2026-08-21 | 41.03% | 2026-09-01 | Source-checked |
| 16 | GPT-5.6 TerraOpenAI | 2026-07-09 | 40.21% | 2026-09-01 | Source-checked |
| 17 | DeepSeek V4 Pro 0813 (Reasoning, Max Effort)DeepSeek | 2026-08-13 | 39.59% | 2026-09-01 | Source-checked |
| 18 | GPT 5.4OpenAI | 2026-03-05 | 39.59% | 2026-09-01 | Source-checked |
| 19 | DeepSeek V4 Flash 0731 (Reasoning, Max Effort)DeepSeek | 2026-07-31 | 39.38% | 2026-09-01 | Source-checked |
| 20 | GPT 5.5OpenAI | 2026-04-23 | 38.97% | 2026-09-01 | Source-checked |
| 21 | Claude Opus 5 (Adaptive Reasoning, Medium Effort)Anthropic | 2026-07-24 | 38.56% | 2026-09-01 | Source-checked |
| 22 | Claude Fable 5Anthropic | 2026-06-09 | 38.14% | 2026-09-01 | Source-checked |
| 23 | Claude Sonnet 5Anthropic | 2026-06-30 | 37.32% | 2026-09-01 | Source-checked |
| 24 | GPT-5.5 (high)OpenAI | 2026-04-23 | 36.7% | 2026-09-01 | Source-checked |
| 25 | Agnes 2.5 Pro BetaOther | 2026-08-26 | 35.67% | 2026-09-01 | Source-checked |
| 26 | Motif 3Other | 2026-08-12 | 35.26% | 2026-09-01 | Source-checked |
| 27 | Muse Spark 1.2Meta | 2026-08-05 | 34.85% | 2026-09-01 | Source-checked |
| 28 | Claude Opus 4.7Anthropic | 2026-04-16 | 34.64% | 2026-09-01 | Source-checked |
| 29 | GLM-5.2Zhipu | 2026-06-13 | 34.64% | 2026-09-01 | Source-checked |
| 30 | Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)Anthropic | 2026-02-17 | 34.43% | 2026-09-01 | Source-checked |
| 31 | Claude Opus 4.8Anthropic | 2026-05-28 | 34.23% | 2026-09-01 | Source-checked |
| 32 | Gemini 3.7 FlashGoogle | 2026-08-13 | 32.78% | 2026-09-01 | Source-checked |
| 33 | Gemini 3.5 FlashGoogle | 2026-05-19 | 32.16% | 2026-09-01 | Source-checked |
| 34 | Muse Spark 1.1Meta | 2026-07-09 | 31.75% | 2026-09-01 | Source-checked |
| 35 | GPT-5.6 LunaOpenAI | 2026-07-09 | 31.13% | 2026-09-01 | Source-checked |
| 36 | Claude Opus 5 (Adaptive Reasoning, Low Effort)Anthropic | 2026-07-24 | 30.31% | 2026-09-01 | Source-checked |
| 37 | Gemini 3.6 Flash (high)Google | 2026-07-21 | 29.9% | 2026-09-01 | Source-checked |
| 38 | GPT-5.5 (medium)OpenAI | 2026-04-23 | 29.9% | 2026-09-01 | Source-checked |
| 39 | Inkling (xhigh)Other | 2026-07-15 | 29.07% | 2026-09-01 | Source-checked |
| 40 | GPT-5.4 NanoOpenAI | 2026-03-17 | 27.42% | 2026-09-01 | Source-checked |
| 41 | Ling 3.0 FlashOther | 2026-08-04 | 27.22% | 2026-09-01 | Source-checked |
| 42 | DeepSeek V4 Flash (Reasoning, High Effort)DeepSeek | 2026-04-24 | 26.19% | 2026-09-01 | Source-checked |
| 43 | DeepSeek V4 Pro (Reasoning, High Effort)DeepSeek | 2026-04-24 | 26.19% | 2026-09-01 | Source-checked |
| 44 | GPT-5.4 MiniOpenAI | 2026-03-17 | 25.57% | 2026-09-01 | Source-checked |
| 45 | Apodex 1.1Other | 2026-08-30 | 25.15% | 2026-09-01 | Source-checked |
| 46 | GPT-5.5 (low)OpenAI | 2026-04-23 | 24.95% | 2026-09-01 | Source-checked |
| 47 | Claude 4.5 Sonnet (Reasoning)Anthropic | 2025-09-29 | 24.54% | 2026-09-01 | Source-checked |
| 48 | Muse GlimmerMeta | 2026-08-10 | 23.51% | 2026-09-01 | Source-checked |
| 49 | Kimi K2.6Moonshot | 2026-04-20 | 23.3% | 2026-09-01 | Source-checked |
| 50 | Solar Pro 4Other | 2026-08-06 | 23.3% | 2026-09-01 | Source-checked |
| 51 | Hy3Other | 2026-07-06 | 22.89% | 2026-09-01 | Source-checked |
| 52 | G9v3-39A5BOther | 2026-08-03 | 22.06% | 2026-09-01 | Source-checked |
| 53 | GPT-5 (high)OpenAI | 2025-08-07 | 22.06% | 2026-09-01 | Source-checked |
| 54 | Solar Open2 250BOther | 2026-08-12 | 21.65% | 2026-09-01 | Source-checked |
| 55 | Gemini 3.1 Pro PreviewGoogle | 2026-02-19 | 21.44% | 2026-09-01 | Source-checked |
| 56 | Gemini 3 Flash Preview (Reasoning)Google | 2025-12-17 | 20.82% | 2026-09-01 | Source-checked |
| 57 | Ling 3.0 TinyOther | 2026-08-06 | 20.82% | 2026-09-01 | Source-checked |
| 58 | Qwen3.6 PlusAlibaba | 2026-04-02 | 20.82% | 2026-09-01 | Source-checked |
| 59 | Kimi K2.7 CodeMoonshot | 2026-06-12 | 20.21% | 2026-09-01 | Source-checked |
| 60 | Inkling SmallOther | 2026-07-30 | 18.76% | 2026-09-01 | Source-checked |
| 61 | Ring-2.6-1TOther | 2026-05-08 | 17.94% | 2026-09-01 | Source-checked |
| 62 | Gemini 3.5 Flash-LiteGoogle | 2026-07-21 | 17.53% | 2026-09-01 | Source-checked |
| 63 | Nex-N2-ProOther | 2026-06-02 | 17.53% | 2026-09-01 | Source-checked |
| 64 | Qwen 3.7 PlusAlibaba | 2026-06-02 | 17.53% | 2026-09-01 | Source-checked |
| 65 | GLM-5.2 (Non-reasoning)Zhipu | 2026-06-16 | 16.7% | 2026-09-01 | Source-checked |
| 66 | Qwen3.6 27B (Reasoning)Alibaba | 2026-04-22 | 16.7% | 2026-09-01 | Source-checked |
| 67 | A.X-K2Other | 2026-08-12 | 16.08% | 2026-09-01 | Source-checked |
| 68 | GPT-5.1OpenAI | 2025-11-13 | 15.88% | 2026-09-01 | Source-checked |
| 69 | MiniMax M3MiniMax | 2026-05-31 | 15.26% | 2026-09-01 | Source-checked |
| 70 | Qwen3.5 122B A10B (Reasoning)Alibaba | 2026-02-24 | 15.26% | 2026-09-01 | Source-checked |
| 71 | Mistral Medium 3.5Mistral | 2026-04-29 | 15.05% | 2026-09-01 | Source-checked |
| 72 | Gemma 4 31B ITGoogle | 2026-03-01 | 14.85% | 2026-09-01 | Source-checked |
| 73 | GPT-5.5 (Non-reasoning)OpenAI | 2026-04-23 | 14.85% | 2026-09-01 | Source-checked |
| 74 | Granite 4.2 30BOther | 2026-08-25 | 14.43% | 2026-09-01 | Source-checked |
| 75 | Kimi K2.5Moonshot | 2026-01-27 | 14.23% | 2026-09-01 | Source-checked |
| 76 | Nemotron 3 Ultra 550B A55B (Reasoning)NVIDIA | 2026-06-04 | 14.23% | 2026-09-01 | Source-checked |
| 77 | GLM-5.1Zhipu | 2026-04-07 | 13.61% | 2026-09-01 | Source-checked |
| 78 | Grok Build 0.1 0616xAI | 2026-06-16 | 13.4% | 2026-09-01 | Source-checked |
| 79 | Qwen3.5 397B A17B (Reasoning)Alibaba | 2026-02-16 | 13.4% | 2026-09-01 | Source-checked |
| 80 | LongCat-2.0Meituan | 2026-06-29 | 13.2% | 2026-09-01 | Source-checked |
| 81 | Agnes 2.5 Pro AlphaOther | 2026-07-24 | 12.37% | 2026-09-01 | Source-checked |
| 82 | Grok 4.3xAI | 2026-04-30 | 12.37% | 2026-09-01 | Source-checked |
| 83 | GLM-4.7 (Reasoning)Zhipu | 2025-12-22 | 12.16% | 2026-09-01 | Source-checked |
| 84 | Step 3.7 FlashStepFun | 2026-06-01 | 11.96% | 2026-09-01 | Source-checked |
| 85 | GPT-5.5 Instant (June 2026)OpenAI | 2026-06-25 | 11.75% | 2026-09-01 | Source-checked |
| 86 | Qwen3.7 MaxAlibaba | 2026-05-19 | 11.75% | 2026-09-01 | Source-checked |
| 87 | K-EXAONE 2.0 0803Other | 2026-08-12 | 11.55% | 2026-09-01 | Source-checked |
| 88 | MiMo-V2.5-ProXiaomi | 2026-04-22 | 9.9% | 2026-09-01 | Source-checked |
| 89 | MiniMax M2.7MiniMax | 2026-04-15 | 9.9% | 2026-09-01 | Source-checked |
| 90 | Gemini 3.1 Flash LiteGoogle | 2026-05-07 | 9.69% | 2026-09-01 | Source-checked |
| 91 | Qwen3.6 27B (Non-reasoning)Alibaba | 2026-04-22 | 9.28% | 2026-09-01 | Source-checked |
| 92 | Qwen3.6 35B A3B (Reasoning)Alibaba | 2026-04-16 | 9.28% | 2026-09-01 | Source-checked |
| 93 | Nemotron 3.5 LightningNVIDIA | 2026-08-11 | 8.87% | 2026-09-01 | Source-checked |
| 94 | MiMo-V2.5Xiaomi | 2026-04-22 | 8.66% | 2026-09-01 | Source-checked |
| 95 | Grok 4.3 (Non-reasoning)xAI | 2026-04-30 | 8.04% | 2026-09-01 | Source-checked |
| 96 | Granite 4.2 8BOther | 2026-08-25 | 7.63% | 2026-09-01 | Source-checked |
| 97 | North Mini CodeCohere | 2026-06-09 | 6.39% | 2026-09-01 | Source-checked |
| 98 | Command ACohere | 2026-03-15 | 5.98% | 2026-09-01 | Source-checked |
| 99 | Mistral Large 3Mistral | 2025-12-02 | 5.77% | 2026-09-01 | Source-checked |
| 100 | Granite 4.2 3BOther | 2026-08-25 | 5.57% | 2026-09-01 | Source-checked |
| 101 | HyperNova 60B 2605Multiverse | 2026-05-06 | 5.57% | 2026-09-01 | Source-checked |
| 102 | Qwen3.6 35B A3B (Non-reasoning)Alibaba | 2026-04-16 | 5.36% | 2026-09-01 | Source-checked |
| 103 | KAT Coder Pro V2Other | 2026-03-27 | 4.54% | 2026-09-01 | Source-checked |
| 104 | Llama 4 MaverickMeta | 2026-04-05 | 3.71% | 2026-09-01 | Source-checked |
| 105 | Llama 4 ScoutMeta | 2026-04-05 | 3.3% | 2026-09-01 | Source-checked |
| 106 | Ling 2.6 FlashOther | 2026-04-21 | 2.89% | 2026-09-01 | Source-checked |
Live — AA τ³-Bench Banking (Index v4.1, 14% weight) uses a harder dual-control scale than retired τ². VerdictPal maps only current `tau_banking` / `tau3_banking` rows onto this page; older τ² high-scale receipts are not mixed in.
No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
Dual-control agent-user simulation for banking tasks: 97 scenarios with knowledge retrieval, backend database state evaluation, pass@1. 14% weight in AA Intelligence Index v4.1.
A strong τ³-Bench Banking result says nothing about:
Backend database state evaluation, pass@1. Task format: 97 banking scenarios; 5 repeats; dual-control simulation with tool use and database checks.
τ³-Bench Banking is currently marked Active in the atlas.
Contamination risk for τ³-Bench Banking is graded Low contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.