Why this benchmark is useful
Historical instruction-following stress test on novel format constraints. Saturated and dropped from AA Intelligence Index v4.1 — glossary honesty, not current frontier ranking.
Instruction-Following Benchmark
Precise instruction-following on novel constraints — tests whether models can obey complex, unseen formatting rules without drifting.
Historical instruction-following stress test on novel format constraints. Saturated and dropped from AA Intelligence Index v4.1 — glossary honesty, not current frontier ranking.
Constraint-satisfaction accuracy only — a perfectly formatted wrong answer still scores. Do not use for separating today's frontier; AA still publishes runs but no longer weights the suite.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | Grok 4.3 (medium)xAI | 2026-04-30 | 83.33% | 2026-09-01 | Source-checked |
| 2 | Grok 4.20 0309 (Reasoning)xAI | 2026-03-10 | 82.93% | 2026-09-01 | Source-checked |
| 3 | MiniMax M3MiniMax | 2026-05-31 | 82.86% | 2026-09-01 | Source-checked |
| 4 | Nemotron 3 Ultra 550B A55B (Reasoning)NVIDIA | 2026-06-04 | 81.36% | 2026-09-01 | Source-checked |
| 5 | Grok 4.3xAI | 2026-04-30 | 81.29% | 2026-09-01 | Source-checked |
| 6 | Grok 4.20 ReasoningxAI | 2026-03-05 | 81.22% | 2026-09-01 | Source-checked |
| 7 | Grok 4.3 (low)xAI | 2026-04-30 | 80.95% | 2026-09-01 | Source-checked |
| 8 | Qwen3.7 MaxAlibaba | 2026-05-19 | 80.54% | 2026-09-01 | Source-checked |
| 9 | MiMo-V2.5-ProXiaomi | 2026-04-22 | 79.86% | 2026-09-01 | Source-checked |
| 10 | Qwen3.5 397B A17B (Reasoning)Alibaba | 2026-02-16 | 78.78% | 2026-09-01 | Source-checked |
| 11 | Gemini 3 Flash Preview (Reasoning)Google | 2025-12-17 | 77.96% | 2026-09-01 | Source-checked |
| 12 | Qwen 3.7 PlusAlibaba | 2026-06-02 | 77.96% | 2026-09-01 | Source-checked |
| 13 | GPT-5.2 Codex (xhigh)OpenAI | 2025-12-11 | 77.62% | 2026-09-01 | Source-checked |
| 14 | Gemini 3.1 Flash LiteGoogle | 2026-05-07 | 77.21% | 2026-09-01 | Source-checked |
| 15 | Gemini 3.1 Pro PreviewGoogle | 2026-02-19 | 77.14% | 2026-09-01 | Source-checked |
| 16 | Qwen3.6 Max PreviewAlibaba | 2026-04-20 | 76.6% | 2026-09-01 | Source-checked |
| 17 | Gemini 3.5 FlashGoogle | 2026-05-19 | 76.33% | 2026-09-01 | Source-checked |
| 18 | GLM-5.1Zhipu | 2026-04-07 | 76.26% | 2026-09-01 | Source-checked |
| 19 | Kimi K2.6Moonshot | 2026-04-20 | 75.99% | 2026-09-01 | Source-checked |
| 20 | GPT-5.4 NanoOpenAI | 2026-03-17 | 75.92% | 2026-09-01 | Source-checked |
| 21 | Muse SparkOther | 2026-01-01 | 75.92% | 2026-09-01 | Source-checked |
| 22 | GPT 5.5OpenAI | 2026-04-23 | 75.85% | 2026-09-01 | Source-checked |
| 23 | MiniMax M2.7MiniMax | 2026-04-15 | 75.71% | 2026-09-01 | Source-checked |
| 24 | Qwen3.5 122B A10B (Reasoning)Alibaba | 2026-02-24 | 75.71% | 2026-09-01 | Source-checked |
| 25 | Gemma 4 31B ITGoogle | 2026-03-01 | 75.58% | 2026-09-01 | Source-checked |
| 26 | Qwen3.5 27B (Reasoning)Alibaba | 2026-02-24 | 75.58% | 2026-09-01 | Source-checked |
| 27 | GPT-5.2OpenAI | 2025-12-11 | 75.44% | 2026-09-01 | Source-checked |
| 28 | GPT-5.3 Codex (xhigh)OpenAI | 2026-02-05 | 75.37% | 2026-09-01 | Source-checked |
| 29 | Qwen3.6 PlusAlibaba | 2026-04-02 | 75.17% | 2026-09-01 | Source-checked |
| 30 | Gemini 3.5 Flash (medium)Google | 2026-05-19 | 74.56% | 2026-09-01 | Source-checked |
| 31 | GPT-5 Codex (high)OpenAI | 2025-09-23 | 74.15% | 2026-09-01 | Source-checked |
| 32 | GPT 5.4OpenAI | 2026-03-05 | 73.95% | 2026-09-01 | Source-checked |
| 33 | Gemma 4 12B (Reasoning)Google | 2026-06-07 | 73.54% | 2026-09-01 | Source-checked |
| 34 | DeepSeek V4 Flash (Reasoning, High Effort)DeepSeek | 2026-04-24 | 73.47% | 2026-09-01 | Source-checked |
| 35 | GLM-5.2Zhipu | 2026-06-13 | 73.33% | 2026-09-01 | Source-checked |
| 36 | GPT-5.4 MiniOpenAI | 2026-03-17 | 73.27% | 2026-09-01 | Source-checked |
| 37 | GLM-5-TurboZhipu | 2026-03-15 | 73.2% | 2026-09-01 | Source-checked |
| 38 | GPT-5 (high)OpenAI | 2025-08-07 | 73.06% | 2026-09-01 | Source-checked |
| 39 | GPT-5.1OpenAI | 2025-11-13 | 72.86% | 2026-09-01 | Source-checked |
| 40 | GPT-5.6 SolOpenAI | 2026-07-09 | 72.65% | 2026-09-01 | Source-checked |
| 41 | GLM-5 (Reasoning)Zhipu | 2026-02-11 | 72.31% | 2026-09-01 | Source-checked |
| 42 | MiMo-V2-Flash (Feb 2026)Xiaomi | 2025-12-16 | 71.84% | 2026-09-01 | Source-checked |
| 43 | GPT-5.5 (high)OpenAI | 2026-04-23 | 71.63% | 2026-09-01 | Source-checked |
| 44 | MiniMax M2.5MiniMax | 2026-04-01 | 71.63% | 2026-09-01 | Source-checked |
| 45 | GPT-5.5 Instant (May 2026)OpenAI | 2026-05-05 | 71.5% | 2026-09-01 | Source-checked |
| 46 | OpenAI o3OpenAI | 2025-04-16 | 71.43% | 2026-09-01 | Source-checked |
| 47 | DeepSeek V4 Pro (Reasoning, High Effort)DeepSeek | 2026-04-24 | 71.29% | 2026-09-01 | Source-checked |
| 48 | GPT-5.6 TerraOpenAI | 2026-07-09 | 71.22% | 2026-09-01 | Source-checked |
| 49 | GPT-5.5 (medium)OpenAI | 2026-04-23 | 70.95% | 2026-09-01 | Source-checked |
| 50 | Qwen3 Max ThinkingAlibaba | 2026-01-26 | 70.75% | 2026-09-01 | Source-checked |
| 51 | GPT-5 (medium)OpenAI | 2025-08-07 | 70.61% | 2026-09-01 | Source-checked |
| 52 | Gemini 3 ProGoogle | 2025-11-18 | 70.41% | 2026-09-01 | Source-checked |
| 53 | OpenAI o1OpenAI | 2024-09-12 | 70.34% | 2026-09-01 | Source-checked |
| 54 | Kimi K2.5Moonshot | 2026-01-27 | 70.2% | 2026-09-01 | Source-checked |
| 55 | GPT-5.1 Codex (high)OpenAI | 2025-11-13 | 70% | 2026-09-01 | Source-checked |
| 56 | MiniMax-M2.1MiniMax | 2025-12-23 | 69.86% | 2026-09-01 | Source-checked |
| 57 | MiMo-V2-ProXiaomi | 2026-03-18 | 68.84% | 2026-09-01 | Source-checked |
| 58 | Mistral Medium 3.5Mistral | 2026-04-29 | 68.78% | 2026-09-01 | Source-checked |
| 59 | Kimi K2 ThinkingMoonshot | 2025-11-06 | 68.1% | 2026-09-01 | Source-checked |
| 60 | GLM-4.7 (Reasoning)Zhipu | 2025-12-22 | 67.89% | 2026-09-01 | Source-checked |
| 61 | Qwen3.6 27B (Reasoning)Alibaba | 2026-04-22 | 67.55% | 2026-09-01 | Source-checked |
| 62 | MiMo-V2-Omni-0327Xiaomi | 2026-03-27 | 67.35% | 2026-09-01 | Source-checked |
| 63 | Step 3.7 FlashStepFun | 2026-06-01 | 67.28% | 2026-09-01 | Source-checked |
| 64 | MiMo-V2.5Xiaomi | 2026-04-22 | 67.14% | 2026-09-01 | Source-checked |
| 65 | KAT Coder Pro V2Other | 2026-03-27 | 66.67% | 2026-09-01 | Source-checked |
| 66 | GPT-5 (low)OpenAI | 2025-08-07 | 66.6% | 2026-09-01 | Source-checked |
| 67 | HyperNova 60B 2605Multiverse | 2026-05-06 | 66.46% | 2026-09-01 | Source-checked |
| 68 | Nex-N2-ProOther | 2026-06-02 | 66.19% | 2026-09-01 | Source-checked |
| 69 | GPT-5.4 (low)OpenAI | 2026-03-05 | 65.92% | 2026-09-01 | Source-checked |
| 70 | GPT-5.2 (medium)OpenAI | 2025-12-11 | 65.24% | 2026-09-01 | Source-checked |
| 71 | GPT-5.5 (low)OpenAI | 2026-04-23 | 64.35% | 2026-09-01 | Source-checked |
| 72 | Qwen3.6 35B A3B (Reasoning)Alibaba | 2026-04-16 | 64.35% | 2026-09-01 | Source-checked |
| 73 | MiMo-V2-Flash (Reasoning)Xiaomi | 2025-12-16 | 64.22% | 2026-09-01 | Source-checked |
| 74 | Claude Fable 5Anthropic | 2026-06-09 | 63.47% | 2026-09-01 | Source-checked |
| 75 | Nemotron 3 Nano Omni 30B A3B ReasoningNVIDIA | 2026-04-29 | 63.2% | 2026-09-01 | Source-checked |
| 76 | Hy3-preview (Reasoning)Other | 2026-04-23 | 63.13% | 2026-09-01 | Source-checked |
| 77 | Kimi K2.7 CodeMoonshot | 2026-06-12 | 63.13% | 2026-09-01 | Source-checked |
| 78 | Claude Opus 4.8Anthropic | 2026-05-28 | 62.24% | 2026-09-01 | Source-checked |
| 79 | GLM 5V Turbo (Reasoning)Zhipu | 2026-04-01 | 61.09% | 2026-09-01 | Source-checked |
| 80 | DeepSeek V3.2 (Reasoning)DeepSeek | 2025-12-01 | 60.68% | 2026-09-01 | Source-checked |
| 81 | DiffusionGemma 26B A4BGoogle | 2026-06-10 | 59.46% | 2026-09-01 | Source-checked |
| 82 | Claude Opus 4.7Anthropic | 2026-04-16 | 58.64% | 2026-09-01 | Source-checked |
| 83 | Claude Opus 4.5 (Reasoning)Anthropic | 2025-11-24 | 57.96% | 2026-09-01 | Source-checked |
| 84 | EXAONE 4.5 33BOther | 2026-04-09 | 57.96% | 2026-09-01 | Source-checked |
| 85 | North Mini CodeCohere | 2026-06-09 | 57.55% | 2026-09-01 | Source-checked |
| 86 | Ling 2.6 FlashOther | 2026-04-21 | 57.41% | 2026-09-01 | Source-checked |
| 87 | Claude 4.5 Sonnet (Reasoning)Anthropic | 2025-09-29 | 57.28% | 2026-09-01 | Source-checked |
| 88 | Ling-2.6-1TOther | 2026-04-23 | 56.87% | 2026-09-01 | Source-checked |
| 89 | Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)Anthropic | 2026-02-17 | 56.6% | 2026-09-01 | Source-checked |
| 90 | LFM2.5-8B-A1BLiquid | 2026-06-08 | 55.65% | 2026-09-01 | Source-checked |
| 91 | Claude 4.1 Opus (Reasoning)Anthropic | 2025-08-05 | 55.44% | 2026-09-01 | Source-checked |
| 92 | GLM-5 (Non-reasoning)Zhipu | 2026-02-11 | 55.17% | 2026-09-01 | Source-checked |
| 93 | Gemini 3 Flash PreviewGoogle | 2025-12-17 | 55.1% | 2026-09-01 | Source-checked |
| 94 | Claude 4 Opus (Reasoning)Anthropic | 2025-05-22 | 53.74% | 2026-09-01 | Source-checked |
| 95 | Grok 4xAI | 2025-07-10 | 53.67% | 2026-09-01 | Source-checked |
| 96 | MiMo-V2-OmniXiaomi | 2026-03-19 | 53.54% | 2026-09-01 | Source-checked |
| 97 | Claude Opus 4.6 (Adaptive Reasoning, Max Effort)Anthropic | 2026-02-05 | 53.13% | 2026-09-01 | Source-checked |
| 98 | Qwen3.5 397B A17B (Non-reasoning)Alibaba | 2026-02-16 | 51.63% | 2026-09-01 | Source-checked |
| 99 | Gemini 3 Pro Preview (low)Google | 2025-11-18 | 49.66% | 2026-09-01 | Source-checked |
| 100 | Grok 4.20 0309 v2 (Non-reasoning)xAI | 2026-04-07 | 49.32% | 2026-09-01 | Source-checked |
| 101 | MiniCPM5-1B (Reasoning)OpenBMB | 2026-06-04 | 49.32% | 2026-09-01 | Source-checked |
| 102 | Hy3-preview (Non-reasoning)Other | 2026-04-23 | 47.96% | 2026-09-01 | Source-checked |
| 103 | Grok 4.3 (Non-reasoning)xAI | 2026-04-30 | 47.62% | 2026-09-01 | Source-checked |
| 104 | Gemini 3.5 Flash (minimal)Google | 2026-05-19 | 47.28% | 2026-09-01 | Source-checked |
| 105 | DeepSeek V4 Flash (Non-reasoning)DeepSeek | 2026-04-24 | 47.21% | 2026-09-01 | Source-checked |
| 106 | GPT-5.5 (Non-reasoning)OpenAI | 2026-04-23 | 46.12% | 2026-09-01 | Source-checked |
| 107 | DeepSeek V4 Pro (Non-reasoning)DeepSeek | 2026-04-24 | 45.78% | 2026-09-01 | Source-checked |
| 108 | Qwen3.6 27B (Non-reasoning)Alibaba | 2026-04-22 | 45.71% | 2026-09-01 | Source-checked |
| 109 | Gemma 4 12B (Non-reasoning)Google | 2026-06-10 | 45.17% | 2026-09-01 | Source-checked |
| 110 | Claude Opus 4.6Anthropic | 2026-02-05 | 44.63% | 2026-09-01 | Source-checked |
| 111 | Ring-2.6-1TOther | 2026-05-08 | 44.56% | 2026-09-01 | Source-checked |
| 112 | Granite 4.1 30BOther | 2026-04-29 | 44.35% | 2026-09-01 | Source-checked |
| 113 | Kimi K2.6 (Non-reasoning)Other | 2026-04-20 | 44.29% | 2026-09-01 | Source-checked |
| 114 | Claude Opus 4.7 (Non-reasoning, High Effort)Anthropic | 2026-04-16 | 43.61% | 2026-09-01 | Source-checked |
| 115 | Claude Opus 4.5Anthropic | 2025-11-24 | 42.99% | 2026-09-01 | Source-checked |
| 116 | Llama 4 MaverickMeta | 2026-04-05 | 42.99% | 2026-09-01 | Source-checked |
| 117 | MiMo-V2.5-Pro (Non-reasoning)Xiaomi | 2026-04-22 | 42.72% | 2026-09-01 | Source-checked |
| 118 | Claude Sonnet 4.6 (Non-reasoning, Low Effort)Anthropic | 2026-02-17 | 42.38% | 2026-09-01 | Source-checked |
| 119 | JT-35B-FlashOther | 2026-05-14 | 42.04% | 2026-09-01 | Source-checked |
| 120 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 41.16% | 2026-09-01 | Source-checked |
| 121 | DeepSeek R1DeepSeek | 2025-01-20 | 39.59% | 2026-09-01 | Source-checked |
| 122 | Llama 4 ScoutMeta | 2026-04-05 | 39.52% | 2026-09-01 | Source-checked |
| 123 | JT-MINIOther | 2026-04-15 | 36.67% | 2026-09-01 | Source-checked |
| 124 | Mistral Large 3Mistral | 2025-12-02 | 36.19% | 2026-09-01 | Source-checked |
| 125 | Qwen3.6 35B A3B (Non-reasoning)Alibaba | 2026-04-16 | 36.19% | 2026-09-01 | Source-checked |
| 126 | MiniCPM5-1B (Non-reasoning)OpenBMB | 2026-05-25 | 35.17% | 2026-09-01 | Source-checked |
| 127 | MiniCPM-V 4.6 1.3BOpenBMB | 2026-05-11 | 26.73% | 2026-09-01 | Source-checked |
Archived — dropped from AA Intelligence Index v4.1 for saturation. AA continues to run it; VerdictPal keeps the glossary entry as an honesty receipt.
No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.
Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.
Precise instruction-following on novel constraints — tests whether models can obey complex, unseen formatting rules without drifting.
A strong IFBench result says nothing about:
Constraint satisfaction accuracy. AA still publishes results but no longer weights IFBench in the composite index. Task format: Novel instruction-following tasks with strict output constraints; accuracy on unseen rule sets.
IFBench is currently marked Saturated in the atlas.
Contamination risk for IFBench is graded Medium contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.
The Pack · Editorial newsletter
One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.