Why this benchmark is useful
It stresses abstraction and multi-step inference. Use it to find models that can hold a problem together after the prompt stops looking like a memorized exam.
Benchmarks / Reasoning
Instruction-Following Benchmark
Precise instruction-following on novel constraints — tests whether models can obey complex, unseen formatting rules without drifting.
It stresses abstraction and multi-step inference. Use it to find models that can hold a problem together after the prompt stops looking like a memorized exam.
Read IFBench as a defunct signal with medium contamination risk. Compare models only when the source uses the same harness, prompting setup, sampling policy, and score unit.
Normalized to this benchmark's axis (0–100). Open the table for raw units.
Scores use this benchmark's own unit and axis, not a universal quality score.
| # | Model | Release date | Score | Provenance | Trust |
|---|---|---|---|---|---|
| 1 | MiniMax M3MiniMax | 2026-05-31 | 82.86% | 2026-07-21 | Source-checked |
| 2 | Nemotron 3 Ultra 550B A55B (Reasoning)NVIDIA | 2026-06-04 | 81.36% | 2026-07-21 | Source-checked |
| 3 | Grok 4.3xAI | 2026-04-30 | 81.29% | 2026-07-21 | Source-checked |
| 4 | Grok 4.20 ReasoningxAI | 2026-03-05 | 81.22% | 2026-07-21 | Source-checked |
| 5 | Qwen3.7 MaxAlibaba | 2026-05-19 | 80.54% | 2026-07-21 | Source-checked |
| 6 | MiMo-V2.5-ProXiaomi | 2026-04-22 | 79.86% | 2026-07-21 | Source-checked |
| 7 | Qwen 3.7 PlusAlibaba | 2026-06-02 | 77.96% | 2026-07-21 | Source-checked |
| 8 | Gemini 3.1 Flash LiteGoogle | 2026-05-07 | 77.21% | 2026-07-21 | Source-checked |
| 9 | Gemini 3.1 Pro PreviewGoogle | 2026-02-19 | 77.14% | 2026-07-21 | Source-checked |
| 10 | DeepSeek V4 Pro (Reasoning, Max Effort)DeepSeek | 2026-04-24 | 76.46% | 2026-07-21 | Source-checked |
| 11 | Gemini 3.5 FlashGoogle | 2026-05-19 | 76.33% | 2026-07-21 | Source-checked |
| 12 | GLM-5.1Zhipu | 2026-04-07 | 76.26% | 2026-07-21 | Source-checked |
| 13 | Kimi K2.6Moonshot | 2026-04-20 | 75.99% | 2026-07-21 | Source-checked |
| 14 | GPT-5.4 NanoOpenAI | 2026-03-17 | 75.92% | 2026-07-21 | Source-checked |
| 15 | Muse SparkOther | 2026-01-01 | 75.92% | 2026-07-21 | Source-checked |
| 16 | GPT 5.5OpenAI | 2026-04-23 | 75.85% | 2026-07-21 | Source-checked |
| 17 | MiniMax M2.7MiniMax | 2026-04-15 | 75.71% | 2026-07-21 | Source-checked |
| 18 | Gemma 4 31B ITGoogle | 2026-03-01 | 75.58% | 2026-07-21 | Source-checked |
| 19 | GPT-5.2OpenAI | 2025-12-11 | 75.44% | 2026-07-21 | Source-checked |
| 20 | GPT-5.3 Codex (xhigh)OpenAI | 2026-02-05 | 75.37% | 2026-07-21 | Source-checked |
| 21 | Gemini 3.5 Flash (medium)Google | 2026-05-19 | 74.56% | 2026-07-21 | Source-checked |
| 22 | GPT 5.4OpenAI | 2026-03-05 | 73.95% | 2026-07-21 | Source-checked |
| 23 | Gemma 4 12B (Reasoning)Google | 2026-06-07 | 73.54% | 2026-07-21 | Source-checked |
| 24 | GLM-5.2Zhipu | 2026-06-13 | 73.33% | 2026-07-21 | Source-checked |
| 25 | GPT-5.4 MiniOpenAI | 2026-03-17 | 73.27% | 2026-07-21 | Source-checked |
| 26 | GPT-5.1OpenAI | 2025-11-13 | 72.86% | 2026-07-21 | Source-checked |
| 27 | GPT-5.6 SolOpenAI | 2026-07-09 | 72.65% | 2026-07-21 | Source-checked |
| 28 | GPT-5.5 (high)OpenAI | 2026-04-23 | 71.63% | 2026-07-21 | Source-checked |
| 29 | MiniMax M2.5MiniMax | 2026-04-01 | 71.63% | 2026-07-21 | Source-checked |
| 30 | OpenAI o3OpenAI | 2025-04-16 | 71.43% | 2026-07-21 | Source-checked |
| 31 | GPT-5.6 TerraOpenAI | 2026-07-09 | 71.22% | 2026-07-21 | Source-checked |
| 32 | GPT-5.5 (medium)OpenAI | 2026-04-23 | 70.95% | 2026-07-21 | Source-checked |
| 33 | Gemini 3 ProGoogle | 2025-11-18 | 70.41% | 2026-07-21 | Source-checked |
| 34 | OpenAI o1OpenAI | 2024-09-12 | 70.34% | 2026-07-21 | Source-checked |
| 35 | Kimi K2.5Moonshot | 2026-01-27 | 70.2% | 2026-07-21 | Source-checked |
| 36 | Step 3.7 FlashStepFun | 2026-06-01 | 67.28% | 2026-07-21 | Source-checked |
| 37 | HyperNova 60B 2605Multiverse | 2026-05-06 | 66.46% | 2026-07-21 | Source-checked |
| 38 | GPT-5.5 (low)OpenAI | 2026-04-23 | 64.35% | 2026-07-21 | Source-checked |
| 39 | Claude Fable 5Anthropic | 2026-06-09 | 63.47% | 2026-07-21 | Source-checked |
| 40 | Kimi K2.7 CodeMoonshot | 2026-06-12 | 63.13% | 2026-07-21 | Source-checked |
| 41 | Claude Opus 4.8Anthropic | 2026-05-28 | 62.24% | 2026-07-21 | Source-checked |
| 42 | Claude Opus 4.7Anthropic | 2026-04-16 | 58.64% | 2026-07-21 | Source-checked |
| 43 | North Mini CodeCohere | 2026-06-09 | 57.55% | 2026-07-21 | Source-checked |
| 44 | Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)Anthropic | 2026-02-17 | 56.6% | 2026-07-21 | Source-checked |
| 45 | LFM2.5-8B-A1BLiquid | 2026-06-08 | 55.65% | 2026-07-21 | Source-checked |
| 46 | Gemini 3 Flash PreviewGoogle | 2025-12-17 | 55.1% | 2026-07-21 | Source-checked |
| 47 | Claude Opus 4.6 (Adaptive Reasoning, Max Effort)Anthropic | 2026-02-05 | 53.13% | 2026-07-21 | Source-checked |
| 48 | MiniCPM5-1B (Reasoning)OpenBMB | 2026-06-04 | 49.32% | 2026-07-21 | Source-checked |
| 49 | Gemma 4 12B (Non-reasoning)Google | 2026-06-10 | 45.17% | 2026-07-21 | Source-checked |
| 50 | Claude Opus 4.6Anthropic | 2026-02-05 | 44.63% | 2026-07-21 | Source-checked |
| 51 | Claude Opus 4.7 (Non-reasoning, High Effort)Anthropic | 2026-04-16 | 43.61% | 2026-07-21 | Source-checked |
| 52 | Claude Opus 4.5Anthropic | 2025-11-24 | 42.99% | 2026-07-21 | Source-checked |
| 53 | Llama 4 MaverickMeta | 2026-04-05 | 42.99% | 2026-07-21 | Source-checked |
| 54 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 41.16% | 2026-07-21 | Source-checked |
| 55 | DeepSeek R1DeepSeek | 2025-01-20 | 39.59% | 2026-07-21 | Source-checked |
| 56 | Llama 4 ScoutMeta | 2026-04-05 | 39.52% | 2026-07-21 | Source-checked |
| 57 | Mistral Large 3Mistral | 2025-12-02 | 36.19% | 2026-07-21 | Source-checked |
Archived — dropped from AA Intelligence Index v4.1 for saturation. AA continues to run it; VerdictPal keeps the glossary entry as an honesty receipt.
No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.
The Pack · Editorial newsletter
One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.