Benchmarks / Reasoning

IFBench

Instruction-Following Benchmark

Precise instruction-following on novel constraints — tests whether models can obey complex, unseen formatting rules without drifting.

What this does not measure
  • Frontier discrimination today — Artificial Analysis removed IFBench from Intelligence Index v4.1 because frontier models saturated it.
  • Factual accuracy — a perfectly formatted wrong answer still scores.
  • Real user tasks — synthetic constraint tests, not production workflows.
Analysis

Why this benchmark is useful

Historical instruction-following stress test on novel format constraints. Saturated and dropped from AA Intelligence Index v4.1 — glossary honesty, not current frontier ranking.

Scope

Coverage map

Task family
Reasoning
Format
Novel instruction-following tasks with strict output constraints; accuracy on unseen rule sets.
Scoring
Constraint satisfaction accuracy. AA still publishes results but no longer weights IFBench in the composite index.
Maintainer
Artificial Analysis
Reading guide

How to read the scores

Constraint-satisfaction accuracy only — a perfectly formatted wrong answer still scores. Do not use for separating today's frontier; AA still publishes runs but no longer weights the suite.

Blind spots

What it does not cover

  • Frontier discrimination today — Artificial Analysis removed IFBench from Intelligence Index v4.1 because frontier models saturated it.
  • Factual accuracy — a perfectly formatted wrong answer still scores.
  • Real user tasks — synthetic constraint tests, not production workflows.
Scores

Evidence ledger

127 rows
127127 rows
83.33%best score
127source-checked
1sources
2026-09-01source date
127

api · api

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. Grok 4.3 (medium)83.33% ·
  2. Grok 4.20 0309 (Reasoning)82.93% ·
  3. MiniMax M382.86% ·
  4. Nemotron 3 Ultra 550B A55B (Reasoning)81.36% ·
  5. Grok 4.381.29% ·

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1Grok 4.3 (medium)xAI2026-04-3083.33%
2026-09-01
Source-checked
2Grok 4.20 0309 (Reasoning)xAI2026-03-1082.93%
2026-09-01
Source-checked
3MiniMax M3MiniMax2026-05-3182.86%
2026-09-01
Source-checked
4Nemotron 3 Ultra 550B A55B (Reasoning)NVIDIA2026-06-0481.36%
2026-09-01
Source-checked
5Grok 4.3xAI2026-04-3081.29%
2026-09-01
Source-checked
6Grok 4.20 ReasoningxAI2026-03-0581.22%
2026-09-01
Source-checked
7Grok 4.3 (low)xAI2026-04-3080.95%
2026-09-01
Source-checked
8Qwen3.7 MaxAlibaba2026-05-1980.54%
2026-09-01
Source-checked
9MiMo-V2.5-ProXiaomi2026-04-2279.86%
2026-09-01
Source-checked
10Qwen3.5 397B A17B (Reasoning)Alibaba2026-02-1678.78%
2026-09-01
Source-checked
11Gemini 3 Flash Preview (Reasoning)Google2025-12-1777.96%
2026-09-01
Source-checked
12Qwen 3.7 PlusAlibaba2026-06-0277.96%
2026-09-01
Source-checked
13GPT-5.2 Codex (xhigh)OpenAI2025-12-1177.62%
2026-09-01
Source-checked
14Gemini 3.1 Flash LiteGoogle2026-05-0777.21%
2026-09-01
Source-checked
15Gemini 3.1 Pro PreviewGoogle2026-02-1977.14%
2026-09-01
Source-checked
16Qwen3.6 Max PreviewAlibaba2026-04-2076.6%
2026-09-01
Source-checked
17Gemini 3.5 FlashGoogle2026-05-1976.33%
2026-09-01
Source-checked
18GLM-5.1Zhipu2026-04-0776.26%
2026-09-01
Source-checked
19Kimi K2.6Moonshot2026-04-2075.99%
2026-09-01
Source-checked
20GPT-5.4 NanoOpenAI2026-03-1775.92%
2026-09-01
Source-checked
21Muse SparkOther2026-01-0175.92%
2026-09-01
Source-checked
22GPT 5.5OpenAI2026-04-2375.85%
2026-09-01
Source-checked
23MiniMax M2.7MiniMax2026-04-1575.71%
2026-09-01
Source-checked
24Qwen3.5 122B A10B (Reasoning)Alibaba2026-02-2475.71%
2026-09-01
Source-checked
25Gemma 4 31B ITGoogle2026-03-0175.58%
2026-09-01
Source-checked
26Qwen3.5 27B (Reasoning)Alibaba2026-02-2475.58%
2026-09-01
Source-checked
27GPT-5.2OpenAI2025-12-1175.44%
2026-09-01
Source-checked
28GPT-5.3 Codex (xhigh)OpenAI2026-02-0575.37%
2026-09-01
Source-checked
29Qwen3.6 PlusAlibaba2026-04-0275.17%
2026-09-01
Source-checked
30Gemini 3.5 Flash (medium)Google2026-05-1974.56%
2026-09-01
Source-checked
31GPT-5 Codex (high)OpenAI2025-09-2374.15%
2026-09-01
Source-checked
32GPT 5.4OpenAI2026-03-0573.95%
2026-09-01
Source-checked
33Gemma 4 12B (Reasoning)Google2026-06-0773.54%
2026-09-01
Source-checked
34DeepSeek V4 Flash (Reasoning, High Effort)DeepSeek2026-04-2473.47%
2026-09-01
Source-checked
35GLM-5.2Zhipu2026-06-1373.33%
2026-09-01
Source-checked
36GPT-5.4 MiniOpenAI2026-03-1773.27%
2026-09-01
Source-checked
37GLM-5-TurboZhipu2026-03-1573.2%
2026-09-01
Source-checked
38GPT-5 (high)OpenAI2025-08-0773.06%
2026-09-01
Source-checked
39GPT-5.1OpenAI2025-11-1372.86%
2026-09-01
Source-checked
40GPT-5.6 SolOpenAI2026-07-0972.65%
2026-09-01
Source-checked
41GLM-5 (Reasoning)Zhipu2026-02-1172.31%
2026-09-01
Source-checked
42MiMo-V2-Flash (Feb 2026)Xiaomi2025-12-1671.84%
2026-09-01
Source-checked
43GPT-5.5 (high)OpenAI2026-04-2371.63%
2026-09-01
Source-checked
44MiniMax M2.5MiniMax2026-04-0171.63%
2026-09-01
Source-checked
45GPT-5.5 Instant (May 2026)OpenAI2026-05-0571.5%
2026-09-01
Source-checked
46OpenAI o3OpenAI2025-04-1671.43%
2026-09-01
Source-checked
47DeepSeek V4 Pro (Reasoning, High Effort)DeepSeek2026-04-2471.29%
2026-09-01
Source-checked
48GPT-5.6 TerraOpenAI2026-07-0971.22%
2026-09-01
Source-checked
49GPT-5.5 (medium)OpenAI2026-04-2370.95%
2026-09-01
Source-checked
50Qwen3 Max ThinkingAlibaba2026-01-2670.75%
2026-09-01
Source-checked
51GPT-5 (medium)OpenAI2025-08-0770.61%
2026-09-01
Source-checked
52Gemini 3 ProGoogle2025-11-1870.41%
2026-09-01
Source-checked
53OpenAI o1OpenAI2024-09-1270.34%
2026-09-01
Source-checked
54Kimi K2.5Moonshot2026-01-2770.2%
2026-09-01
Source-checked
55GPT-5.1 Codex (high)OpenAI2025-11-1370%
2026-09-01
Source-checked
56MiniMax-M2.1MiniMax2025-12-2369.86%
2026-09-01
Source-checked
57MiMo-V2-ProXiaomi2026-03-1868.84%
2026-09-01
Source-checked
58Mistral Medium 3.5Mistral2026-04-2968.78%
2026-09-01
Source-checked
59Kimi K2 ThinkingMoonshot2025-11-0668.1%
2026-09-01
Source-checked
60GLM-4.7 (Reasoning)Zhipu2025-12-2267.89%
2026-09-01
Source-checked
61Qwen3.6 27B (Reasoning)Alibaba2026-04-2267.55%
2026-09-01
Source-checked
62MiMo-V2-Omni-0327Xiaomi2026-03-2767.35%
2026-09-01
Source-checked
63Step 3.7 FlashStepFun2026-06-0167.28%
2026-09-01
Source-checked
64MiMo-V2.5Xiaomi2026-04-2267.14%
2026-09-01
Source-checked
65KAT Coder Pro V2Other2026-03-2766.67%
2026-09-01
Source-checked
66GPT-5 (low)OpenAI2025-08-0766.6%
2026-09-01
Source-checked
67HyperNova 60B 2605Multiverse2026-05-0666.46%
2026-09-01
Source-checked
68Nex-N2-ProOther2026-06-0266.19%
2026-09-01
Source-checked
69GPT-5.4 (low)OpenAI2026-03-0565.92%
2026-09-01
Source-checked
70GPT-5.2 (medium)OpenAI2025-12-1165.24%
2026-09-01
Source-checked
71GPT-5.5 (low)OpenAI2026-04-2364.35%
2026-09-01
Source-checked
72Qwen3.6 35B A3B (Reasoning)Alibaba2026-04-1664.35%
2026-09-01
Source-checked
73MiMo-V2-Flash (Reasoning)Xiaomi2025-12-1664.22%
2026-09-01
Source-checked
74Claude Fable 5Anthropic2026-06-0963.47%
2026-09-01
Source-checked
75Nemotron 3 Nano Omni 30B A3B ReasoningNVIDIA2026-04-2963.2%
2026-09-01
Source-checked
76Hy3-preview (Reasoning)Other2026-04-2363.13%
2026-09-01
Source-checked
77Kimi K2.7 CodeMoonshot2026-06-1263.13%
2026-09-01
Source-checked
78Claude Opus 4.8Anthropic2026-05-2862.24%
2026-09-01
Source-checked
79GLM 5V Turbo (Reasoning)Zhipu2026-04-0161.09%
2026-09-01
Source-checked
80DeepSeek V3.2 (Reasoning)DeepSeek2025-12-0160.68%
2026-09-01
Source-checked
81DiffusionGemma 26B A4BGoogle2026-06-1059.46%
2026-09-01
Source-checked
82Claude Opus 4.7Anthropic2026-04-1658.64%
2026-09-01
Source-checked
83Claude Opus 4.5 (Reasoning)Anthropic2025-11-2457.96%
2026-09-01
Source-checked
84EXAONE 4.5 33BOther2026-04-0957.96%
2026-09-01
Source-checked
85North Mini CodeCohere2026-06-0957.55%
2026-09-01
Source-checked
86Ling 2.6 FlashOther2026-04-2157.41%
2026-09-01
Source-checked
87Claude 4.5 Sonnet (Reasoning)Anthropic2025-09-2957.28%
2026-09-01
Source-checked
88Ling-2.6-1TOther2026-04-2356.87%
2026-09-01
Source-checked
89Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)Anthropic2026-02-1756.6%
2026-09-01
Source-checked
90LFM2.5-8B-A1BLiquid2026-06-0855.65%
2026-09-01
Source-checked
91Claude 4.1 Opus (Reasoning)Anthropic2025-08-0555.44%
2026-09-01
Source-checked
92GLM-5 (Non-reasoning)Zhipu2026-02-1155.17%
2026-09-01
Source-checked
93Gemini 3 Flash PreviewGoogle2025-12-1755.1%
2026-09-01
Source-checked
94Claude 4 Opus (Reasoning)Anthropic2025-05-2253.74%
2026-09-01
Source-checked
95Grok 4xAI2025-07-1053.67%
2026-09-01
Source-checked
96MiMo-V2-OmniXiaomi2026-03-1953.54%
2026-09-01
Source-checked
97Claude Opus 4.6 (Adaptive Reasoning, Max Effort)Anthropic2026-02-0553.13%
2026-09-01
Source-checked
98Qwen3.5 397B A17B (Non-reasoning)Alibaba2026-02-1651.63%
2026-09-01
Source-checked
99Gemini 3 Pro Preview (low)Google2025-11-1849.66%
2026-09-01
Source-checked
100Grok 4.20 0309 v2 (Non-reasoning)xAI2026-04-0749.32%
2026-09-01
Source-checked
101MiniCPM5-1B (Reasoning)OpenBMB2026-06-0449.32%
2026-09-01
Source-checked
102Hy3-preview (Non-reasoning)Other2026-04-2347.96%
2026-09-01
Source-checked
103Grok 4.3 (Non-reasoning)xAI2026-04-3047.62%
2026-09-01
Source-checked
104Gemini 3.5 Flash (minimal)Google2026-05-1947.28%
2026-09-01
Source-checked
105DeepSeek V4 Flash (Non-reasoning)DeepSeek2026-04-2447.21%
2026-09-01
Source-checked
106GPT-5.5 (Non-reasoning)OpenAI2026-04-2346.12%
2026-09-01
Source-checked
107DeepSeek V4 Pro (Non-reasoning)DeepSeek2026-04-2445.78%
2026-09-01
Source-checked
108Qwen3.6 27B (Non-reasoning)Alibaba2026-04-2245.71%
2026-09-01
Source-checked
109Gemma 4 12B (Non-reasoning)Google2026-06-1045.17%
2026-09-01
Source-checked
110Claude Opus 4.6Anthropic2026-02-0544.63%
2026-09-01
Source-checked
111Ring-2.6-1TOther2026-05-0844.56%
2026-09-01
Source-checked
112Granite 4.1 30BOther2026-04-2944.35%
2026-09-01
Source-checked
113Kimi K2.6 (Non-reasoning)Other2026-04-2044.29%
2026-09-01
Source-checked
114Claude Opus 4.7 (Non-reasoning, High Effort)Anthropic2026-04-1643.61%
2026-09-01
Source-checked
115Claude Opus 4.5Anthropic2025-11-2442.99%
2026-09-01
Source-checked
116Llama 4 MaverickMeta2026-04-0542.99%
2026-09-01
Source-checked
117MiMo-V2.5-Pro (Non-reasoning)Xiaomi2026-04-2242.72%
2026-09-01
Source-checked
118Claude Sonnet 4.6 (Non-reasoning, Low Effort)Anthropic2026-02-1742.38%
2026-09-01
Source-checked
119JT-35B-FlashOther2026-05-1442.04%
2026-09-01
Source-checked
120Claude Sonnet 4.6Anthropic2026-02-1741.16%
2026-09-01
Source-checked
121DeepSeek R1DeepSeek2025-01-2039.59%
2026-09-01
Source-checked
122Llama 4 ScoutMeta2026-04-0539.52%
2026-09-01
Source-checked
123JT-MINIOther2026-04-1536.67%
2026-09-01
Source-checked
124Mistral Large 3Mistral2025-12-0236.19%
2026-09-01
Source-checked
125Qwen3.6 35B A3B (Non-reasoning)Alibaba2026-04-1636.19%
2026-09-01
Source-checked
126MiniCPM5-1B (Non-reasoning)OpenBMB2026-05-2535.17%
2026-09-01
Source-checked
127MiniCPM-V 4.6 1.3BOpenBMB2026-05-1126.73%
2026-09-01
Source-checked
Method

What it covers

Archived — dropped from AA Intelligence Index v4.1 for saturation. AA continues to run it; VerdictPal keeps the glossary entry as an honesty receipt.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

Receipts

Sources and further reading

Questions about this benchmark

Every answer below is assembled from the dated fields on this page. Nothing is written separately for search.

What does IFBench measure?

Precise instruction-following on novel constraints — tests whether models can obey complex, unseen formatting rules without drifting.

What does a high IFBench score not prove?

A strong IFBench result says nothing about:

  • Frontier discrimination today — Artificial Analysis removed IFBench from Intelligence Index v4.1 because frontier models saturated it.
  • Factual accuracy — a perfectly formatted wrong answer still scores.
  • Real user tasks — synthetic constraint tests, not production workflows.

How is IFBench scored?

Constraint satisfaction accuracy. AA still publishes results but no longer weights IFBench in the composite index. Task format: Novel instruction-following tasks with strict output constraints; accuracy on unseen rule sets.

Is IFBench saturated?

IFBench is currently marked Saturated in the atlas.

Can IFBench results be contaminated by training data?

Contamination risk for IFBench is graded Medium contamination. Treat every row on this page as a public claim with a source and a date, not as a controlled experiment.

The Pack · Editorial newsletter

New tools in your inbox. Free.

One short email when a tool ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.