Benchmarks / Reasoning

IFBench

Instruction-Following Benchmark

Precise instruction-following on novel constraints — tests whether models can obey complex, unseen formatting rules without drifting.

Artificial AnalysisReasoningSolidDefunctMedium contamination riskSince 2024
What this does not measure
  • Frontier discrimination today — Artificial Analysis removed IFBench from Intelligence Index v4.1 because frontier models saturated it.
  • Factual accuracy — a perfectly formatted wrong answer still scores.
  • Real user tasks — synthetic constraint tests, not production workflows.
Analysis

Why this benchmark is useful

It stresses abstraction and multi-step inference. Use it to find models that can hold a problem together after the prompt stops looking like a memorized exam.

Scope

Coverage map

Task family
Reasoning
Format
Novel instruction-following tasks with strict output constraints; accuracy on unseen rule sets.
Scoring
Constraint satisfaction accuracy. AA still publishes results but no longer weights IFBench in the composite index.
Maintainer
Artificial Analysis
Reading guide

How to read the scores

Read IFBench as a defunct signal with medium contamination risk. Compare models only when the source uses the same harness, prompting setup, sampling policy, and score unit.

Blind spots

What it does not cover

  • Frontier discrimination today — Artificial Analysis removed IFBench from Intelligence Index v4.1 because frontier models saturated it.
  • Factual accuracy — a perfectly formatted wrong answer still scores.
  • Real user tasks — synthetic constraint tests, not production workflows.
Scores

Evidence ledger

57 rows
5757 rows
82.86%best score
57source-checked
1sources
2026-07-21source date
57

api · api

Distribution

Where the rows land

0255075100

Normalized to this benchmark's axis (0–100). Open the table for raw units.

Timeline

Newest receipts

  1. MiniMax M382.86% ·
  2. Nemotron 3 Ultra 550B A55B (Reasoning)81.36% ·
  3. Grok 4.381.29% ·
  4. Grok 4.20 Reasoning81.22% ·
  5. Qwen3.7 Max80.54% ·

Scores use this benchmark's own unit and axis, not a universal quality score.

#ModelRelease dateScoreProvenanceTrust
1MiniMax M3MiniMax2026-05-3182.86%
2026-07-21
Source-checked
2Nemotron 3 Ultra 550B A55B (Reasoning)NVIDIA2026-06-0481.36%
2026-07-21
Source-checked
3Grok 4.3xAI2026-04-3081.29%
2026-07-21
Source-checked
4Grok 4.20 ReasoningxAI2026-03-0581.22%
2026-07-21
Source-checked
5Qwen3.7 MaxAlibaba2026-05-1980.54%
2026-07-21
Source-checked
6MiMo-V2.5-ProXiaomi2026-04-2279.86%
2026-07-21
Source-checked
7Qwen 3.7 PlusAlibaba2026-06-0277.96%
2026-07-21
Source-checked
8Gemini 3.1 Flash LiteGoogle2026-05-0777.21%
2026-07-21
Source-checked
9Gemini 3.1 Pro PreviewGoogle2026-02-1977.14%
2026-07-21
Source-checked
10DeepSeek V4 Pro (Reasoning, Max Effort)DeepSeek2026-04-2476.46%
2026-07-21
Source-checked
11Gemini 3.5 FlashGoogle2026-05-1976.33%
2026-07-21
Source-checked
12GLM-5.1Zhipu2026-04-0776.26%
2026-07-21
Source-checked
13Kimi K2.6Moonshot2026-04-2075.99%
2026-07-21
Source-checked
14GPT-5.4 NanoOpenAI2026-03-1775.92%
2026-07-21
Source-checked
15Muse SparkOther2026-01-0175.92%
2026-07-21
Source-checked
16GPT 5.5OpenAI2026-04-2375.85%
2026-07-21
Source-checked
17MiniMax M2.7MiniMax2026-04-1575.71%
2026-07-21
Source-checked
18Gemma 4 31B ITGoogle2026-03-0175.58%
2026-07-21
Source-checked
19GPT-5.2OpenAI2025-12-1175.44%
2026-07-21
Source-checked
20GPT-5.3 Codex (xhigh)OpenAI2026-02-0575.37%
2026-07-21
Source-checked
21Gemini 3.5 Flash (medium)Google2026-05-1974.56%
2026-07-21
Source-checked
22GPT 5.4OpenAI2026-03-0573.95%
2026-07-21
Source-checked
23Gemma 4 12B (Reasoning)Google2026-06-0773.54%
2026-07-21
Source-checked
24GLM-5.2Zhipu2026-06-1373.33%
2026-07-21
Source-checked
25GPT-5.4 MiniOpenAI2026-03-1773.27%
2026-07-21
Source-checked
26GPT-5.1OpenAI2025-11-1372.86%
2026-07-21
Source-checked
27GPT-5.6 SolOpenAI2026-07-0972.65%
2026-07-21
Source-checked
28GPT-5.5 (high)OpenAI2026-04-2371.63%
2026-07-21
Source-checked
29MiniMax M2.5MiniMax2026-04-0171.63%
2026-07-21
Source-checked
30OpenAI o3OpenAI2025-04-1671.43%
2026-07-21
Source-checked
31GPT-5.6 TerraOpenAI2026-07-0971.22%
2026-07-21
Source-checked
32GPT-5.5 (medium)OpenAI2026-04-2370.95%
2026-07-21
Source-checked
33Gemini 3 ProGoogle2025-11-1870.41%
2026-07-21
Source-checked
34OpenAI o1OpenAI2024-09-1270.34%
2026-07-21
Source-checked
35Kimi K2.5Moonshot2026-01-2770.2%
2026-07-21
Source-checked
36Step 3.7 FlashStepFun2026-06-0167.28%
2026-07-21
Source-checked
37HyperNova 60B 2605Multiverse2026-05-0666.46%
2026-07-21
Source-checked
38GPT-5.5 (low)OpenAI2026-04-2364.35%
2026-07-21
Source-checked
39Claude Fable 5Anthropic2026-06-0963.47%
2026-07-21
Source-checked
40Kimi K2.7 CodeMoonshot2026-06-1263.13%
2026-07-21
Source-checked
41Claude Opus 4.8Anthropic2026-05-2862.24%
2026-07-21
Source-checked
42Claude Opus 4.7Anthropic2026-04-1658.64%
2026-07-21
Source-checked
43North Mini CodeCohere2026-06-0957.55%
2026-07-21
Source-checked
44Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)Anthropic2026-02-1756.6%
2026-07-21
Source-checked
45LFM2.5-8B-A1BLiquid2026-06-0855.65%
2026-07-21
Source-checked
46Gemini 3 Flash PreviewGoogle2025-12-1755.1%
2026-07-21
Source-checked
47Claude Opus 4.6 (Adaptive Reasoning, Max Effort)Anthropic2026-02-0553.13%
2026-07-21
Source-checked
48MiniCPM5-1B (Reasoning)OpenBMB2026-06-0449.32%
2026-07-21
Source-checked
49Gemma 4 12B (Non-reasoning)Google2026-06-1045.17%
2026-07-21
Source-checked
50Claude Opus 4.6Anthropic2026-02-0544.63%
2026-07-21
Source-checked
51Claude Opus 4.7 (Non-reasoning, High Effort)Anthropic2026-04-1643.61%
2026-07-21
Source-checked
52Claude Opus 4.5Anthropic2025-11-2442.99%
2026-07-21
Source-checked
53Llama 4 MaverickMeta2026-04-0542.99%
2026-07-21
Source-checked
54Claude Sonnet 4.6Anthropic2026-02-1741.16%
2026-07-21
Source-checked
55DeepSeek R1DeepSeek2025-01-2039.59%
2026-07-21
Source-checked
56Llama 4 ScoutMeta2026-04-0539.52%
2026-07-21
Source-checked
57Mistral Large 3Mistral2025-12-0236.19%
2026-07-21
Source-checked
Method

What it covers

Archived — dropped from AA Intelligence Index v4.1 for saturation. AA continues to run it; VerdictPal keeps the glossary entry as an honesty receipt.

Score ceiling

Where it breaks down

No ceiling note is recorded yet. Treat clustering near the top as a warning that the benchmark may no longer separate frontier models.

Receipts

Sources and further reading

The Pack · Editorial newsletter

New cards in your inbox. Free.

One short email when a card ships or changes status. No tracking, no third-party analytics. Unsubscribe in one click.