Skip to main content

Benchmark profile

Instruction Following Benchmark (IFBench)

IFBench evaluates precise instruction-following generalization on 58 challenging, verifiable out-of-domain constraints. Unlike IFEval which tests familiar constraint types, IFBench specifically measures how well models follow novel instructions they haven't been optimized for, exposing overfitting to common instruction patterns.

Data verified

MAI-Thinking-1 leads the IFBench leaderboard on BenchLM's July 2026 update with 85%, ahead of Nemotron 3 Ultra (81.7%) and Grok 4.3 (81.3%), across 16 tracked models.

Top models on IFBench — July 23, 2026

As of July 23, 2026, MAI-Thinking-1 leads the IFBench leaderboard with 85% , followed by Nemotron 3 Ultra (81.7%) and Grok 4.3 (81.3%).

16 modelsInstruction Following65% of category scoreCurrentUpdated July 23, 2026

Leaderboard (16 models)

Score
1
MAI-Thinking-1Microsoft · Closed
85%
2
Nemotron 3 UltraNVIDIA · Open weight
81.7%
3
Grok 4.3xAI · Closed
81.3%
4
InklingThinking Machines Lab · Open weight
79.8%
5
Qwen3.7 MaxAlibaba · Closed
79.1%
6
Qwen3.7 PlusAlibaba · Closed
79.1%
7
Gemini 3.5 FlashGoogle · Closed
76.3%
8
Qwen3.6 PlusAlibaba · Closed
75.8%
9
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weight
74.2%
10
Hy3 PreviewTencent · Open weight
63.1%
11
Claude Opus 4.5Anthropic · Closed
58%
12
Ling 2.6 FlashInclusionAI · Open weight
57%
13
LFM2.5-8B-A1BLiquidAI · Open weight
56.5%
14
ZAYA1-8BZyphra · Open weight
52.6%
15
MiniCPM5-1BOpenBMB · Open weight
46.7%
16
LFM2.5-230MLiquidAI · Open weight
38.4%

According to BenchLM.ai, MAI-Thinking-1 leads the IFBench benchmark with a score of 85%, followed by Nemotron 3 Ultra (81.7%) and Grok 4.3 (81.3%). There is significant spread across the leaderboard, making this benchmark effective at differentiating model capabilities.

16 models have been evaluated on IFBench. The benchmark falls in the Instruction Following category. This category carries a 5% weight in BenchLM.ai's overall scoring system. Within that category, IFBench contributes 65% of the category score, so strong performance here directly affects a model's overall ranking.

About IFBench

Year

2025

Tasks

58

BenchLM freshness & provenance

Version

IFBench 2025

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does IFBench measure?

IFBench evaluates precise instruction-following generalization on 58 challenging, verifiable out-of-domain constraints. Unlike IFEval which tests familiar constraint types, IFBench specifically measures how well models follow novel instructions they haven't been optimized for, exposing overfitting to common instruction patterns.

Which model scores highest on IFBench?

MAI-Thinking-1 by Microsoft currently leads with a score of 85% on IFBench.

How many models are evaluated on IFBench?

16 AI models have been evaluated on IFBench on BenchLM.

Last updated: July 23, 2026 · BenchLM version IFBench 2025

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.