Benchmark profile
Instruction Following Benchmark (IFBench)
IFBench evaluates precise instruction-following generalization on 58 challenging, verifiable out-of-domain constraints. Unlike IFEval which tests familiar constraint types, IFBench specifically measures how well models follow novel instructions they haven't been optimized for, exposing overfitting to common instruction patterns.
Data verifiedMAI-Thinking-1 leads the IFBench leaderboard on BenchLM's July 2026 update with 85%, ahead of Nemotron 3 Ultra (81.7%) and Grok 4.3 (81.3%), across 16 tracked models.
Top models on IFBench — July 23, 2026
As of July 23, 2026, MAI-Thinking-1 leads the IFBench leaderboard with 85% , followed by Nemotron 3 Ultra (81.7%) and Grok 4.3 (81.3%).
MAI-Thinking-1
Microsoft
mai-thinking-1
Nemotron 3 Ultra
NVIDIA
nemotron-3-ultra-500b
Grok 4.3
xAI
grok-4-3
Leaderboard (16 models)
ScoreAccording to BenchLM.ai, MAI-Thinking-1 leads the IFBench benchmark with a score of 85%, followed by Nemotron 3 Ultra (81.7%) and Grok 4.3 (81.3%). There is significant spread across the leaderboard, making this benchmark effective at differentiating model capabilities.
16 models have been evaluated on IFBench. The benchmark falls in the Instruction Following category. This category carries a 5% weight in BenchLM.ai's overall scoring system. Within that category, IFBench contributes 65% of the category score, so strong performance here directly affects a model's overall ranking.
About IFBench
Year
2025
Tasks
58
BenchLM freshness & provenance
Version
IFBench 2025
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does IFBench measure?
IFBench evaluates precise instruction-following generalization on 58 challenging, verifiable out-of-domain constraints. Unlike IFEval which tests familiar constraint types, IFBench specifically measures how well models follow novel instructions they haven't been optimized for, exposing overfitting to common instruction patterns.
Which model scores highest on IFBench?
MAI-Thinking-1 by Microsoft currently leads with a score of 85% on IFBench.
How many models are evaluated on IFBench?
16 AI models have been evaluated on IFBench on BenchLM.
Compare Top Models on IFBench
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.