Benchmark profile
Artificial Analysis IFBench (AA-IFBench)
A display-only Artificial Analysis IFBench score.
Data verifiedBenchmark score on AA-IFBench — July 23, 2026
BenchLM mirrors the published score view for AA-IFBench. MiniMax M3 leads the public snapshot at 82.9% , followed by Nemotron 3 Ultra (81.4%) and Grok 4.3 (81.3%). BenchLM does not use these results to rank models overall.
MiniMax M3
MiniMax
minimax-m3
Nemotron 3 Ultra
NVIDIA
nemotron-3-ultra-500b
Grok 4.3
xAI
grok-4-3
Benchmark score table (146 models)
ScoreThe published AA-IFBench snapshot places MiniMax M3 first at 82.9%. The third row is 1.6 points behind. The broader top-10 range is 5.3 points, so many of the published results sit in a relatively narrow band.
146 models have been evaluated on AA-IFBench. The benchmark falls in the Instruction Following category. This category carries a 5% weight in BenchLM.ai's overall scoring system. AA-IFBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About AA-IFBench
Year
2026
Tasks
Verifiable instruction constraints
Format
Constraint satisfaction accuracy
Difficulty
Instruction precision
BenchLM stores the Artificial Analysis IFBench result separately from the weighted IFBench lane so AA refreshes remain display-only.
BenchLM freshness & provenance
Version
AA-IFBench 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does AA-IFBench measure?
A display-only Artificial Analysis IFBench score.
Which model scores highest on AA-IFBench?
MiniMax M3 by MiniMax currently leads with a score of 82.9% on AA-IFBench.
How many models are evaluated on AA-IFBench?
146 AI models have been evaluated on AA-IFBench on BenchLM.
Compare Top Models on AA-IFBench
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.