Benchmark profile
Vals-hosted MMLU-Pro mirror (Vals MMLU-Pro mirror)
Vals AI hosted MMLU-Pro view with subject-level task splits.
Data verifiedHow BenchLM shows Vals MMLU-Pro mirror
BenchLM mirrors the public Vals AI Vals MMLU-Pro mirror leaderboard captured from https://www.vals.ai/benchmarks/mmlu_pro and updated by Vals on July 19, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.
Vals MMLU-Pro mirror is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.
Vals MMLU-Pro mirror score on Vals MMLU-Pro mirror — July 19, 2026
BenchLM mirrors the published vals mmlu-pro mirror score view for Vals MMLU-Pro mirror. Claude Fable 5 leads the public snapshot at 91.50% , followed by Gemini 3.1 Pro Preview (90.99%) and Gemini 3 Pro Preview (90.10%). BenchLM does not use these results to rank models overall.
Claude Fable 5
Anthropic
anthropic/claude-fable-5
Gemini 3.1 Pro Preview
google/gemini-3.1-pro-preview
Gemini 3 Pro Preview
google/gemini-3-pro-preview
Vals MMLU-Pro mirror score table (122 models)
ScoreThe published Vals MMLU-Pro mirror snapshot places Claude Fable 5 first at 91.50%. The third row is 1.40 points behind. The broader top-10 range is 2.40 points, so many of the published results sit in a relatively narrow band.
122 models have been evaluated on Vals MMLU-Pro mirror. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Vals MMLU-Pro mirror is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About Vals MMLU-Pro mirror
Year
2026
Tasks
MMLU-Pro subject splits
Format
Accuracy score
Difficulty
Professional academic reasoning
BenchLM keeps this Vals-hosted MMLU-Pro table separate from canonical MMLU-Pro source records.
BenchLM freshness & provenance
Version
Vals MMLU-Pro mirror 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does Vals MMLU-Pro mirror measure?
Vals AI hosted MMLU-Pro view with subject-level task splits.
Which model leads the published Vals MMLU-Pro mirror snapshot?
Claude Fable 5 currently leads the published Vals MMLU-Pro mirror snapshot with 91.50% vals mmlu-pro mirror score. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on Vals MMLU-Pro mirror?
122 AI models are included in BenchLM's mirrored Vals MMLU-Pro mirror snapshot, based on the public leaderboard captured on July 19, 2026.
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.