Affiliate partnerships never influence scores, rankings, or model coverage. See the affiliate disclosure for details.
Evidence strength is separate from score
Supported positions have sufficiently diverse direct evidence. Estimated positions use the available calibrated evidence and carry wider uncertainty. Both remain visible in one ranking.
Freshness is explicit
Every benchmark now carries BenchLM freshness metadata: version, refresh cadence, staleness state, saturation state, and whether the benchmark is weighted or display-only.
Independent signals share one scale
BenchAlign calibrates benchmark protocols, human-preference signals, and external composite indices before combining them. Raw percentages are never averaged across tests with different difficulty. Runtime metrics remain separate and were updated 2026-07-21.
Six pillars and familiar evidence lenses
The product ontology has six pillars: Agentic Work, Software Engineering, Reasoning and Mathematics, Knowledge and Factuality, Multimodal and Documents, and Communication and Language. Existing category URLs remain indexable evidence lenses beneath those pillars.
Agentic
3 weighted benchmarks and 62 display-only benchmarks.
Coding
5 weighted benchmarks and 43 display-only benchmarks.
Reasoning
3 weighted benchmarks and 22 display-only benchmarks.
Multimodal
3 weighted benchmarks and 54 display-only benchmarks.
Knowledge
5 weighted benchmarks and 29 display-only benchmarks.
Multilingual
1 weighted benchmark and 10 display-only benchmarks.
Instruction Following
2 weighted benchmarks and 2 display-only benchmarks.
Math
5 weighted benchmarks and 22 display-only benchmarks.
Missing data increases uncertainty
BenchAlign estimates capability from the evidence a model does have instead of assigning zero to missing tests or averaging only the tests a provider chose to report. Sparse models can rank, but they receive an Estimated label and a wider uncertainty interval until independent evidence accumulates. Ranking order uses unrounded scores.
Method bench-align-v5.2-2026-07-17 · frozen input 43ff16a45bec1c88
Benchmarks by category
Agentic
65 tracked benchmarks
Display only
Subfamilies
Coding
48 tracked benchmarks
Display only
Reasoning
25 tracked benchmarks
Weighted
Subfamilies
Multimodal
57 tracked benchmarks
Weighted
Display only
Subfamilies
Knowledge
34 tracked benchmarks
Display only
Subfamilies
Multilingual
11 tracked benchmarks
Instruction Following
4 tracked benchmarks
Display only
Math
27 tracked benchmarks
Future tracked families
BenchLM tracks a small number of important benchmark families that are intentionally not weighted yet because exact-source density and cross-model coverage are still too thin for defensible ranking use.
Calibration approach
External consensus signals
Overall capability is anchored by multiple independent source families when available. Direct benchmark evidence contributes as a bounded residual, and no single source may dominate a Supported position. Agentic and Coding use their own relevant external and benchmark evidence.
Runtime metrics
Runtime metrics (tokens/sec, time-to-first-token) stay separate from ranking and are shown as operational metadata only. They do not affect overall or category scores.
Source refresh: 2026-07-21
BenchLM defaults and caveats
BenchLM uses benchmark freshness as a product layer, not as a claim about an official benchmark maintainer. A benchmark marked Current means BenchLM still treats it as a strong differentiator. A benchmark marked Stale means it is still useful for context but is no longer relied on heavily to separate frontier models.
Public BenchLM benchmark tables default to exact-source rows only. Generated benchmark values remain excluded from BenchAlign evidence.
Overall, Agentic, and Coding use the frozen BenchAlign v5.2 launch artifact. Other category pages remain evidence lenses on the legacy category score until their own validation is complete.
BenchLM does not estimate runtime metrics when no sourced runtime snapshot is available. The leaderboard and model pages show N/A instead. Pricing sort uses the average of input and output token price so models can be ranked on a single cost column while the full input/output pair remains visible.
Last benchmark dataset refresh: July 23, 2026. For raw benchmark exploration, use the benchmark directory. For current provider rollups, use provider pages.