Method update, July 14: The disclosure counts below come from the July 4 audit cohort. BenchAlign v5 has since replaced conservative fixed imputation with uncertainty-aware estimation and Supported/Estimated evidence labels. The sourcing imbalance remains relevant; the old calculation does not.
What AI Labs Don't Publish: The Benchmark Disclosure Gap
In the July 4 audit cohort, 29 of 30 leading models had citable coding results and 10 had citable multilingual results. Between those two numbers sits every incentive that shapes what AI labs disclose. The cohort is now historical; the disclosure pattern is still worth reading category by category, lab by lab, with the receipts attached.
Twenty-nine publish coding. Ten publish multilingual.
We only count a row on the verified board when a primary source is attached: a model card, a technical report, a launch page with the exact number for the exact variant. That sourcing bar makes absence visible, and the absences are not random.
Labs benchmark what sells.
Coding and agentic results move API revenue, so the top 30 publish them almost without exception (29 of 30 each). Competition math moves Twitter for an afternoon. Multilingual performance matters enormously to users in eighty countries and to approximately zero enterprise procurement decks written in English. The disclosure rates below are the cleanest picture of lab priorities we know how to draw, precisely because nobody composed it on purpose. Each lab made its own small, local, defensible choices about which evaluations to run and which results to ship. The aggregate is the tell.
The disclosure ladder
Sourcing rates across the top 30 verified models, by category, as of July 2026:
| Category | Top-30 models with sourced results |
|---|---|
| Coding | 29 of 30 |
| Agentic | 29 of 30 |
| Knowledge | 29 of 30 |
| Multimodal | 18 of 30 |
| Competition math | 17 of 30 |
| Reasoning | 13 of 30 |
| Instruction following | 13 of 30 |
| Multilingual | 10 of 30 |
The ladder has a shape: the further a category sits from "demo that closes an enterprise deal," the fewer labs publish it.
The top rungs are unsurprising. Coding and agentic scores are the currency of the API market, and knowledge benchmarks (MMLU-Pro, GPQA and their successors) are the legacy lingua franca every launch post still speaks. One model in thirty lacking a sourced row on any of those rungs is noise, not pattern.
The middle of the ladder is where it gets interesting. Reasoning at 13 of 30 partly reflects harness fragmentation: long-context and abstract-reasoning suites have less standardization than coding, so labs publish idiosyncratic subsets that fail exact-variant matching. Multimodal at 18 of 30 is the gap with the clearest lab-level structure, and it gets its own section below.
Instruction following at 13 of 30 is the quietly damning row. IFEval and its successors are cheap to run, utterly standard, and measure the single property every one of these models is marketed on. Every API landing page promises a model that does what it is told; more than half the makers decline to publish the standard measurement of exactly that. There is no harness-fragmentation excuse on this rung.
And the bottom rung is multilingual, at 10 of 30. Two thirds of the top models on a leaderboard read worldwide carry no citable evidence about how they perform outside English. The evaluation suites exist and are cheap; the results either were not run or were not shipped.
What "sourced" means here, and what the math rung taught us
Our verification pipeline attaches a status to every benchmark value in the catalog. The bar for a weighted, verified row is specific: a primary document (model card, technical report, launch page, or equivalent) stating the number for the exact model variant we track, at retrievable URL, with provider-published sources preferred when multiple candidates exist. Aggregator-reported values get display-only status: visible on model pages, never weighted into rankings. Generated estimates are excluded from ranking entirely, and a build-time validator proves stripping them changes nothing public.
Math is also where we owe readers a correction that proves the method matters. An earlier draft of this analysis put competition math at 4 of 30, the bottom of the ladder by a mile. Investigating that number as part of our July 6 benchmark-weight review showed it was partly an artifact of our own key selection: we were still weighting the 2025 contest season while labs had moved on to publishing 2026 results. Rotating the weighted set to the current season lifted math to 17 of 30 overnight. The remaining gap is real, and the exact-variant rule still bites (launch decks love reporting AIME numbers for "preview" builds and max-compute configurations that never ship under the same name), but the size of the swing is a warning label for every disclosure statistic on this page, ours included: what counts as "reported" depends on who is doing the counting, and with which keys.
We did not soften the sourcing rule to fill that audit's columns. A board that accepts adjacent-checkpoint numbers as exact evidence is estimation with better marketing.
The display-only tier, where unweighted numbers live
One mechanism deserves a paragraph before the lab tables, because it explains several absences that would otherwise look stranger than they are. Alongside the weighted benchmarks, we track a display-only tier: roughly thirty benchmark families shown on model pages but excluded from every ranking computation. Aggregator-reported indexes, arena-style preference Elos, and benchmarks whose cross-model coverage is still too thin for defensible weighting all live there.
Qwen3.7 Max's entire multimodal record is one display-only row, a Design Arena Elo of 1303 synced from an aggregator. The number is real and readers can see it; it simply cannot buy ranking credit, because a preference rating from a third party is not a sourced result from the lab. The display-only tier is our compromise between showing everything we know and weighting only what survives verification, and the boundary between the tiers is itself part of the disclosure story: a lab whose category coverage consists entirely of display-only rows has published signals, not evidence.
Reasoning at 13 of 30 illustrates the other soft failure mode. The category's harnesses are fragmented: long-context suites, abstract-reasoning sets, and multi-step evaluation frameworks each have several competing versions, and labs publish whichever subset their internal tooling supports. Many reasoning rows fail our exact-variant matching not because a lab hid a number but because the number describes a benchmark version we cannot reconcile with the one we track. Fragmentation and selectivity produce identical gaps in the table, and only the lab knows which one it is.
Who skips what
The multimodal gap clusters by lab, and naming the cluster requires precision about what is being claimed. The following statement is about our catalog, not about model capability: these models have no weighted multimodal benchmark results that meet the sourcing bar.
| Lab | Top-30 models without sourced multimodal rows |
|---|---|
| DeepSeek | V4 Pro (Max), V4 Pro (High), V4 Pro, V4 Flash (Max), V4 Flash (High) |
| Z.AI | GLM-5.2, GLM-5.1, GLM-5 |
| Alibaba | Qwen3.7 Max, Qwen3.5-27B |
| NVIDIA | Nemotron 3 Ultra |
| Microsoft | MAI-Thinking-1 |
Anthropic, OpenAI, Google, Moonshot, and MiniMax all have sourced multimodal coverage in the same cohort, so this is not a case of an impossible evaluation. The strangest entry is Alibaba, which published 15 multimodal rows for Qwen3.7 Plus and zero for the Max tier above it. The harness exists, the results do not, and readers can draw their own conclusion about which direction those results probably point.
For DeepSeek the pattern spans five models of one generation, which reads more like a policy than an oversight. For NVIDIA and Microsoft, single text-focused flagships, the charitable product-scope explanation carries more weight.
Volume is not the problem
Before anyone files this under a lazy geopolitical narrative, the depth numbers point the other way.
| Lab | Models in top 30 | Average sourced benchmarks per model |
|---|---|---|
| Moonshot AI | 2 | 34 |
| Alibaba | 8 | 33 |
| Anthropic | 6 | 24 |
| OpenAI | 2 | 24 |
| DeepSeek | 5 | 23 |
| Z.AI | 3 | 21 |
| 1 | 20 | |
| NVIDIA | 1 | 18 |
| MiniMax | 1 | 15 |
| Microsoft | 1 | 14 |
The labs with missing multimodal categories are among the heaviest publishers in the catalog. Alibaba averages 33 sourced rows per top-30 model, ahead of every Western lab on the board; Moonshot leads outright at 34. The disclosure gap is category-shaped, not volume-shaped. High-volume labs flood the categories they like and go silent in the ones they do not, which is arguably a stronger signal than thin coverage everywhere would be. A lab that publishes 33 results has an evaluation team, a harness budget, and an editorial process. What that lab omits, it omits on purpose or by a priority ranking that amounts to the same thing.
The same table torpedoes the reverse narrative too. Anthropic and OpenAI, at 24 rows apiece, publish fewer numbers than the top Chinese labs while covering categories more completely. Neither disclosure culture dominates the other; they fail differently.
Two readings of one table
The charitable reading: evaluation is expensive, multimodal harnesses are genuinely painful to run, and a lab shipping a text-first model may reasonably deprioritize categories far from its product. Some of the gap is logistics, and the NVIDIA and Microsoft rows probably live mostly here.
The uncharitable reading: labs run these evaluations internally, look at the results, and publish the subset that flatters. Selective disclosure is invisible on a leaderboard that averages whatever exists, which until this month included ours. The Qwen3.7 Max case sits awkwardly for the charitable reading, because the same lab, same quarter, same modality stack published the full multimodal suite for the cheaper tier.
We cannot see inside the labs, so we decline to pick between the readings. We can change which reading pays.
What silence cost under the July 4 method
The July 4 method scored every model against the full rubric and filled an unsourced category with a conservative blend. The audit below quantified the effect at that snapshot. BenchAlign v5 no longer applies this fixed fill, so the table is an audit artifact rather than a projection of today's scores:
| Model | Missing coverage | Projected gain if sourced at own level |
|---|---|---|
| Qwen3.7 Max | multimodal, entire category | +1.9 points |
| GLM-5.2 | multimodal, entire category | +1.8 |
| DeepSeek V4 Pro (Max) | multimodal, entire category | +1.5 |
| Claude Opus 4.8 | MMMU-Pro, single row | +1.3 |
| Claude Mythos 5 | OfficeQA Pro, single row | +1.1 |
Those projections assume each model performs at its own covered level, so they are sourcing priorities rather than predictions. Note the bottom two rows: the invoice is ecumenical, and Anthropic's flagships appear on it for single missing benchmarks. Nobody in the top five of the verified board has a clean sheet.
The direction was the point: publishing needed to be worth more than silence. BenchAlign v5 keeps that goal but expresses the cost as wider uncertainty and, when evidence is thin, an Estimated label rather than a fixed imputation debit.
How to read any leaderboard after this
The disclosure ladder is not a quirk of ours; it is the shape of the public evidence every ranking is built from. Three habits transfer to reading anyone's table, including ours.
Check the denominator before the rank. A #4 built on broad independent evidence and a #5 built on a narrow disclosure set are different claims wearing the same font. BenchAlign prints an evidence label and retains provenance for exactly this reason.
Ask which categories are load-bearing. A model ranked mostly on coding, agentic, and knowledge evidence has been measured on the subjects every lab studies for. The interesting information is often in the empty columns, even when the estimator can place the model with uncertainty.
A worked example: Claude Mythos 5 ranks first overall with Supported evidence. The label is stronger than a sparse launch-day estimate, but it does not make provenance irrelevant. Readers should still inspect which sources and categories carry the result.
Treat single-number rankings as compressed arguments, not measurements. Every composite score encodes editorial decisions about weights, missing data, and protocol. The honest sites publish the decisions; the rest publish the number. We keep our methodology page current and our scoring constants in version control, and we would extend the same skepticism to us that we are recommending toward everyone else.
The rows we are waiting on
A standing offer to every lab in the tables above: publish missing results somewhere citable, and the data pipeline can attach them in a refresh cycle. The evaluation is yours to run; the uncertainty is ours to update.
The row we are watching hardest is Qwen3.7 Max multimodal, worth 1.9 points and the difference between a ranking that survives arguments and one that merely wins them. Alibaba already proved it can publish this category. The next disclosure decision is theirs, and either way, the board will say what the evidence says.
Reader questions
Frequently asked questions
01Which benchmark category do AI labs report least?
In BenchLM's July 4 audit cohort, 10 of the top 30 rows had citable multilingual results, versus 29 of 30 for coding and agentic work. The cohort is a historical disclosure snapshot, not the current ranking table.
02Why is there no multimodal score for Qwen3.7 Max on BenchLM?
Alibaba has not published weighted multimodal benchmark results for Qwen3.7 Max that meet our sourcing bar; the only multimodal signal available is a display-only Design Arena rating. Qwen3.7 Plus, the same family's lower tier, has 15 sourced multimodal rows, so the harness clearly exists.
03Do missing benchmarks lower a model's BenchLM score?
Yes. BenchAlign v5 preserves uncertainty around missing results and labels rows Supported or Estimated. Publishing independent, citable results can narrow that uncertainty and can change the score; missing results are not converted to zeros.
04Which AI lab publishes the most benchmark results?
In the July 4 audit cohort, Moonshot AI averaged 34 sourced benchmarks per model, Alibaba 33, and Anthropic and OpenAI 24. Those figures describe that audit snapshot; the live catalog continues to change.
05What counts as a sourced benchmark result on BenchLM?
A number published somewhere citable for the exact model variant: a model card, technical report, benchmark site, launch page, or equivalent record with traceable provenance. BenchAlign uses source reliability controls; no source receives automatic trust merely because it belongs to a provider or aggregator.
06How can an AI lab get missing scores added to BenchLM?
Publish the results somewhere citable: a model card, technical report, or launch page with exact numbers for the exact model variant. BenchLM's verification pipeline attaches primary sources to every weighted row, and the verification-impact audit shows which missing rows would move each model most.
Share or save