Skip to main content

Why Chinese LLM Rankings Disagree

The provisional lane picked Qwen3.7 Max; BenchAlign v5 picked MiMo-V2.5-Pro. A dated audit of the contract, identity, and creator-filter split.

Published
Last updated
Reading time
7 min
External sources
0
Tags: chinese, benchmarks, methodology, rankingData and scoring methodology
In this article5 sections

Chinese LLM rankings disagree when they use different scoring contracts or answer different deployment questions. The live Chinese-model ranking owns the current decision; this article records why an older path produced a different first row.

As of July 2026, the Chinese overall leader is Kimi K3 (81, Supported). Require downloadable weights and the first row becomes MiniMax M3 (69.8, Supported). Narrow the work to coding and it changes again: Kimi K3 (78, Supported).

That is the point of this audit. “Best” is not one stable question, and a score is not interpretable without its contract.

The same filter produced two winners

On July 14, I ran the Chinese-lab filter through both ranking paths in the repository. The legacy provisional path and the public BenchAlign v5 path returned entirely different top threes:

Table 1
Path observed on July 14 First row Second row Third row
Legacy provisional Qwen3.7 Max — 83 GLM-5.2 — 81 Qwen3.7 Plus — 77
BenchAlign v5 MiMo-V2.5-Pro — 70.22 MiniMax M3 — 69.8 GLM-5.1 — 67.76

Those point totals use different contracts. An 83 in the provisional lane does not beat a 70.22 in BenchAlign v5. The numbers should never have appeared as two answers to the same URL-level question.

The mismatch was not a new model release. The article's live tokens already used BenchAlign v5, while /best/chinese-models still followed the older provisional branch. The two paths also carried different creator lists. The repair puts the live decision page and its crawler-facing Markdown mirror on one shared China filter and one public ranking lane.

Why the contracts disagreed

The provisional table used weighted display scores from the catalog. BenchAlign v5 uses an approved snapshot, representative capability groups, evidence routes, and conditional uncertainty. Its Supported and Estimated labels describe the evidence behind a position; they are not a second leaderboard.

Model identity matters too. A provider may publish Base, High, or Max configurations under one family name. The provisional path could rank one configuration as its own row. BenchAlign v5 can instead use a representative capability group when the evidence does not support treating every harness setting as a separate all-purpose model.

That is why the method has to travel with the number. The methodology explains the public contract. The state-of-benchmarks review shows how the same evidence rule affects other ranking surfaces.

One table cannot answer three decisions

The overall table is useful when broad benchmark performance is the constraint. It is the wrong final answer when the job is narrower.

Table 2
Decision Current owner What changes the shortlist
Broad Chinese-model order Chinese-model ranking Overall evidence, capability-group identity, score interval
Coding work Coding leaderboard Coding evidence, repository language, tool use, evidence status
Downloadable weights Open-weight ranking License, hardware, serving stack, acceptable score trade-off

The live tokens above show the consequence. The overall, open-weight, and coding questions do not share one first row. A team that requires downloadable weights has already disqualified a proprietary overall leader before latency or price enters the discussion.

This is also why “open weight” must not be shortened to “open source.” Downloadable parameters can still carry license restrictions, and an available checkpoint says nothing about the cost of running it well.

The lead is not decisive

The July 14 BenchAlign v5 snapshot put MiMo-V2.5-Pro at 70.22 and MiniMax M3 at 69.80, a 0.42-point gap. Their conditional 90% intervals overlap: 63.24–77.19 for MiMo and 65.84–73.75 for MiniMax.

That makes the ordering a point estimate, not proof of a practical win. The intervals are conditional on the current evidence and model; they do not capture every source of benchmark error or workload shift.

That distinction matters because the first two rows can trade places after a modest evidence refresh without either model changing. A useful shortlist should survive that movement; a one-name recommendation often will not.

There are three more limits:

  • “Chinese model” means the originating lab in this filter. It does not measure Chinese-language quality.
  • The overall score does not settle license terms, serving cost, throughput, data residency, or regional API availability.
  • A current overall leader can lose once coding, open-weight access, latency, or a private evaluation becomes the constraint.

How the two URLs divide the work

The live ranking owns the current contract, scores, and model order. It regenerates with the public data and carries the decision-oriented title.

This article is the dated audit trail. Its self-canonical URL stays indexable because the contract split, model-identity issue, and filter mismatch are useful analysis. It no longer republishes a second top-ten table or a competing “best by use case” list.

Use the live page to choose where to start. Use this article to understand why another ranking may disagree, then test the finalist on the workload that matters.


Reader questions

Frequently asked questions

01Which BenchLM page has the current Chinese model ranking?

The live decision page is /best/chinese-models. It rebuilds the Chinese-lab slice from the active public ranking lane instead of preserving a copied winner. This article records why an older provisional path disagreed; it should not be used as a second current leaderboard or as the current model order.

02Why did two BenchLM Chinese-model rankings disagree?

The two paths used different scoring contracts and different creator filters. One used legacy provisional weighted scores; the other used BenchAlign v5, which models evidence coverage and uncertainty. Their point totals are not on the same scale, so the higher provisional number did not prove a better model.

03Does the top overall Chinese model also lead coding?

No. The overall and coding surfaces weight different evidence, and the first row can change. The live coding page is the right owner for a coding decision. Treat an overall winner as a broad starting point, then check category evidence, access, latency, and cost for the actual workload.

04Does open weight mean open source?

No. Open weight means downloadable parameters are available under a model-specific license. It does not guarantee an OSI-approved license, unrestricted commercial use, reproducible training data, or inexpensive deployment. Read the license and estimate serving requirements before treating an open-weight row as the operational default.

Share or save

Share on XShare on LinkedIn

Keep reading

All research

New models drop every week. Join 2,000+ readers for one email a week on what moved, why, and what still needs proof.