BenchLM recommendation
Best LLMs for Translation in 2026
As of July 23, 2026, the top model in best llms for translation on the BenchLM leaderboard is Qwen3.7 Max with a score of 100.
Last verified: July 23, 2026
Translation quality tracks the multilingual category: benchmarks that test comprehension and generation across languages. The frontier models at the top of this table are effectively tied on major-language pairs; the gaps show up in lower-resource languages, idiom, and domain terminology.
Unless noted otherwise, ranking surfaces on this page use BenchLM's provisional leaderboard lane rather than the stricter sourced-only verified leaderboard.
Bottom line: Claude Fable 5, Claude Mythos 5, and Gemini 3.1 Pro are tied at the top of the multilingual category — pick by price and ecosystem for major-language translation, and test low-resource pairs yourself.
Qwen3.7 Max leads this ranking with a score of 100, followed by Claude Opus 4.5 (82.9) and Qwen3.7 Plus (78.9). There is a significant gap between the leading models and the rest of the field.
The best open-weight option is Qwen3.5 397B (ranked #5 with a score of 69.7). While proprietary models lead, open-weight options are within striking distance for teams willing to trade a few points of performance for full model control.
This ranking is based on provisional weighted averages across the scoring benchmarks in multilingual tracked by BenchLM.ai. For detailed model profiles, click any model name below. To compare two specific models head-to-head, use the "vs #" links.
What changed
Claude Fable 5 tied at the top of the multilingual category.
Gemini 3.1 Pro matches the leaders at a fraction of the price — the value translation pick.
Claude Mythos 5 tied for the top multilingual score.
How to choose
High-stakes translation?
Claude Fable 5 — top multilingual score with strong instruction following
High-volume translation pipelines?
Gemini 3.1 Pro — leader-tier quality at $2/$12
Full multilingual benchmark detail?
The multilingual category page lists every benchmark
Self-hosted translation?
Qwen rows are the strongest open-weight multilingual picks
Full Rankings (12 models)
Key Takeaways
The top model is Qwen3.7 Max by Alibaba with a provisional score of 100.
The best open-weight model is Qwen3.5 397B at position #5.
12 models are included in this ranking.
Score in Context
What these scores mean
The multilingual score blends cross-language comprehension and generation benchmarks like MGSM and MMLU-ProX. It is the closest measured proxy for translation strength BenchLM tracks.
Known limitations
Benchmarks over-represent high-resource languages. For low-resource pairs, dialects, or domain terminology (legal, medical), run your own evaluation set — leaderboard gaps do not transfer reliably.
Best LLMs for Translation FAQ
What is the best LLM for translation?
Claude Fable 5, Claude Mythos 5, and Gemini 3.1 Pro are tied at the top of BenchLM's multilingual category. For most translation work Gemini 3.1 Pro is the practical pick — leader-tier quality at $2/$12 per million tokens, roughly a fifth of Fable 5's price.
Are LLMs better than Google Translate?
For context-heavy translation — documents, marketing copy, anything where tone matters — frontier LLMs generally produce more natural output because they use surrounding context and follow style instructions. For quick single sentences, dedicated translation tools remain faster and cheaper.
What is the best open-source model for translation?
Alibaba's Qwen rows are the strongest open-weight multilingual performers BenchLM tracks, which fits their heavily multilingual training focus. See the Qwen rankings and the open-source leaderboard for current scores per model.
How should I evaluate translation quality myself?
Build a 30-50 segment test set from your real content across your target pairs, translate with 2-3 shortlisted models, and have a native speaker rank blind. Benchmark scores shortlist correctly, but domain terminology and tone preferences are yours to verify.
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.