Skip to main content

Benchmark profile

SWE Multilingual

A multilingual software-engineering benchmark for real-world code issue resolution across multiple programming languages.

Data verified

Benchmark score on SWE Multilingual — July 23, 2026

BenchLM mirrors the published score view for SWE Multilingual. Claude Opus 4.8 leads the public snapshot at 84.4% , followed by Composer 2.5 (79.8%) and Ornith-1.0-397B (78.9%). BenchLM does not use these results to rank models overall.

29 modelsCodingCurrentDisplay onlyUpdated July 23, 2026

Benchmark score table (29 models)

Score
1
Claude Opus 4.8Anthropic · Closed
84.4%
2
Composer 2.5Cursor · Closed
79.8%
3
Ornith-1.0-397BDeepReinforce AI · Open weight
78.9%
4
Laguna S 2.1Poolside · Open weight
78.5%
5
Claude Sonnet 5Anthropic · Closed
78.3%
6
Qwen3.7 MaxAlibaba · Closed
78.3%
7
Grok 4.5xAI · Closed
78%
8
SWE-1.7Cognition · Closed
77.8%
9
Claude Opus 4.5Anthropic · Closed
77.5%
10
Kimi K2.6Moonshot AI · Open weight
76.7%
11
MiniMax M2.7MiniMax · Open weight
76.5%
12
DeepSeek V4 Pro (Max)DeepSeek · Open weight
76.2%
13
Qwen3.7 PlusAlibaba · Closed
75.8%
14
DeepSeek V4 Pro (High)DeepSeek · Open weight
74.1%
15
Qwen3.6 PlusAlibaba · Closed
73.8%
16
Composer 2Cursor · Closed
73.7%
17
GLM-5Z.AI · Open weight
73.3%
18
DeepSeek V4 Flash (Max)DeepSeek · Open weight
73.3%
19
Kimi K2.5Moonshot AI · Open weight
73%
20
Qwen3.6-27BAlibaba · Open weight
71.3%
21
DeepSeek V4 Flash (High)DeepSeek · Open weight
70.2%
22
DeepSeek V4 ProDeepSeek · Open weight
69.8%
23
DeepSeek V4 FlashDeepSeek · Open weight
69.7%
24
Ornith-1.0-35BDeepReinforce AI · Open weight
69.3%
25
Nemotron 3 UltraNVIDIA · Open weight
67.7%
26
Qwen3.6-35B-A3BAlibaba · Open weight
67.2%
27
Laguna M.1Poolside · Closed
63.1%
28
Laguna XS.2Poolside · Open weight
57.7%
29
Ornith-1.0-9BDeepReinforce AI · Open weight
52%

The published SWE Multilingual snapshot places Claude Opus 4.8 first at 84.4%. The third row is 5.5 points behind. The broader top-10 range is 7.7 points, so many of the published results sit in a relatively narrow band.

29 models have been evaluated on SWE Multilingual. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. SWE Multilingual is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About SWE Multilingual

Year

2026

Tasks

Multilingual software-engineering tasks

Format

Repository task completion

Difficulty

Professional software engineering

MiniMax reports SWE Multilingual as a coding benchmark focused on multilingual software-engineering tasks beyond single-language Python issue fixing.

BenchLM freshness & provenance

Version

SWE Multilingual 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does SWE Multilingual measure?

A multilingual software-engineering benchmark for real-world code issue resolution across multiple programming languages.

Which model scores highest on SWE Multilingual?

Claude Opus 4.8 by Anthropic currently leads with a score of 84.4% on SWE Multilingual.

How many models are evaluated on SWE Multilingual?

29 AI models have been evaluated on SWE Multilingual on BenchLM.

Last updated: July 23, 2026 · BenchLM version SWE Multilingual 2026

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.