Skip to main content

Benchmark profile

Vals-hosted SWE-bench mirror (Vals SWE-bench mirror)

Vals AI hosted SWE-bench view for solving production software engineering tasks.

Data verified

How BenchLM shows Vals SWE-bench mirror

BenchLM mirrors the public Vals AI Vals SWE-bench mirror leaderboard captured from https://www.vals.ai/benchmarks/swebench and updated by Vals on July 17, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

Vals SWE-bench mirror is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

72 Vals rows5 task viewspublic datasetTasks: Overall, 1-4 hours, 15 min - 1 hour, <15 min fix, >4 hoursDisplay only

Vals SWE-bench score on Vals SWE-bench mirror — July 17, 2026

BenchLM mirrors the published vals swe-bench score view for Vals SWE-bench mirror. GPT-5.6 Sol leads the public snapshot at 96.20% , followed by Claude Fable 5 (95.00%) and Kimi K3 (93.40%). BenchLM does not use these results to rank models overall.

72 modelsExternal benchmark mirrorsCurrentDisplay onlyUpdated July 17, 2026

Vals SWE-bench score table (72 models)

Score
1
96.20%
2
95.00%
3
Kimi K3Moonshot AI
93.40%
4
93.00%
5
88.60%
6
Grok 4.5SpaceXAI
86.60%
8
GLM 5.2Zhipu AI
82.80%
9
GPT-5.5OpenAI
82.60%
10
82.00%
11
82.00%
12
79.60%
13
79.60%
14
78.80%
16
GPT-5.4OpenAI
78.20%
17
78.20%
18
Kimi K2.7 CodeMoonshot AI
78.20%
19
78.00%
20
InklingThinkingmachines
77.60%
21
77.40%
22
77.40%
23
76.40%
25
76.40%
26
GLM 5.1Zhipu AI
76.40%
27
76.20%
28
Kimi K2.6Moonshot AI
76.20%
29
GPT-5.2OpenAI
75.80%
30
75.20%
32
MiniMax M3MiniMax
75.00%
33
74.80%
34
74.40%
35
74.20%
36
74.00%
37
73.80%
38
73.40%
39
73.00%
40
72.80%
41
72.40%
42
72.20%
43
Grok 4.3SpaceXAI
71.40%
44
71.40%
45
71.20%
46
Mimo V2.5Xiaomi
71.00%
47
70.00%
49
70.00%
50
69.80%
51
GPT-5.1OpenAI
69.80%
52
GLM 4.7Zhipu AI
69.40%
53
GPT-5OpenAI
69.00%
55
68.80%
56
67.60%
58
66.40%
59
64.40%
61
Devstral 2512Mistral AI
62.80%
62
60.80%
63
Kimi K2 ThinkingMoonshot AI
60.20%
64
Grok 4 0709SpaceXAI
57.80%
65
Laguna M.1Poolside
57.60%
66
Laguna Xs.2Poolside
55.20%
67
54.40%
68
45.40%
69
41.40%
70
41.40%
71
GPT Oss 120bFireworks AI
33.60%
72
7.80%

The published Vals SWE-bench mirror snapshot places GPT-5.6 Sol first at 96.20%. The third row is 2.80 points behind. The broader top-10 range is 14.20 points, so the table still separates the published systems.

72 models have been evaluated on Vals SWE-bench mirror. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Vals SWE-bench mirror is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Vals SWE-bench mirror

Year

2026

Tasks

Software engineering issue-resolution tasks

Format

Accuracy score

Difficulty

Production software engineering

BenchLM keeps this separate from its canonical SWE-bench Verified page so Vals-hosted results remain secondary context rather than source-of-record data.

BenchLM freshness & provenance

Version

Vals SWE-bench mirror 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does Vals SWE-bench mirror measure?

Vals AI hosted SWE-bench view for solving production software engineering tasks.

Which model leads the published Vals SWE-bench mirror snapshot?

GPT-5.6 Sol currently leads the published Vals SWE-bench mirror snapshot with 96.20% vals swe-bench score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Vals SWE-bench mirror?

72 AI models are included in BenchLM's mirrored Vals SWE-bench mirror snapshot, based on the public leaderboard captured on July 17, 2026.

Last updated: July 17, 2026 · mirrored from the public benchmark leaderboard

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.