Skip to main content

Benchmark profile

Vals-hosted Terminal-Bench 2.0 mirror (Vals Terminal-Bench 2.0 mirror)

Vals AI hosted Terminal-Bench 2.0 view with easy, medium, and hard task splits.

Data verified

How BenchLM shows Vals Terminal-Bench 2.0 mirror

BenchLM mirrors the public Vals AI Vals Terminal-Bench 2.0 mirror leaderboard captured from https://www.vals.ai/benchmarks/terminal-bench-2 and updated by Vals on June 4, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

Vals Terminal-Bench 2.0 mirror is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

67 Vals rows4 task viewspublic datasetTasks: Overall, Easy, Medium, HardDisplay only

Vals Terminal-Bench 2.0 mirror score on Vals Terminal-Bench 2.0 mirror — June 4, 2026

BenchLM mirrors the published vals terminal-bench 2.0 mirror score view for Vals Terminal-Bench 2.0 mirror. GPT-5.5 leads the public snapshot at 73.20% , followed by Claude Opus 4.8 (70.04%) and Claude Opus 4.7 (68.54%). BenchLM does not use these results to rank models overall.

67 modelsExternal benchmark mirrorsCurrentDisplay onlyUpdated June 4, 2026

Vals Terminal-Bench 2.0 mirror score table (67 models)

Score
1
GPT-5.5OpenAI
73.20%
2
70.04%
3
68.54%
4
67.42%
6
64.05%
7
59.55%
8
59.55%
9
59.18%
10
58.43%
11
58.43%
12
GPT-5.4OpenAI
58.43%
13
Kimi K2.6Moonshot AI
57.30%
14
56.18%
15
55.06%
17
GLM 5.1Zhipu AI
53.93%
19
GPT-5.2OpenAI
51.69%
20
51.69%
21
49.44%
22
47.19%
23
MiniMax M3MiniMax
46.07%
24
44.94%
25
GPT-5.1OpenAI
44.94%
26
44.94%
27
44.94%
28
Grok 4.3SpaceXAI
43.45%
30
41.57%
31
41.57%
32
40.45%
33
40.45%
34
39.89%
35
39.33%
37
GLM 4.7Zhipu AI
38.20%
38
37.08%
39
GPT-5OpenAI
37.08%
40
Kimi K2 ThinkingMoonshot AI
37.08%
41
35.95%
42
DeepSeek V3p2Fireworks AI
34.83%
43
Laguna M.1Poolside
31.46%
44
30.34%
45
30.34%
46
29.21%
47
Laguna Xs.2Poolside
28.09%
48
GLM 4.6Zhipu AI
28.09%
49
Grok 4 0709SpaceXAI
28.09%
50
26.97%
51
25.84%
52
24.72%
54
24.72%
55
Qwen3 MaxAlibaba
24.72%
56
DeepSeek V3p1Fireworks AI
22.47%
58
Qwen3 MaxAlibaba
20.23%
59
GPT Oss 120bFireworks AI
19.10%
61
17.98%
62
16.85%
63
GPT-4.1OpenAI
14.61%
64
13.48%
65
8.99%
66
2.25%

The published Vals Terminal-Bench 2.0 mirror snapshot places GPT-5.5 first at 73.20%. The third row is 4.66 points behind. The broader top-10 range is 14.77 points, so the table still separates the published systems.

67 models have been evaluated on Vals Terminal-Bench 2.0 mirror. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Vals Terminal-Bench 2.0 mirror is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Vals Terminal-Bench 2.0 mirror

Year

2026

Tasks

Terminal task difficulty splits

Format

Accuracy score

Difficulty

Terminal-based agent execution

BenchLM mirrors this Vals-hosted Terminal-Bench view as display-only secondary context.

BenchLM freshness & provenance

Version

Vals Terminal-Bench 2.0 mirror 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does Vals Terminal-Bench 2.0 mirror measure?

Vals AI hosted Terminal-Bench 2.0 view with easy, medium, and hard task splits.

Which model leads the published Vals Terminal-Bench 2.0 mirror snapshot?

GPT-5.5 currently leads the published Vals Terminal-Bench 2.0 mirror snapshot with 73.20% vals terminal-bench 2.0 mirror score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Vals Terminal-Bench 2.0 mirror?

67 AI models are included in BenchLM's mirrored Vals Terminal-Bench 2.0 mirror snapshot, based on the public leaderboard captured on June 4, 2026.

Last updated: June 4, 2026 · mirrored from the public benchmark leaderboard

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.