Skip to main content

Benchmark profile

Vals Terminal-Bench 2.1 (Terminal-Bench 2.1)

State-of-the-art set of difficult terminal-based tasks

Data verified

How BenchLM shows Vals Terminal-Bench 2.1 mirror

BenchLM mirrors the public Vals AI Vals Terminal-Bench 2.1 mirror leaderboard captured from https://www.vals.ai/benchmarks/terminal-bench-2-1 and updated by Vals on July 16, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

Vals Terminal-Bench 2.1 mirror is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

42 Vals rows4 task viewspublic datasetTasks: Overall, Easy, Medium, HardDisplay only

Vals Terminal-Bench 2.1 mirror score on Terminal-Bench 2.1 — July 16, 2026

BenchLM mirrors the published vals terminal-bench 2.1 mirror score view for Terminal-Bench 2.1. GPT-5.6 Sol leads the public snapshot at 85.77% , followed by Kimi K3 (80.90%) and Claude Fable 5 (80.52%). BenchLM does not use these results to rank models overall.

42 modelsExternal benchmark mirrorsCurrentDisplay onlyUpdated July 16, 2026

Vals Terminal-Bench 2.1 mirror score table (42 models)

Score
1
85.77%
2
Kimi K3Moonshot AI
80.90%
3
80.52%
4
79.03%
5
GPT-5.5OpenAI
76.40%
6
74.53%
7
74.16%
8
73.41%
9
71.91%
12
69.29%
13
68.54%
14
Grok 4.5SpaceXAI
67.79%
15
GLM 5.2Zhipu AI
67.79%
16
Kimi K2.7 CodeMoonshot AI
67.04%
17
61.05%
18
Mimo V2.5Xiaomi
60.67%
19
58.43%
20
57.30%
21
57.30%
22
57.30%
23
57.30%
24
GLM 5.1Zhipu AI
56.93%
25
54.68%
27
Kimi K2.6Moonshot AI
53.56%
28
MiniMax M3MiniMax
53.56%
29
53.18%
30
52.81%
32
50.19%
33
48.69%
34
InklingThinkingmachines
47.57%
35
44.20%
37
Grok 4.3SpaceXAI
41.95%
38
41.57%
39
38.95%
41
Laguna M.1Poolside
34.08%
42
Laguna Xs.2Poolside
25.84%

The published Terminal-Bench 2.1 snapshot places GPT-5.6 Sol first at 85.77%. The third row is 5.24 points behind. The broader top-10 range is 14.98 points, so the table still separates the published systems.

42 models have been evaluated on Terminal-Bench 2.1. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Terminal-Bench 2.1 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Terminal-Bench 2.1

Year

2026

Tasks

Terminal-based task execution

Format

Accuracy score

Difficulty

Frontier terminal-agent execution

BenchLM mirrors the public Vals AI Terminal-Bench 2.1 leaderboard as display-only external evidence. The captured snapshot preserves overall scores, task-level scores where Vals publishes them, uncertainty, latency, and cost-per-test metadata. It is excluded from BenchLM weighted rankings.

BenchLM freshness & provenance

Version

Terminal-Bench 2.1 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does Terminal-Bench 2.1 measure?

State-of-the-art set of difficult terminal-based tasks

Which model leads the published Terminal-Bench 2.1 snapshot?

GPT-5.6 Sol currently leads the published Terminal-Bench 2.1 snapshot with 85.77% vals terminal-bench 2.1 mirror score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Terminal-Bench 2.1?

42 AI models are included in BenchLM's mirrored Terminal-Bench 2.1 snapshot, based on the public leaderboard captured on July 16, 2026.

Last updated: July 16, 2026 · mirrored from the public benchmark leaderboard

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.