Benchmark profile
Vals-hosted Terminal-Bench 2.0 mirror (Vals Terminal-Bench 2.0 mirror)
Vals AI hosted Terminal-Bench 2.0 view with easy, medium, and hard task splits.
Data verifiedHow BenchLM shows Vals Terminal-Bench 2.0 mirror
BenchLM mirrors the public Vals AI Vals Terminal-Bench 2.0 mirror leaderboard captured from https://www.vals.ai/benchmarks/terminal-bench-2 and updated by Vals on June 4, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.
Vals Terminal-Bench 2.0 mirror is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.
Vals Terminal-Bench 2.0 mirror score on Vals Terminal-Bench 2.0 mirror — June 4, 2026
BenchLM mirrors the published vals terminal-bench 2.0 mirror score view for Vals Terminal-Bench 2.0 mirror. GPT-5.5 leads the public snapshot at 73.20% , followed by Claude Opus 4.8 (70.04%) and Claude Opus 4.7 (68.54%). BenchLM does not use these results to rank models overall.
GPT-5.5
OpenAI
openai/gpt-5.5
Claude Opus 4.8
Anthropic
anthropic/claude-opus-4-8
Claude Opus 4.7
Anthropic
anthropic/claude-opus-4-7
Vals Terminal-Bench 2.0 mirror score table (67 models)
ScoreThe published Vals Terminal-Bench 2.0 mirror snapshot places GPT-5.5 first at 73.20%. The third row is 4.66 points behind. The broader top-10 range is 14.77 points, so the table still separates the published systems.
67 models have been evaluated on Vals Terminal-Bench 2.0 mirror. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Vals Terminal-Bench 2.0 mirror is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About Vals Terminal-Bench 2.0 mirror
Year
2026
Tasks
Terminal task difficulty splits
Format
Accuracy score
Difficulty
Terminal-based agent execution
BenchLM mirrors this Vals-hosted Terminal-Bench view as display-only secondary context.
BenchLM freshness & provenance
Version
Vals Terminal-Bench 2.0 mirror 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does Vals Terminal-Bench 2.0 mirror measure?
Vals AI hosted Terminal-Bench 2.0 view with easy, medium, and hard task splits.
Which model leads the published Vals Terminal-Bench 2.0 mirror snapshot?
GPT-5.5 currently leads the published Vals Terminal-Bench 2.0 mirror snapshot with 73.20% vals terminal-bench 2.0 mirror score. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on Vals Terminal-Bench 2.0 mirror?
67 AI models are included in BenchLM's mirrored Vals Terminal-Bench 2.0 mirror snapshot, based on the public leaderboard captured on June 4, 2026.
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.