Benchmark profile
Vals-hosted SWE-bench mirror (Vals SWE-bench mirror)
Vals AI hosted SWE-bench view for solving production software engineering tasks.
Data verifiedHow BenchLM shows Vals SWE-bench mirror
BenchLM mirrors the public Vals AI Vals SWE-bench mirror leaderboard captured from https://www.vals.ai/benchmarks/swebench and updated by Vals on July 17, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.
Vals SWE-bench mirror is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.
Vals SWE-bench score on Vals SWE-bench mirror — July 17, 2026
BenchLM mirrors the published vals swe-bench score view for Vals SWE-bench mirror. GPT-5.6 Sol leads the public snapshot at 96.20% , followed by Claude Fable 5 (95.00%) and Kimi K3 (93.40%). BenchLM does not use these results to rank models overall.
GPT-5.6 Sol
OpenAI
openai/gpt-5.6-sol
Claude Fable 5
Anthropic
anthropic/claude-fable-5
Kimi K3
Moonshot AI
kimi/kimi-k3
Vals SWE-bench score table (72 models)
ScoreThe published Vals SWE-bench mirror snapshot places GPT-5.6 Sol first at 96.20%. The third row is 2.80 points behind. The broader top-10 range is 14.20 points, so the table still separates the published systems.
72 models have been evaluated on Vals SWE-bench mirror. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Vals SWE-bench mirror is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About Vals SWE-bench mirror
Year
2026
Tasks
Software engineering issue-resolution tasks
Format
Accuracy score
Difficulty
Production software engineering
BenchLM keeps this separate from its canonical SWE-bench Verified page so Vals-hosted results remain secondary context rather than source-of-record data.
BenchLM freshness & provenance
Version
Vals SWE-bench mirror 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does Vals SWE-bench mirror measure?
Vals AI hosted SWE-bench view for solving production software engineering tasks.
Which model leads the published Vals SWE-bench mirror snapshot?
GPT-5.6 Sol currently leads the published Vals SWE-bench mirror snapshot with 96.20% vals swe-bench score. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on Vals SWE-bench mirror?
72 AI models are included in BenchLM's mirrored Vals SWE-bench mirror snapshot, based on the public leaderboard captured on July 17, 2026.
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.