Benchmark profile
Senior SWE-Bench
A Snorkel AI benchmark of senior-level software engineering tasks emphasizing under-specified feature work, bug/performance investigation, and taste-based correctness.
How BenchLM shows Senior SWE-Bench
BenchLM mirrors the official Senior SWE-Bench leaderboard published by Snorkel AI for the v2026.06 release. The benchmark contains 100 long-horizon tasks sourced from real pull requests across 12 production repositories, with 50 tasks public and 50 held private to mitigate contamination.
This page ranks models by the primary published metric: tasteful solve rate (pass@1), scored by a taste judge and a validation agent judge that Snorkel calibrated against reviews from its senior software engineering expert network. Snorkel removes runs with detected reward hacking, such as agents searching GitHub for the original pull request, from published scores.
Senior SWE-Bench is display only on BenchLM. The published rows are long-horizon agent-harness results with judge-based scoring rather than normalized model-only comparisons, and half the task suite is private, so BenchLM does not use these scores as weighted ranking inputs.
Tasteful solve rate (pass@1) on Senior SWE-Bench — v2026.06 release
BenchLM mirrors the published tasteful solve rate (pass@1) view for Senior SWE-Bench. Claude Opus 4.8 leads the public snapshot at 24% , followed by Claude Sonnet 5 (19.4%) and GPT-5.5 (16%). BenchLM does not use these results to rank models overall.
Claude Opus 4.8
Anthropic
Claude Sonnet 5
Anthropic
GPT-5.5
OpenAI
Tasteful solve rate (pass@1) table (9 models)
ScoreThe published Senior SWE-Bench snapshot places Claude Opus 4.8 first at 24%. The third row is 8.0 points behind. The broader top-10 range is 21.0 points, so the table still separates the published systems.
9 models have been evaluated on Senior SWE-Bench. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. Senior SWE-Bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About Senior SWE-Bench
Year
2026
Tasks
Senior-level repository tasks
Format
Agentic software-engineering evaluation
Difficulty
Professional senior engineering
BenchLM tracks Senior SWE-Bench as source metadata for now. The public page describes 50 public and 50 private tasks and renders a model comparison chart, but the crawler output does not expose exact aggregate model scores for each row.
BenchLM freshness & provenance
Version
Senior SWE-Bench v2026.06
Refresh cadence
Quarterly
Staleness state
Current
Question availability
50 of 100 tasks public as a Harbor dataset
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does Senior SWE-Bench measure?
A Snorkel AI benchmark of senior-level software engineering tasks emphasizing under-specified feature work, bug/performance investigation, and taste-based correctness.
Which model leads the published Senior SWE-Bench snapshot?
Claude Opus 4.8 currently leads the published Senior SWE-Bench snapshot with 24% tasteful solve rate (pass@1). BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on Senior SWE-Bench?
9 AI models are included in BenchLM's mirrored Senior SWE-Bench snapshot, based on the public leaderboard captured on v2026.06 release.
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.