Skip to main content

Benchmark profile

Senior SWE-Bench

A Snorkel AI benchmark of senior-level software engineering tasks emphasizing under-specified feature work, bug/performance investigation, and taste-based correctness.

How BenchLM shows Senior SWE-Bench

BenchLM mirrors the official Senior SWE-Bench leaderboard published by Snorkel AI for the v2026.06 release. The benchmark contains 100 long-horizon tasks sourced from real pull requests across 12 production repositories, with 50 tasks public and 50 held private to mitigate contamination.

This page ranks models by the primary published metric: tasteful solve rate (pass@1), scored by a taste judge and a validation agent judge that Snorkel calibrated against reviews from its senior software engineering expert network. Snorkel removes runs with detected reward hacking, such as agents searching GitHub for the original pull request, from published scores.

Senior SWE-Bench is display only on BenchLM. The published rows are long-horizon agent-harness results with judge-based scoring rather than normalized model-only comparisons, and half the task suite is private, so BenchLM does not use these scores as weighted ranking inputs.

9 model rows100 tasks (50 public)12 source repositoriesTasteful solve rate (pass@1)Display only

Tasteful solve rate (pass@1) on Senior SWE-Bench — v2026.06 release

BenchLM mirrors the published tasteful solve rate (pass@1) view for Senior SWE-Bench. Claude Opus 4.8 leads the public snapshot at 24% , followed by Claude Sonnet 5 (19.4%) and GPT-5.5 (16%). BenchLM does not use these results to rank models overall.

9 modelsCodingCurrentDisplay onlyUpdated v2026.06 release

Tasteful solve rate (pass@1) table (9 models)

Score
1
Claude Opus 4.8Anthropic · Closed
24%
2
Claude Sonnet 5Anthropic · Closed
19.4%
3
GPT-5.5OpenAI · Closed
16%
4
Claude Opus 4.7Anthropic · Closed
14.1%
5
GLM-5.2Z.AI · Open weight
12.5%
6
Kimi K2.6Moonshot AI · Open weight
8.2%
7
Claude Sonnet 4.6Anthropic · Closed
8.2%
8
Gemini 3.1 ProGoogle · Closed
6.1%
9
Gemini 3.5 FlashGoogle · Closed
3%

The published Senior SWE-Bench snapshot places Claude Opus 4.8 first at 24%. The third row is 8.0 points behind. The broader top-10 range is 21.0 points, so the table still separates the published systems.

9 models have been evaluated on Senior SWE-Bench. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. Senior SWE-Bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Senior SWE-Bench

Year

2026

Tasks

Senior-level repository tasks

Format

Agentic software-engineering evaluation

Difficulty

Professional senior engineering

BenchLM tracks Senior SWE-Bench as source metadata for now. The public page describes 50 public and 50 private tasks and renders a model comparison chart, but the crawler output does not expose exact aggregate model scores for each row.

BenchLM freshness & provenance

Version

Senior SWE-Bench v2026.06

Refresh cadence

Quarterly

Staleness state

Current

Question availability

50 of 100 tasks public as a Harbor dataset

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does Senior SWE-Bench measure?

A Snorkel AI benchmark of senior-level software engineering tasks emphasizing under-specified feature work, bug/performance investigation, and taste-based correctness.

Which model leads the published Senior SWE-Bench snapshot?

Claude Opus 4.8 currently leads the published Senior SWE-Bench snapshot with 24% tasteful solve rate (pass@1). BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Senior SWE-Bench?

9 AI models are included in BenchLM's mirrored Senior SWE-Bench snapshot, based on the public leaderboard captured on v2026.06 release.

Last updated: v2026.06 release · mirrored from the public benchmark leaderboard

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.