Skip to main content
BenchLM
Data

DeepSWE

We show this table for reference; we do not rank on it.

Data verified 36 confirmed releases in the last 30 daysFollow model changes

A long-horizon software engineering benchmark from Datacurve for measuring frontier coding agents on original tasks drawn from active open-source repositories.

Pass@1 on DeepSWE — September 1, 2026

We mirror the published pass@1 view for DeepSWE. Gemini 3.8 Flash leads the public snapshot at 73.8%, followed by Claude Opus 5 (73.6%) and GPT-6 Astra (73.2%). We do not use these results to rank models overall.

28 modelsCoding15% of Coding reference weightCurrentUpdated September 1, 2026

Pass@1 table (28 models)

Score
1
Gemini 3.8 FlashGoogle · Closedmini-swe-agent · high reasoning
73.8%
2
Claude Opus 5Anthropic · Closedmini-swe-agent · max reasoning
73.6%
3
GPT-6 AstraOpenAI · Closedmini-swe-agent · max reasoning
73.2%
4
GPT-5.6 SolOpenAI · Closedmini-swe-agent · max reasoning
72.7%
5
Claude Fable 5Anthropic · Closedmini-swe-agent · max reasoning
69.7%
6
GPT-5.6 TerraOpenAI · Closedmini-swe-agent · max reasoning
69.6%
7
GLM-5.3Z.AI · Open weightmini-swe-agent · max reasoning
69.0%
8
Kimi K3Moonshot AI · Closedmini-swe-agent · max reasoning
68.5%
9
GPT-5.6 LunaOpenAI · Closedmini-swe-agent · max reasoning
67.2%
10
GPT-5.5OpenAI · Closedmini-swe-agent · xhigh reasoning
67.0%
11
Grok 4.6xAI · Closedmini-swe-agent · xhigh reasoning
66.7%
12
Gemini 3.7 FlashGoogle · Closedmini-swe-agent · high reasoning
65.3%
13
GLM-5.3-FlashZ.AI · Open weightmini-swe-agent · max reasoning
63.4%
14
DeepSeek V4 Pro 0813DeepSeek · Open weightmini-swe-agent · max reasoning
62.8%
15
Claude Opus 4.8Anthropic · Closedmini-swe-agent · max reasoning
59.0%
16
Qwen3.8 MaxAlibaba · Open weightmini-swe-agent · xhigh reasoning
57.5%
17
Muse Spark 1.2Meta · Closedmini-swe-agent · xhigh reasoning
54.9%
18
Claude Sonnet 5Anthropic · Closedmini-swe-agent · max reasoning
53.8%
19
Grok 4.5xAI · Closedmini-swe-agent · high reasoning
53.8%
20
DeepSeek V4 Flash 0731DeepSeek · Open weightmini-swe-agent · max reasoning
53.3%
21
Muse Spark 1.1Meta · Closedmini-swe-agent · xhigh reasoning
53.3%
22
GPT-5.4OpenAI · Closedmini-swe-agent · xhigh reasoning
51.8%
23
Gemini 3.6 FlashGoogle · Closedmini-swe-agent · high reasoning
46.7%
24
GLM-5.2Z.AI · Open weightmini-swe-agent · max reasoning
43.8%
25
Gemini 3.5 FlashGoogle · Closedmini-swe-agent · high reasoning
36.1%
26
Kimi K2.7 CodeMoonshot AI · Open weightmini-swe-agent
30.5%
27
Claude Sonnet 4.6Anthropic · Closedmini-swe-agent · high reasoning
29.9%
28
Gemini 3.1 ProGoogle · Closedmini-swe-agent · high reasoning
11.7%

How we show DeepSWE

We mirror the public DeepSWE leaderboard JSON from Datacurve. The snapshot shows the best available mini-swe-agent configuration per model, while preserving 70 underlying effort-level rows in the source metadata.

DeepSWE evaluates coding agents on 113 original, long-horizon software engineering tasks across 91 repositories and 5 languages, using isolated task environments and program-based verifiers.

Each row keeps the harness and settings the source published. BenchLM shows the published table for reference; the per-model scores it stores feed the ranking.

Snapshot

28 best-per-model rows113 tasks91 repositories5 languages70 effort rows in sourceScored

The published DeepSWE snapshot places Gemini 3.8 Flash first at 73.8%. The third row is 0.6 points behind. The broader top-10 range is 6.8 points, so many of the published results sit in a relatively narrow band.

28 models have been evaluated on DeepSWE. The benchmark falls in the Coding category. BenchLM shows the published table for reference; the per-model scores it stores feed the ranking. BenchAlign v5.8 gives DeepSWE 15% of the Coding reference weight, so it moves the Coding leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.

About DeepSWE

Year

2026

Tasks

113 software engineering tasks across 91 repositories and 5 languages

Format

Pass@1 with confidence interval, cost, time, and token metadata

Difficulty

Long-horizon software engineering

DeepSWE includes original tasks with isolated environments and program-based verifiers. BenchLM mirrors the public DeepSWE leaderboard JSON, using the best available mini-swe-agent configuration per model and preserving cost, time, token, and effort-level source metadata. Each row combines a model, agent harness, and reasoning-effort setting rather than a pure model-only benchmark score.

Freshness and provenance

Version

DeepSWE 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does DeepSWE measure?

A long-horizon software engineering benchmark from Datacurve for measuring frontier coding agents on original tasks drawn from active open-source repositories.

Which model leads the published DeepSWE snapshot?

Gemini 3.8 Flash currently leads the published DeepSWE snapshot with 73.8% pass@1. BenchLM shows the published table for reference; the per-model scores it stores feed the ranking.

How many models are evaluated on DeepSWE?

The September 1, 2026 snapshot contains 28 AI models.

Last updated: September 1, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.