Skip to main content

Benchmark profile

EdgeBench

A systems and software-engineering benchmark from ByteDance Seed that evaluates agents on long-horizon edge tasks using time-budgeted learning curves rather than a single static pass rate.

How BenchLM shows EdgeBench

BenchLM mirrors the official EdgeBench results published by ByteDance Seed with the July 2, 2026 release. EdgeBench contains 134 real-world, day-scale tasks across 6 domains, built by domain experts averaging 57.2 hours per task, and 51 tasks are publicly released together with the SForge evaluation harness.

This page ranks models by the primary published metric: average score after 12 hours of agent interaction on the full 134-task suite. The mirrored snapshot also preserves each model's score on the 51-task open-source subset.

EdgeBench is display only on BenchLM. The published rows measure long-horizon agent runs inside the SForge harness rather than normalized model-only comparisons, and most of the task suite is not public, so BenchLM does not use these scores as weighted ranking inputs.

5 model rows134 tasks (51 public)6 task domainsScore @12hDisplay only

Score @12h on EdgeBench — July 2, 2026 release

BenchLM mirrors the published score @12h view for EdgeBench. Claude Opus 4.8 leads the public snapshot at 51.3% , followed by GPT-5.5 (48.4%) and GPT-5.4 (39.3%). BenchLM does not use these results to rank models overall.

5 modelsCodingCurrentDisplay onlyUpdated July 2, 2026 release

Score @12h table (5 models)

Score
1
Claude Opus 4.8Anthropic · Closed
51.3%
2
GPT-5.5OpenAI · Closed
48.4%
3
GPT-5.4OpenAI · Closed
39.3%
4
GLM-5.1Z.AI · Open weight
37.4%
5
DeepSeek V4 ProDeepSeek · Open weight
31%

The published EdgeBench snapshot places Claude Opus 4.8 first at 51.3%. The third row is 12.0 points behind. The broader top-10 range is 20.3 points, so the table still separates the published systems.

5 models have been evaluated on EdgeBench. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. EdgeBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About EdgeBench

Year

2026

Tasks

Systems and software-engineering tasks

Format

Time-budgeted agent learning curves

Difficulty

Long-horizon engineering

BenchLM tracks EdgeBench as source metadata for now. The reviewed site, paper, GitHub repository, and Hugging Face dataset describe tasks and learning-curve methodology, but do not provide a stable aggregate model leaderboard suitable for scored model rows.

BenchLM freshness & provenance

Version

EdgeBench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

51 of 134 tasks public with the SForge harness

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does EdgeBench measure?

A systems and software-engineering benchmark from ByteDance Seed that evaluates agents on long-horizon edge tasks using time-budgeted learning curves rather than a single static pass rate.

Which model leads the published EdgeBench snapshot?

Claude Opus 4.8 currently leads the published EdgeBench snapshot with 51.3% score @12h. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on EdgeBench?

5 AI models are included in BenchLM's mirrored EdgeBench snapshot, based on the public leaderboard captured on July 2, 2026 release.

Last updated: July 2, 2026 release · mirrored from the public benchmark leaderboard

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.