Benchmark profile
EdgeBench
A systems and software-engineering benchmark from ByteDance Seed that evaluates agents on long-horizon edge tasks using time-budgeted learning curves rather than a single static pass rate.
How BenchLM shows EdgeBench
BenchLM mirrors the official EdgeBench results published by ByteDance Seed with the July 2, 2026 release. EdgeBench contains 134 real-world, day-scale tasks across 6 domains, built by domain experts averaging 57.2 hours per task, and 51 tasks are publicly released together with the SForge evaluation harness.
This page ranks models by the primary published metric: average score after 12 hours of agent interaction on the full 134-task suite. The mirrored snapshot also preserves each model's score on the 51-task open-source subset.
EdgeBench is display only on BenchLM. The published rows measure long-horizon agent runs inside the SForge harness rather than normalized model-only comparisons, and most of the task suite is not public, so BenchLM does not use these scores as weighted ranking inputs.
Score @12h on EdgeBench — July 2, 2026 release
BenchLM mirrors the published score @12h view for EdgeBench. Claude Opus 4.8 leads the public snapshot at 51.3% , followed by GPT-5.5 (48.4%) and GPT-5.4 (39.3%). BenchLM does not use these results to rank models overall.
Claude Opus 4.8
Anthropic
GPT-5.5
OpenAI
GPT-5.4
OpenAI
Score @12h table (5 models)
ScoreThe published EdgeBench snapshot places Claude Opus 4.8 first at 51.3%. The third row is 12.0 points behind. The broader top-10 range is 20.3 points, so the table still separates the published systems.
5 models have been evaluated on EdgeBench. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. EdgeBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About EdgeBench
Year
2026
Tasks
Systems and software-engineering tasks
Format
Time-budgeted agent learning curves
Difficulty
Long-horizon engineering
BenchLM tracks EdgeBench as source metadata for now. The reviewed site, paper, GitHub repository, and Hugging Face dataset describe tasks and learning-curve methodology, but do not provide a stable aggregate model leaderboard suitable for scored model rows.
BenchLM freshness & provenance
Version
EdgeBench 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
51 of 134 tasks public with the SForge harness
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does EdgeBench measure?
A systems and software-engineering benchmark from ByteDance Seed that evaluates agents on long-horizon edge tasks using time-budgeted learning curves rather than a single static pass rate.
Which model leads the published EdgeBench snapshot?
Claude Opus 4.8 currently leads the published EdgeBench snapshot with 51.3% score @12h. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on EdgeBench?
5 AI models are included in BenchLM's mirrored EdgeBench snapshot, based on the public leaderboard captured on July 2, 2026 release.
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.