Benchmark profile
SWE-Atlas Refactoring
A Scale SWE-Atlas software-engineering agent benchmark focused on refactoring tasks.
How BenchLM shows SWE-Atlas Refactoring
BenchLM mirrors the public Scale SWE-Atlas Refactoring leaderboard from July 21, 2026 snapshot. The source reports 14 agent/model rows with confidence intervals and harness labels such as Claude Code, Codex, Gemini CLI, and Mini-SWE-Agent.
SWE-Atlas Refactoring is display only on BenchLM. It is useful evidence about software-engineering agents, but the rows mix base model quality with agent harness choices, so BenchLM keeps it out of weighted model-only rankings.
Refactoring score on SWE-Atlas Refactoring — July 21, 2026 snapshot
BenchLM mirrors the published refactoring score view for SWE-Atlas Refactoring. Fable-5 (Claude Code) xHigh leads the public snapshot at 54.8% , followed by Claude Opus 4.7 (Adaptive) (48.6%) and Opus 4.8 (Claude Code)\n (46.7%). BenchLM does not use these results to rank models overall.
Fable-5 (Claude Code) xHigh
Anthropic
Claude Opus 4.7 (Adaptive)
Anthropic
Opus-4.7 (Claude Code)
Opus 4.8 (Claude Code)\n
Anthropic
Refactoring score table (14 models)
ScoreThe published SWE-Atlas Refactoring snapshot places Fable-5 (Claude Code) xHigh first at 54.8%. The third row is 8.1 points behind. The broader top-10 range is 22.5 points, so the table still separates the published systems.
14 models have been evaluated on SWE-Atlas Refactoring. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. SWE-Atlas Refactoring is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About SWE-Atlas Refactoring
Year
2026
Tasks
SWE-Atlas refactoring tasks
Format
Refactoring score with confidence intervals
Difficulty
Real-world software-engineering agent tasks
BenchLM mirrors the public Scale SWE-Atlas Refactoring leaderboard as a display-only agentic software-engineering benchmark. The source compares model-agent combinations such as Claude Code, Codex, Gemini CLI, and Mini-SWE-Agent.
BenchLM freshness & provenance
Version
SWE-Atlas Refactoring 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does SWE-Atlas Refactoring measure?
A Scale SWE-Atlas software-engineering agent benchmark focused on refactoring tasks.
Which model leads the published SWE-Atlas Refactoring snapshot?
Fable-5 (Claude Code) xHigh currently leads the published SWE-Atlas Refactoring snapshot with 54.8% refactoring score. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on SWE-Atlas Refactoring?
14 AI models are included in BenchLM's mirrored SWE-Atlas Refactoring snapshot, based on the public leaderboard captured on July 21, 2026 snapshot.
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.