Model comparison
GPT-5.6 Terra vs Muse Spark
Head-to-head evidence from 24 shared benchmark results across 7 categories. Overall scores shown here use the public BenchAlign v5 ranking lane.
Public leaderboard positions: GPT-5.6 Terra #11 (Estimated); Muse Spark #13 (Supported). Intervals and evidence labels describe ranking uncertainty, not a guarantee for a specific workload.
Evidence parity. GPT-5.6 Terra and Muse Spark share 24 comparable benchmark results. 5 of 8 categories are comparable. 20 results are unique to GPT-5.6 Terra; 15 to Muse Spark.
Updated July 23, 2026- Shared results
- 24
- GPT-5.6 Terra only
- 20
- Muse Spark only
- 15
- Comparable categories
- 5 / 8
Pick GPT-5.6 Terra if you want the stronger benchmark profile. Muse Spark only becomes the better choice if coding is the priority.
Confidence note. This is a partial-evidence comparison with 24 shared benchmark results across 7 evidence categories; 5 of 8 categories currently have scoreable aggregates for both models. Treat the verdict as directional until coverage is more balanced.
Why this result
GPT-5.6 Terra has the cleaner BenchAlign overall profile here, landing at 72.57 versus 71.04. It is a real lead, but still close enough that category-level strengths matter more than the headline number.
GPT-5.6 Terra's sharpest advantage is in mathematics, where it averages 80.8 against 32.9. The single biggest benchmark swing on the page is FrontierMath v2 (Tier 4), 68.300% to 14.600%. Muse Spark does hit back in coding, so the answer changes if that is the part of the workload you care about most.
GPT-5.6 Terra gives you the larger context window at 1M, compared with 262K for Muse Spark.
Category breakdown
Exact category averages are shown below. Not measured means BenchLM does not have enough sourced public coverage for that model and category.
| Category | GPT-5.6 Terra | Δ | Muse Spark |
|---|---|---|---|
| Math | GPT-5.6 Terra80.8 | Margin← 47.9 | Muse Spark32.9 |
| Knowledge | GPT-5.6 Terra92.9 | Margin← 42.5 | Muse Spark50.4 |
| Agentic | GPT-5.6 Terra87.4 | Margin← 28.4 | Muse Spark59.0 |
| Coding | GPT-5.6 Terra63.4 | Margin→ 4.4 | Muse Spark67.8 |
| Multimodal | GPT-5.6 Terra80.7 | Margin→ 1.8 | Muse Spark82.5 |
| Reasoning | GPT-5.6 TerraNot measured | MarginNo overlap | Muse Spark42.5 |
Decisive benchmark drivers
The largest measured benchmark gaps in this matchup, with exact reported values.
More
- Source ↗
FrontierMath v2 (Tier 4)
MathA 68.300%B 14.600%Winner: GPT-5.6 TerraΔ 53.7FrontierMath v2 (Tier 4): GPT-5.6 Terra scored 68.300%; Muse Spark scored 14.600%. GPT-5.6 Terra wins this benchmark. - Source ↗
FrontierMath v2 (Tiers 1-3)
MathA 84.900%B 39.000%Winner: GPT-5.6 TerraΔ 45.9FrontierMath v2 (Tiers 1-3): GPT-5.6 Terra scored 84.900%; Muse Spark scored 39.000%. GPT-5.6 Terra wins this benchmark. - Source ↗
Terminal-Bench 2.0
AgenticA 87.4%B 59%Winner: GPT-5.6 TerraΔ 28.4Terminal-Bench 2.0: GPT-5.6 Terra scored 87.4%; Muse Spark scored 59%. GPT-5.6 Terra wins this benchmark. - Source ↗
SWE-bench Pro
CodingA 63.4%B 52.4%Winner: GPT-5.6 TerraΔ 11SWE-bench Pro: GPT-5.6 Terra scored 63.4%; Muse Spark scored 52.4%. GPT-5.6 Terra wins this benchmark. - Source ↗
MMMU-Pro
MultimodalA 80.7%B 80.4%Winner: GPT-5.6 TerraΔ 0.3MMMU-Pro: GPT-5.6 Terra scored 80.7%; Muse Spark scored 80.4%. GPT-5.6 Terra wins this benchmark.
Operational comparison
Runtime and commercial metrics are compared only when both models have a complete sourced value.
| Metric | GPT-5.6 Terra | Muse Spark | Comparison |
|---|---|---|---|
| Input / output priceUSD per 1M tokens | GPT-5.6 Terra$2.5 input / $15 output | Muse SparkNot available | A complete price comparison is not available. |
| Generation speedtokens per second | GPT-5.6 TerraNot available | Muse SparkNot available | A complete speed comparison is not available. |
| First-answer latencyseconds to first token | GPT-5.6 TerraNot available | Muse SparkNot available | A complete latency comparison is not available. |
| Context windowmaximum listed tokens | GPT-5.6 Terra1M | Muse Spark262K | GPT-5.6 Terra lists the larger context window. |
Benchmark Deep Dive
AgenticGPT-5.6 Terra wins19 benchmarks
| Benchmark | GPT-5.6 Terra | Muse Spark | Result |
|---|---|---|---|
| Terminal-Bench 2.0Source | 87.4% | 59% | GPT-5.6 Terra leads |
| BrowseCompSource | 87.5% | — | Not comparable |
| OSWorld 2.0Source | 50.2% | — | Not comparable |
| CyberGymSource | 81.8% | 43.5% | GPT-5.6 Terra leads |
| ExploitGymSource | 23.2% | — | Not comparable |
| ToolathlonSource | 53.1% | — | Not comparable |
| AA Agentic IndexSource | 47.4% | 28.7% | GPT-5.6 Terra leads |
| τ²-bench resultsSource | 86.3% | 91.5% | Muse Spark leads |
| GDPval-AASource | 54.1% | 32.2% | GPT-5.6 Terra leads |
| GDPval-AASource | 1581 | 1144 | GPT-5.6 Terra leads |
| AA Harvey LABSource | 85.2% | — | Not comparable |
| AA ITBenchSource | 51.0% | — | Not comparable |
| AA Tau3 BankingSource | 31.8% | — | Not comparable |
| AA AutomationBenchSource | 45.6% | — | Not comparable |
| terminalBenchHardSource | 57.6% | — | Not comparable |
| aaTerminalBench21Source | 88% | — | Not comparable |
| APEX-Agents-AASource | 38.9% | — | Not comparable |
| DeepSearchQASource | — | 74.8% | Not comparable |
| Claw-EvalSource | — | 63.8% | Not comparable |
CodingMuse Spark wins10 benchmarks
| Benchmark | GPT-5.6 Terra | Muse Spark | Result |
|---|---|---|---|
| SWE-bench ProSource | 63.4% | 52.4% | GPT-5.6 Terra leads |
| Terminal-Bench 2.0Source | 87.4% | — | Not comparable |
| deepSweSource | 69.6% | — | Not comparable |
| FrontierCode 1.1 ExtendedSource | 55.8% | — | Not comparable |
| cursorBench32Source | 64.9% | — | Not comparable |
| AA Coding IndexSource | 76.7% | 58.6% | GPT-5.6 Terra leads |
| AA-SciCodeSource | 53.9% | 51.5% | GPT-5.6 Terra leads |
| SWE-bench VerifiedSource | — | 77.4% | Not comparable |
| LiveCodeBench ProSource | — | 80.0% | Not comparable |
| Vibe Code BenchSource | — | 19.67% | Not comparable |
Reasoning4 benchmarks
KnowledgeGPT-5.6 Terra wins13 benchmarks
| Benchmark | GPT-5.6 Terra | Muse Spark | Result |
|---|---|---|---|
| GPQASource | 92.9% | — | Not comparable |
| GPQA-DSource | 92.9% | 89.5% | GPT-5.6 Terra leads |
| HealthBench ProfessionalSource | 57.7% | — | Not comparable |
| HealthBench HardSource | 32.7% | 42.8% | Muse Spark leads |
| Artificial Analysis Intelligence IndexSource | 55.0% | 43.1% | GPT-5.6 Terra leads |
| AA-GPQA DiamondSource | 92.5% | 88.4% | GPT-5.6 Terra leads |
| AA-HLESource | 41.8% | 39.9% | GPT-5.6 Terra leads |
| AA-Omniscience IndexSource | -0.2% | 4.1% | Muse Spark leads |
| AA-Omniscience AccuracySource | 45.9% | 44.6% | GPT-5.6 Terra leads |
| AA-Omniscience Hallucination RateSource | 85.2% | 73.2% | Muse Spark leads |
| HLESource | — | 50.4% | Not comparable |
| HLE w/o toolsSource | — | 42.8% | Not comparable |
| MedXpertQA (Text)Source | — | 52.6% | Not comparable |
MathGPT-5.6 Terra wins3 benchmarks
MultimodalMuse Spark wins9 benchmarks
| Benchmark | GPT-5.6 Terra | Muse Spark | Result |
|---|---|---|---|
| MMMU-ProSource | 80.7% | 80.4% | GPT-5.6 Terra leads |
| MMMU-Pro w/ PythonSource | 82% | — | Not comparable |
| AA-MMMU-ProSource | 80.7% | 80.5% | GPT-5.6 Terra leads |
| CharXivSource | — | 86.4% | Not comparable |
| ERQASource | — | 64.7% | Not comparable |
| SimpleVQASource | — | 71.3% | Not comparable |
| ScreenSpot ProSource | — | 84.1% | Not comparable |
| ZeroBenchSource | — | 33.0% | Not comparable |
| MedXpertQA (MM)Source | — | 78.4% | Not comparable |
Inst. Following1 benchmarks
| Benchmark | GPT-5.6 Terra | Muse Spark | Result |
|---|---|---|---|
| AA-IFBenchSource | 71.2% | 75.9% | Muse Spark leads |
Frequently Asked Questions (6)
Which is better, GPT-5.6 Terra or Muse Spark?
GPT-5.6 Terra is ahead on BenchLM's BenchAlign leaderboard, 72.57 to 71.04. The biggest single separator in this matchup is FrontierMath v2 (Tier 4), where the scores are 68.300% and 14.600%.
Which is better for knowledge tasks, GPT-5.6 Terra or Muse Spark?
GPT-5.6 Terra has the edge for knowledge tasks in this comparison, averaging 92.9 versus 50.4. Inside this category, AA-Omniscience Hallucination Rate is the benchmark that creates the most daylight between them.
Which is better for coding, GPT-5.6 Terra or Muse Spark?
Muse Spark has the edge for coding in this comparison, averaging 67.8 versus 63.4. Inside this category, AA Coding Index is the benchmark that creates the most daylight between them.
Which is better for math, GPT-5.6 Terra or Muse Spark?
GPT-5.6 Terra has the edge for math in this comparison, averaging 80.8 versus 32.9. Inside this category, FrontierMath v2 (Tier 4) is the benchmark that creates the most daylight between them.
Which is better for agentic tasks, GPT-5.6 Terra or Muse Spark?
GPT-5.6 Terra has the edge for agentic tasks in this comparison, averaging 87.4 versus 59. Inside this category, GDPval-AA is the benchmark that creates the most daylight between them.
Which is better for multimodal and grounded tasks, GPT-5.6 Terra or Muse Spark?
Muse Spark has the edge for multimodal and grounded tasks in this comparison, averaging 82.5 versus 80.7. Inside this category, MMMU-Pro is the benchmark that creates the most daylight between them.
Related Comparisons
Explore More
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.