Model comparison
GPT-5.6 Luna vs MiMo-V2-Pro
Head-to-head evidence from 9 shared benchmark results across 3 categories. Overall scores shown here use the public BenchAlign v5 ranking lane.
Public leaderboard positions: GPT-5.6 Luna #22 (Estimated); MiMo-V2-Pro #17 (Supported). Intervals and evidence labels describe ranking uncertainty, not a guarantee for a specific workload.
Evidence parity. GPT-5.6 Luna and MiMo-V2-Pro share 9 comparable benchmark results. 1 of 8 categories are comparable. 32 results are unique to GPT-5.6 Luna; 6 to MiMo-V2-Pro.
Updated July 23, 2026- Shared results
- 9
- GPT-5.6 Luna only
- 32
- MiMo-V2-Pro only
- 6
- Comparable categories
- 1 / 8
Pick MiMo-V2-Pro if you want the stronger benchmark profile. GPT-5.6 Luna only becomes the better choice if its workflow or ecosystem matters more than the raw scoreboard.
Confidence note. This is a partial-evidence comparison with 9 shared benchmark results across 3 evidence categories; 1 of 8 categories currently have scoreable aggregates for both models. Treat the verdict as directional until coverage is more balanced.
Why this result
MiMo-V2-Pro has the cleaner BenchAlign overall profile here, landing at 67.78 versus 67.17. It is a real lead, but still close enough that category-level strengths matter more than the headline number.
MiMo-V2-Pro's sharpest advantage is in coding, where it averages 78 against 62.7.
Category breakdown
Exact category averages are shown below. Not measured means BenchLM does not have enough sourced public coverage for that model and category.
| Category | GPT-5.6 Luna | Δ | MiMo-V2-Pro |
|---|---|---|---|
| Coding | GPT-5.6 Luna62.7 | Margin→ 15.3 | MiMo-V2-Pro78.0 |
| Agentic | GPT-5.6 Luna84.1 | MarginNo overlap | MiMo-V2-ProNot measured |
| Knowledge | GPT-5.6 Luna92.3 | MarginNo overlap | MiMo-V2-ProNot measured |
| Math | GPT-5.6 Luna73.6 | MarginNo overlap | MiMo-V2-ProNot measured |
| Multimodal | GPT-5.6 Luna78.4 | MarginNo overlap | MiMo-V2-ProNot measured |
Operational comparison
Runtime and commercial metrics are compared only when both models have a complete sourced value.
| Metric | GPT-5.6 Luna | MiMo-V2-Pro | Comparison |
|---|---|---|---|
| Input / output priceUSD per 1M tokens | GPT-5.6 Luna$1 input / $6 output | MiMo-V2-ProNot available | A complete price comparison is not available. |
| Generation speedtokens per second | GPT-5.6 LunaNot available | MiMo-V2-ProNot available | A complete speed comparison is not available. |
| First-answer latencyseconds to first token | GPT-5.6 LunaNot available | MiMo-V2-ProNot available | A complete latency comparison is not available. |
| Context windowmaximum listed tokens | GPT-5.6 Luna1M | MiMo-V2-Pro1M | Listed context windows are equal. |
Benchmark Deep Dive
Agentic19 benchmarks
| Benchmark | GPT-5.6 Luna | MiMo-V2-Pro | Result |
|---|---|---|---|
| Terminal-Bench 2.0Source | 84.7% | — | Not comparable |
| BrowseCompSource | 83.3% | — | Not comparable |
| OSWorld 2.0Source | 45.6% | — | Not comparable |
| CyberGymSource | 77.9% | — | Not comparable |
| ExploitGymSource | 12.4% | — | Not comparable |
| ToolathlonSource | 53.4% | — | Not comparable |
| AA Agentic IndexSource | 45.6% | — | Not comparable |
| GDPval-AASource | 54.2% | — | Not comparable |
| GDPval-AASource | 1584 | — | Not comparable |
| AA Harvey LABSource | 87.9% | — | Not comparable |
| AA ITBenchSource | 40.3% | — | Not comparable |
| AA Tau3 BankingSource | 27.2% | — | Not comparable |
| AA AutomationBenchSource | 42.2% | — | Not comparable |
| aaTerminalBench21Source | 80.9% | — | Not comparable |
| APEX-Agents-AASource | 35.8% | — | Not comparable |
| Claw-EvalSource | — | 57.8% | Not comparable |
| τ²-bench resultsSource | — | 95% | Not comparable |
| Gert LabsSource | — | 36.68% | Not comparable |
| ResearchClawBenchSource | — | 15.3% | Not comparable |
CodingMiMo-V2-Pro wins8 benchmarks
| Benchmark | GPT-5.6 Luna | MiMo-V2-Pro | Result |
|---|---|---|---|
| SWE-bench ProSource | 62.7% | — | Not comparable |
| Terminal-Bench 2.0Source | 84.7% | — | Not comparable |
| deepSweSource | 67.2% | — | Not comparable |
| FrontierCode 1.1 ExtendedSource | 55.1% | — | Not comparable |
| cursorBench32Source | 61.1% | — | Not comparable |
| AA Coding IndexSource | 71.5% | — | Not comparable |
| AA-SciCodeSource | 52.5% | 42.5% | GPT-5.6 Luna leads |
| SWE-bench VerifiedSource | — | 78% | Not comparable |
Reasoning3 benchmarks
Knowledge10 benchmarks
| Benchmark | GPT-5.6 Luna | MiMo-V2-Pro | Result |
|---|---|---|---|
| GPQASource | 92.3% | — | Not comparable |
| GPQA-DSource | 92.3% | — | Not comparable |
| HealthBench ProfessionalSource | 55.7% | — | Not comparable |
| HealthBench HardSource | 32.0% | — | Not comparable |
| Artificial Analysis Intelligence IndexSource | 51.2% | 40.3% | GPT-5.6 Luna leads |
| AA-GPQA DiamondSource | 91.1% | 87.0% | GPT-5.6 Luna leads |
| AA-HLESource | 37.2% | 28.3% | GPT-5.6 Luna leads |
| AA-Omniscience IndexSource | -11.2% | 4.9% | MiMo-V2-Pro leads |
| AA-Omniscience AccuracySource | 41.5% | 26.8% | GPT-5.6 Luna leads |
| AA-Omniscience Hallucination RateSource | 90.1% | 29.9% | MiMo-V2-Pro leads |
Math3 benchmarks
Multimodal3 benchmarks
Inst. Following1 benchmarks
| Benchmark | GPT-5.6 Luna | MiMo-V2-Pro | Result |
|---|---|---|---|
| AA-IFBenchSource | — | 68.8% | Not comparable |
Frequently Asked Questions (2)
Which is better, GPT-5.6 Luna or MiMo-V2-Pro?
MiMo-V2-Pro is ahead on BenchLM's BenchAlign leaderboard, 67.78 to 67.17.
Which is better for coding, GPT-5.6 Luna or MiMo-V2-Pro?
MiMo-V2-Pro has the edge for coding in this comparison, averaging 78 versus 62.7. Inside this category, AA-SciCode is the benchmark that creates the most daylight between them.
Related Comparisons
Explore More
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.