Model comparison
DeepSeek V4 Flash (High) vs MiMo-V2-Flash
Head-to-head evidence from 18 shared benchmark results across 5 categories. Overall scores shown here use the public BenchAlign v5 ranking lane.
Public leaderboard positions: DeepSeek V4 Flash (High) #92 (Estimated); MiMo-V2-Flash #91 (Supported). Intervals and evidence labels describe ranking uncertainty, not a guarantee for a specific workload.
Evidence parity. DeepSeek V4 Flash (High) and MiMo-V2-Flash share 18 comparable benchmark results. 2 of 8 categories are comparable. 20 results are unique to DeepSeek V4 Flash (High); 1 to MiMo-V2-Flash.
Updated July 23, 2026- Shared results
- 18
- DeepSeek V4 Flash (High) only
- 20
- MiMo-V2-Flash only
- 1
- Comparable categories
- 2 / 8
Pick MiMo-V2-Flash if you want the stronger benchmark profile. DeepSeek V4 Flash (High) only becomes the better choice if you need the larger 1M context window.
Confidence note. This is a partial-evidence comparison with 18 shared benchmark results across 5 evidence categories; 2 of 8 categories currently have scoreable aggregates for both models. Treat the verdict as directional until coverage is more balanced.
Why this result
MiMo-V2-Flash has the cleaner BenchAlign overall profile here, landing at 54.06 versus 53.95. It is a real lead, but still close enough that category-level strengths matter more than the headline number.
MiMo-V2-Flash's sharpest advantage is in knowledge, where it averages 84.7 against 52.1. The single biggest benchmark swing on the page is SWE-bench Verified, 78.6% to 73.4%.
DeepSeek V4 Flash (High) is also the more expensive model on tokens at $0.14 input / $0.28 output per 1M tokens, versus $0.00 input / $0.00 output per 1M tokens for MiMo-V2-Flash. That is roughly Infinityx on output cost alone. DeepSeek V4 Flash (High) gives you the larger context window at 1M, compared with 256K for MiMo-V2-Flash.
Category breakdown
Exact category averages are shown below. Not measured means BenchLM does not have enough sourced public coverage for that model and category.
| Category | DeepSeek V4 Flash (High) | Δ | MiMo-V2-Flash |
|---|---|---|---|
| Knowledge | DeepSeek V4 Flash (High)52.1 | Margin→ 32.6 | MiMo-V2-Flash84.7 |
| Coding | DeepSeek V4 Flash (High)68.5 | Margin→ 4.9 | MiMo-V2-Flash73.4 |
| Agentic | DeepSeek V4 Flash (High)55.3 | MarginNo overlap | MiMo-V2-FlashNot measured |
| Math | DeepSeek V4 Flash (High)91.9 | MarginNo overlap | MiMo-V2-FlashNot measured |
Decisive benchmark drivers
The largest measured benchmark gaps in this matchup, with exact reported values.
More
- Source ↗
SWE-bench Verified
CodingA 78.6%B 73.4%Winner: DeepSeek V4 Flash (High)Δ 5.2SWE-bench Verified: DeepSeek V4 Flash (High) scored 78.6%; MiMo-V2-Flash scored 73.4%. DeepSeek V4 Flash (High) wins this benchmark. - Source ↗
GPQA
KnowledgeA 87.4%B 83.7%Winner: DeepSeek V4 Flash (High)Δ 3.7GPQA: DeepSeek V4 Flash (High) scored 87.4%; MiMo-V2-Flash scored 83.7%. DeepSeek V4 Flash (High) wins this benchmark. - Source ↗
MMLU-Pro
KnowledgeA 86.4%B 84.9%Winner: DeepSeek V4 Flash (High)Δ 1.5MMLU-Pro: DeepSeek V4 Flash (High) scored 86.4%; MiMo-V2-Flash scored 84.9%. DeepSeek V4 Flash (High) wins this benchmark.
Operational comparison
Runtime and commercial metrics are compared only when both models have a complete sourced value.
| Metric | DeepSeek V4 Flash (High) | MiMo-V2-Flash | Comparison |
|---|---|---|---|
| Input / output priceUSD per 1M tokens | DeepSeek V4 Flash (High)$0.14 input / $0.28 output | MiMo-V2-Flash$0 input / $0 output | MiMo-V2-Flash has the lower combined listed price. |
| Generation speedtokens per second | DeepSeek V4 Flash (High)Not available | MiMo-V2-Flash129 tok/s | A complete speed comparison is not available. |
| First-answer latencyseconds to first token | DeepSeek V4 Flash (High)Not available | MiMo-V2-Flash2.14 s | A complete latency comparison is not available. |
| Context windowmaximum listed tokens | DeepSeek V4 Flash (High)1M | MiMo-V2-Flash256K | DeepSeek V4 Flash (High) lists the larger context window. |
Benchmark Deep Dive
Agentic9 benchmarks
| Benchmark | DeepSeek V4 Flash (High) | MiMo-V2-Flash | Result |
|---|---|---|---|
| Terminal-Bench 2.0Source | 56.6% | — | Not comparable |
| BrowseCompSource | 53.5% | — | Not comparable |
| HLE w/ toolsSource | 40.3% | — | Not comparable |
| MCP AtlasSource | 67.4% | — | Not comparable |
| ToolathlonSource | 43.5% | — | Not comparable |
| τ²-bench resultsSource | 95.6% | 83.9% | DeepSeek V4 Flash (High) leads |
| AA Agentic IndexSource | 28.2% | 12.0% | DeepSeek V4 Flash (High) leads |
| GDPval-AASource | 32.4% | 16.7% | DeepSeek V4 Flash (High) leads |
| GDPval-AASource | 1147 | 833 | DeepSeek V4 Flash (High) leads |
CodingMiMo-V2-Flash wins7 benchmarks
| Benchmark | DeepSeek V4 Flash (High) | MiMo-V2-Flash | Result |
|---|---|---|---|
| CodeforcesSource | 2816.0 | — | Not comparable |
| SWE-bench VerifiedSource | 78.6% | 73.4% | DeepSeek V4 Flash (High) leads |
| SWE-bench ProSource | 52.3% | — | Not comparable |
| SWE MultilingualSource | 70.2% | — | Not comparable |
| Terminal-Bench 2.0Source | 56.6% | — | Not comparable |
| AA-SciCodeSource | 42.0% | 25.9% | DeepSeek V4 Flash (High) leads |
| AA Coding IndexSource | 52.0% | 49.8% | DeepSeek V4 Flash (High) leads |
Reasoning4 benchmarks
KnowledgeMiMo-V2-Flash wins12 benchmarks
| Benchmark | DeepSeek V4 Flash (High) | MiMo-V2-Flash | Result |
|---|---|---|---|
| MMLU-ProSource | 86.4% | 84.9% | DeepSeek V4 Flash (High) leads |
| SimpleQASource | 28.9% | — | Not comparable |
| Chinese-SimpleQASource | 73.2% | — | Not comparable |
| GPQASource | 87.4% | 83.7% | DeepSeek V4 Flash (High) leads |
| GPQA-DSource | 87.4% | — | Not comparable |
| HLESource | 29.4% | — | Not comparable |
| Artificial Analysis Intelligence IndexSource | 37.5% | 24.7% | DeepSeek V4 Flash (High) leads |
| AA-GPQA DiamondSource | 86.7% | 65.6% | DeepSeek V4 Flash (High) leads |
| AA-HLESource | 27.8% | 8.0% | DeepSeek V4 Flash (High) leads |
| AA-Omniscience IndexSource | -22.3% | -48.5% | DeepSeek V4 Flash (High) leads |
| AA-Omniscience AccuracySource | 35.5% | 15.2% | DeepSeek V4 Flash (High) leads |
| AA-Omniscience Hallucination RateSource | 89.7% | 75.1% | MiMo-V2-Flash leads |
Math5 benchmarks
Multimodal1 benchmarks
| Benchmark | DeepSeek V4 Flash (High) | MiMo-V2-Flash | Result |
|---|---|---|---|
| Design Arena WebsiteSource | 1238 | — | Not comparable |
Inst. Following1 benchmarks
| Benchmark | DeepSeek V4 Flash (High) | MiMo-V2-Flash | Result |
|---|---|---|---|
| AA-IFBenchSource | 73.5% | 39.9% | DeepSeek V4 Flash (High) leads |
Frequently Asked Questions (3)
Which is better, DeepSeek V4 Flash (High) or MiMo-V2-Flash?
MiMo-V2-Flash is ahead on BenchLM's BenchAlign leaderboard, 54.06 to 53.95. The biggest single separator in this matchup is SWE-bench Verified, where the scores are 78.6% and 73.4%.
Which is better for knowledge tasks, DeepSeek V4 Flash (High) or MiMo-V2-Flash?
MiMo-V2-Flash has the edge for knowledge tasks in this comparison, averaging 84.7 versus 52.1. Inside this category, AA-Omniscience Index is the benchmark that creates the most daylight between them.
Which is better for coding, DeepSeek V4 Flash (High) or MiMo-V2-Flash?
MiMo-V2-Flash has the edge for coding in this comparison, averaging 73.4 versus 68.5. Inside this category, AA-SciCode is the benchmark that creates the most daylight between them.
Related Comparisons
- DeepSeek V4 Flash (High) vs DeepSeek V4 Pro (Max)
- DeepSeek V4 Flash (High) vs DeepSeek V4 Pro (High)
- DeepSeek V4 Flash (High) vs DeepSeek V4 Flash (Max)
- DeepSeek V4 Flash (High) vs DeepSeek V4 Pro
- DeepSeek V4 Flash (High) vs DeepSeek V4 Pro Base
- DeepSeek V4 Flash (High) vs DeepSeek V4 Flash Base
- DeepSeek V4 Flash (High) vs DeepSeek V4 Flash
- DeepSeek V4 Flash (High) vs Claude Mythos 5
Explore More
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.