Model comparison
GPT-5.4 mini vs Grok 4.20
Head-to-head evidence from 4 shared benchmark results across 4 categories. Overall scores shown here use the public BenchAlign v5 ranking lane.
Public leaderboard positions: GPT-5.4 mini #75 (Estimated); Grok 4.20 #88 (Estimated). Intervals and evidence labels describe ranking uncertainty, not a guarantee for a specific workload.
Evidence parity. GPT-5.4 mini and Grok 4.20 share 4 comparable benchmark results. 2 of 8 categories are comparable. 26 results are unique to GPT-5.4 mini; 14 to Grok 4.20.
Updated July 23, 2026- Shared results
- 4
- GPT-5.4 mini only
- 26
- Grok 4.20 only
- 14
- Comparable categories
- 2 / 8
Pick GPT-5.4 mini if you want the stronger benchmark profile. Grok 4.20 only becomes the better choice if you need the larger 2M context window.
Confidence note. This is a partial-evidence comparison with 4 shared benchmark results across 4 evidence categories; 2 of 8 categories currently have scoreable aggregates for both models. Treat the verdict as directional until coverage is more balanced.
Why this result
GPT-5.4 mini has the cleaner BenchAlign overall profile here, landing at 56.77 versus 54.68. It is a real lead, but still close enough that category-level strengths matter more than the headline number.
GPT-5.4 mini's sharpest advantage is in agentic, where it averages 65.7 against 47.1. The single biggest benchmark swing on the page is Terminal-Bench 2.0, 60% to 47.1%.
Grok 4.20 is also the more expensive model on tokens at $2.00 input / $6.00 output per 1M tokens, versus $0.75 input / $4.50 output per 1M tokens for GPT-5.4 mini. Grok 4.20 gives you the larger context window at 2M, compared with 400K for GPT-5.4 mini.
Category breakdown
Exact category averages are shown below. Not measured means BenchLM does not have enough sourced public coverage for that model and category.
| Category | GPT-5.4 mini | Δ | Grok 4.20 |
|---|---|---|---|
| Agentic | GPT-5.4 mini65.7 | Margin← 18.6 | Grok 4.2047.1 |
| Multimodal | GPT-5.4 mini76.6 | Margin← 6.5 | Grok 4.2070.1 |
| Coding | GPT-5.4 miniNot measured | MarginNo overlap | Grok 4.2067.1 |
| Reasoning | GPT-5.4 miniNot measured | MarginNo overlap | Grok 4.2053.3 |
| Knowledge | GPT-5.4 mini47.8 | MarginNo overlap | Grok 4.20Not measured |
| Math | GPT-5.4 mini21.7 | MarginNo overlap | Grok 4.20Not measured |
Decisive benchmark drivers
The largest measured benchmark gaps in this matchup, with exact reported values.
More
- Source ↗
Terminal-Bench 2.0
AgenticA 60%B 47.1%Winner: GPT-5.4 miniΔ 12.9Terminal-Bench 2.0: GPT-5.4 mini scored 60%; Grok 4.20 scored 47.1%. GPT-5.4 mini wins this benchmark. - Source ↗
MMMU-Pro
MultimodalA 76.6%B 75.2%Winner: GPT-5.4 miniΔ 1.4MMMU-Pro: GPT-5.4 mini scored 76.6%; Grok 4.20 scored 75.2%. GPT-5.4 mini wins this benchmark.
Operational comparison
Runtime and commercial metrics are compared only when both models have a complete sourced value.
| Metric | GPT-5.4 mini | Grok 4.20 | Comparison |
|---|---|---|---|
| Input / output priceUSD per 1M tokens | GPT-5.4 mini$0.75 input / $4.5 output | Grok 4.20$2 input / $6 output | GPT-5.4 mini has the lower combined listed price. |
| Generation speedtokens per second | GPT-5.4 mini201 tok/s | Grok 4.20233 tok/s | Grok 4.20 has the higher measured throughput. |
| First-answer latencyseconds to first token | GPT-5.4 mini3.85 s | Grok 4.2010.33 s | GPT-5.4 mini reaches the first token sooner. |
| Context windowmaximum listed tokens | GPT-5.4 mini400K | Grok 4.202M | Grok 4.20 lists the larger context window. |
Benchmark Deep Dive
AgenticGPT-5.4 mini wins11 benchmarks
| Benchmark | GPT-5.4 mini | Grok 4.20 | Result |
|---|---|---|---|
| Terminal-Bench 2.0Source | 60% | 47.1% | GPT-5.4 mini leads |
| OSWorld-VerifiedSource | 72.1% | — | Not comparable |
| MCP AtlasSource | 57.7% | — | Not comparable |
| ToolathlonSource | 42.9% | — | Not comparable |
| τ²-bench resultsSource | 83.3% | — | Not comparable |
| AA Agentic IndexSource | 30.2% | — | Not comparable |
| APEX-Agents-AASource | 28.2% | — | Not comparable |
| GDPval-AASource | 33.6% | — | Not comparable |
| GDPval-AASource | 1171 | — | Not comparable |
| DeepSearchQASource | — | 62.8% | Not comparable |
| Gert LabsSource | — | 38.36% | Not comparable |
Coding7 benchmarks
| Benchmark | GPT-5.4 mini | Grok 4.20 | Result |
|---|---|---|---|
| Vibe Code BenchSource | 47.97% | 4.06% | GPT-5.4 mini leads |
| AA Coding IndexSource | 56.1% | — | Not comparable |
| AA-SciCodeSource | 49.9% | — | Not comparable |
| FrontierCode 1.1 MainSource | 27.0% | — | Not comparable |
| LiveCodeBench ProSource | — | 74.2% | Not comparable |
| SWE-bench VerifiedSource | — | 76.7% | Not comparable |
| SWE-bench ProSource | — | 51.8% | Not comparable |
Reasoning3 benchmarks
Knowledge12 benchmarks
| Benchmark | GPT-5.4 mini | Grok 4.20 | Result |
|---|---|---|---|
| GPQASource | 88% | — | Not comparable |
| HLESource | 41.5% | — | Not comparable |
| HLE w/o toolsSource | 28.2% | 31.6% | Grok 4.20 leads |
| Artificial Analysis Intelligence IndexSource | 40.0% | — | Not comparable |
| AA-GPQA DiamondSource | 87.5% | — | Not comparable |
| AA-HLESource | 26.6% | — | Not comparable |
| AA-Omniscience IndexSource | -18.7% | — | Not comparable |
| AA-Omniscience AccuracySource | 37.5% | — | Not comparable |
| AA-Omniscience Hallucination RateSource | 89.8% | — | Not comparable |
| GPQA-DSource | — | 88.5% | Not comparable |
| HealthBench HardSource | — | 20.3% | Not comparable |
| MedXpertQA (Text)Source | — | 50.2% | Not comparable |
Math2 benchmarks
MultimodalGPT-5.4 mini wins8 benchmarks
| Benchmark | GPT-5.4 mini | Grok 4.20 | Result |
|---|---|---|---|
| MMMU-ProSource | 76.6% | 75.2% | GPT-5.4 mini leads |
| MMMU-Pro w/ PythonSource | 78% | — | Not comparable |
| AA-MMMU-ProSource | 73.3% | — | Not comparable |
| CharXivSource | — | 60.9% | Not comparable |
| ERQASource | — | 54.1% | Not comparable |
| SimpleVQASource | — | 57.4% | Not comparable |
| MedXpertQA (MM)Source | — | 65.8% | Not comparable |
| Design Arena WebsiteSource | — | 1257 | Not comparable |
Inst. Following1 benchmarks
| Benchmark | GPT-5.4 mini | Grok 4.20 | Result |
|---|---|---|---|
| AA-IFBenchSource | 73.3% | — | Not comparable |
Frequently Asked Questions (3)
Which is better, GPT-5.4 mini or Grok 4.20?
GPT-5.4 mini is ahead on BenchLM's BenchAlign leaderboard, 56.77 to 54.68. The biggest single separator in this matchup is Terminal-Bench 2.0, where the scores are 60% and 47.1%.
Which is better for agentic tasks, GPT-5.4 mini or Grok 4.20?
GPT-5.4 mini has the edge for agentic tasks in this comparison, averaging 65.7 versus 47.1. Inside this category, Terminal-Bench 2.0 is the benchmark that creates the most daylight between them.
Which is better for multimodal and grounded tasks, GPT-5.4 mini or Grok 4.20?
GPT-5.4 mini has the edge for multimodal and grounded tasks in this comparison, averaging 76.6 versus 70.1. Inside this category, MMMU-Pro is the benchmark that creates the most daylight between them.
Related Comparisons
Explore More
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.