Model comparison
Inkling vs Step 3.7 Flash
Head-to-head evidence from 18 shared benchmark results across 5 categories. Overall scores shown here use the public BenchAlign v5 ranking lane.
Public leaderboard positions: Inkling #20 (Supported); Step 3.7 Flash #110 (Estimated). Intervals and evidence labels describe ranking uncertainty, not a guarantee for a specific workload.
Evidence parity. Inkling and Step 3.7 Flash share 18 comparable benchmark results. 2 of 8 categories are comparable. 14 results are unique to Inkling; 11 to Step 3.7 Flash.
Updated July 23, 2026- Shared results
- 18
- Inkling only
- 14
- Step 3.7 Flash only
- 11
- Comparable categories
- 2 / 8
Pick Inkling if you want the stronger benchmark profile. Step 3.7 Flash only becomes the better choice if you want the cheaper token bill or you want the stronger reasoning-first profile.
Confidence note. This is a partial-evidence comparison with 18 shared benchmark results across 5 evidence categories; 2 of 8 categories currently have scoreable aggregates for both models. Treat the verdict as directional until coverage is more balanced.
Why this result
Inkling is clearly ahead on the BenchAlign aggregate, 67.54 to 50.87. The gap is large enough that you do not need to squint at the spreadsheet to see the difference.
Inkling's sharpest advantage is in coding, where it averages 68.6 against 56.3. The single biggest benchmark swing on the page is Terminal-Bench 2.0, 63.8% to 59.5%.
Inkling is also the more expensive model on tokens at $1.87 input / $4.68 output per 1M tokens, versus $0.20 input / $1.15 output per 1M tokens for Step 3.7 Flash. That is roughly 4.1x on output cost alone. Step 3.7 Flash is the reasoning model in the pair, while Inkling is not. That usually helps on harder chain-of-thought-heavy tests, but it can also mean more latency and more token spend in real use. Inkling gives you the larger context window at 1M, compared with 256K for Step 3.7 Flash.
Category breakdown
Exact category averages are shown below. Not measured means BenchLM does not have enough sourced public coverage for that model and category.
| Category | Inkling | Δ | Step 3.7 Flash |
|---|---|---|---|
| Coding | Inkling68.6 | Margin← 12.3 | Step 3.7 Flash56.3 |
| Agentic | Inkling69.4 | Margin← 3.0 | Step 3.7 Flash66.4 |
| Knowledge | Inkling51.6 | MarginNo overlap | Step 3.7 FlashNot measured |
| Math | Inkling97.1 | MarginNo overlap | Step 3.7 FlashNot measured |
| Multimodal | Inkling76.5 | MarginNo overlap | Step 3.7 FlashNot measured |
| Inst. Following | Inkling79.8 | MarginNo overlap | Step 3.7 FlashNot measured |
Decisive benchmark drivers
The largest measured benchmark gaps in this matchup, with exact reported values.
More
- Source ↗
Terminal-Bench 2.0
AgenticA 63.8%B 59.5%Winner: InklingΔ 4.3Terminal-Bench 2.0: Inkling scored 63.8%; Step 3.7 Flash scored 59.5%. Inkling wins this benchmark. - Source ↗
SWE-bench Pro
CodingA 54.3%B 56.3%Winner: Step 3.7 FlashΔ 2SWE-bench Pro: Inkling scored 54.3%; Step 3.7 Flash scored 56.3%. Step 3.7 Flash wins this benchmark. - Source ↗
BrowseComp
AgenticA 77.1%B 75.8%Winner: InklingΔ 1.3BrowseComp: Inkling scored 77.1%; Step 3.7 Flash scored 75.8%. Inkling wins this benchmark.
Operational comparison
Runtime and commercial metrics are compared only when both models have a complete sourced value.
| Metric | Inkling | Step 3.7 Flash | Comparison |
|---|---|---|---|
| Input / output priceUSD per 1M tokens | Inkling$1.87 input / $4.68 output | Step 3.7 Flash$0.2 input / $1.15 output | Step 3.7 Flash has the lower combined listed price. |
| Generation speedtokens per second | InklingNot available | Step 3.7 FlashNot available | A complete speed comparison is not available. |
| First-answer latencyseconds to first token | InklingNot available | Step 3.7 FlashNot available | A complete latency comparison is not available. |
| Context windowmaximum listed tokens | Inkling1M | Step 3.7 Flash256K | Inkling lists the larger context window. |
Benchmark Deep Dive
AgenticInkling wins16 benchmarks
| Benchmark | Inkling | Step 3.7 Flash | Result |
|---|---|---|---|
| Terminal-Bench 2.0Source | 63.8% | 59.5% | Inkling leads |
| BrowseCompSource | 77.1% | 75.8% | Inkling leads |
| MCP AtlasSource | 74.1% | — | Not comparable |
| Design Arena Agentic Web DevSource | 1257 | — | Not comparable |
| AA Agentic IndexSource | 32.3% | 21.5% | Inkling leads |
| GDPval-AASource | 36.9% | 25.9% | Inkling leads |
| GDPval-AASource | 1239 | 1017 | Inkling leads |
| AA BriefcaseSource | 836 | — | Not comparable |
| AA Tau3 BankingSource | 23.7% | — | Not comparable |
| DeepSearchQASource | — | 92.8% | Not comparable |
| ToolathlonSource | — | 49.5% | Not comparable |
| Claw-EvalSource | — | 67.1% | Not comparable |
| HLE w/ toolsSource | — | 47.2% | Not comparable |
| Gert LabsSource | — | 51.57% | Not comparable |
| τ²-bench resultsSource | — | 98.5% | Not comparable |
| APEX-Agents-AASource | — | 14.8% | Not comparable |
CodingInkling wins5 benchmarks
Reasoning2 benchmarks
Knowledge10 benchmarks
| Benchmark | Inkling | Step 3.7 Flash | Result |
|---|---|---|---|
| GPQASource | 87.9% | — | Not comparable |
| GPQA-DSource | 87.9% | — | Not comparable |
| HLESource | 46% | — | Not comparable |
| HLE w/o toolsSource | 30% | — | Not comparable |
| Artificial Analysis Intelligence IndexSource | 40.7% | 30.3% | Inkling leads |
| AA-GPQA DiamondSource | 87.2% | 80.9% | Inkling leads |
| AA-HLESource | 29.7% | 19.9% | Inkling leads |
| AA-Omniscience IndexSource | 2.1% | -37.5% | Inkling leads |
| AA-Omniscience AccuracySource | 40.0% | 25.4% | Inkling leads |
| AA-Omniscience Hallucination RateSource | 63.1% | 84.4% | Inkling leads |
Math1 benchmarks
| Benchmark | Inkling | Step 3.7 Flash | Result |
|---|---|---|---|
| AIME26Source | 97.1% | — | Not comparable |
Multimodal7 benchmarks
| Benchmark | Inkling | Step 3.7 Flash | Result |
|---|---|---|---|
| MMMU-ProSource | 73.5% | — | Not comparable |
| CharXivSource | 82% | — | Not comparable |
| CharXiv w/o toolsSource | 78.1% | — | Not comparable |
| AA-MMMU-ProSource | 73.5% | 75.3% | Step 3.7 Flash leads |
| SimpleVQASource | — | 79.2% | Not comparable |
| V*Source | — | 95.3% | Not comparable |
| Design Arena WebsiteSource | — | 1211 | Not comparable |
Frequently Asked Questions (3)
Which is better, Inkling or Step 3.7 Flash?
Inkling is ahead on BenchLM's BenchAlign leaderboard, 67.54 to 50.87. The biggest single separator in this matchup is Terminal-Bench 2.0, where the scores are 63.8% and 59.5%.
Which is better for coding, Inkling or Step 3.7 Flash?
Inkling has the edge for coding in this comparison, averaging 68.6 versus 56.3. Inside this category, AA Coding Index is the benchmark that creates the most daylight between them.
Which is better for agentic tasks, Inkling or Step 3.7 Flash?
Inkling has the edge for agentic tasks in this comparison, averaging 69.4 versus 66.4. Inside this category, GDPval-AA is the benchmark that creates the most daylight between them.
Related Comparisons
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.