Skip to main content

Best Budget LLMs in 2026: GPT-5.4 Mini, Nano, MiniMax M2.7, and Every Cheap Model Ranked

Which budget LLM should you use in 2026? We rank GPT-5.4 mini, GPT-5.4 nano, MiniMax M2.7, Claude Haiku 4.5, Gemini Flash, DeepSeek, and more by benchmarks and price.

Published
Last updated
Reading time
12 min
External sources
0
Tags: budget, comparison, pricing, guideData and scoring methodology
In this article7 sections

The budget leaderboard moved. Gemini 3.5 Flash now leads the models priced at or below $1.50 per million input tokens. MiniMax M3 leads the open-weight options. GPT-5.4 nano ranks above GPT-5.4 mini overall, an order that looks strange until evidence breadth enters the calculation.

This guide ranks every major LLM under $1.50 per million input tokens by benchmark performance, with pricing breakdowns and use-case recommendations. All scores come from our leaderboard and pricing page.

The budget tier landscape (July 2026)

There are more than 15 models priced at or below $1.50/M input. This table is generated from the current pricing catalog and BenchAlign ranking, so a price or score refresh changes the post with the leaderboard.

Table 1
Rank Model Type License Score Evidence Input / 1M
1 MiniMax M3 Non-Reasoning Open Weight 69.8 Supported $0.3
2 GLM-5.1 Reasoning Open Weight 67.7 Supported $1.4
3 GPT-5.6 Luna Reasoning Proprietary 67.2 Estimated $1
4 GLM-5-Turbo Reasoning Proprietary 66.9 Supported $1.2
5 GPT-5.4 nano Reasoning Proprietary 66.8 Supported $0.2
6 GLM-5 Non-Reasoning Open Weight 66.1 Supported $1
7 Grok 4.3 Reasoning Proprietary 65.1 Supported $1.25
8 Gemini 3.5 Flash Reasoning Proprietary 64.8 Estimated $1.5
9 MiniMax M2.7 Non-Reasoning Open Weight 64.1 Supported $0.3
10 GLM-5.2 Reasoning Open Weight 64 Estimated $1.4
11 GLM-5V-Turbo Non-Reasoning Proprietary 63.5 Supported $1.2
12 DeepSeek V4 Pro Non-Reasoning Open Weight 60.7 Supported $0.435
13 Gemini 3 Flash Non-Reasoning Proprietary 60.5 Supported $0.5
14 GLM-5 (Reasoning) Reasoning Open Weight 59.8 Estimated $1
15 Kimi K2.5 Non-Reasoning Open Weight 59.7 Supported $0.6

The score is not a value judgment by itself. It is the capability estimate; the adjacent evidence label shows whether the row has enough independent support to carry normally or needs the wider uncertainty attached to Estimated evidence.

GPT-5.4 mini and nano: OpenAI on a budget

GPT-5.4 mini is OpenAI's reasoning model at budget pricing: $0.75/M input, 3.3x cheaper than GPT-5.4. Its current public score appears in the live table above; the model has a 400K context window.

Where mini stands out:

  • Agentic tasks: OSWorld-Verified 72.2 (vs full GPT-5.4's 85), Terminal-Bench 2.0 at 60, Tau2Bench 93.4. These are strong agentic scores for a budget model: OSWorld 72.2 beats Claude Haiku 4.5 (57) and Gemini 3 Flash (53) comfortably.
  • Knowledge: GPQA 88 is solid, above Gemini 3 Flash (69) and Claude Haiku 4.5 (67), though below Gemini 3.1 Pro (97). HLE 41.5 is a standout on a benchmark where most budget models score in single digits.
  • Multimodal: MMMU-Pro 76.6 is competitive, trailing only Gemini 3.1 Pro (95) and Claude Haiku 4.5 (82) in the budget tier.

Where mini falls short:

Coding is the big one: SWE-bench Pro 54.4 is a long way down from GPT-5.4's 85, so for a coding-focused workload the quality gap is real. MRCR v2 at 40.7 shows the long-context reasoning ceiling (full GPT-5.4 scores 97 here), and with no AIME or HMMT scores published yet, math capability is hard to evaluate at all.

The pitch: GPT-5.4 mini makes sense only after it wins a trial on the work you actually send. The broad data no longer supports calling it the default budget model.

One tier further down, GPT-5.4 nano costs $0.20/M input, 12.5x cheaper than full GPT-5.4. It lands at 66.72 overall with Supported evidence and materially outperforms the older GPT-5 nano budget row.

Key scores:

Table 2
Benchmark GPT-5.4 nano GPT-5 nano GPT-5.4 mini
GPQA 82.8 71.2 88
HLE 37.7 41.5
SWE-bench Pro 52.4 22 54.4
Terminal-Bench 2.0 46.3 38 60
OSWorld-Verified 39 30 72.1
MMMU-Pro 66.1 58 76.6

GPT-5.4 nano beats GPT-5 nano on every available benchmark, especially coding (SWE-bench Pro 52.4 vs 22) and knowledge (GPQA 82.8 vs 71.2). The gap is large enough that GPT-5.4 nano effectively replaces GPT-5 nano for anything beyond the cheapest possible classification tasks.

The cost math: At $0.20/M input, nano processes 5 million input tokens per dollar. For a classification pipeline handling 100M tokens/month, GPT-5.4 nano costs $20/month. GPT-5.4 mini would cost $75/month for the same volume. That 3.75x multiplier matters at scale.

Where nano makes sense: High-volume tasks where cost dominates: classification, tagging, simple extraction, content filtering. For anything requiring strong reasoning or coding, the step up to mini ($0.75) is worth the extra cost.

MiniMax M2.7: the coding wildcard

MiniMax M2.7 is the surprise of this batch. At $0.30/M input (cheaper than both GPT-5.4 mini and nano for quality coding) it posts the highest SWE-bench Pro score in the budget tier: 56.22.

Table 3
Benchmark MiniMax M2.7 GPT-5.4 mini GPT-5.4 nano Claude Haiku 4.5
SWE-bench Pro 56.22 54.4 52.4 46
Terminal-Bench 2.0 57 60 46.3 53
SWE-Multilingual 76.5
MLE-Bench-Lite 66.6
Toolathlon 46.3 42.9 35.5

MiniMax M2.7 beats GPT-5.4 mini on SWE-bench Pro by nearly 2 points while costing 2.5x less on input tokens. On SWE-Multilingual (76.5) and MLE-Bench-Lite (66.6), it shows strong coding breadth that the OpenAI budget models haven't been tested on yet.

The caveat: MiniMax M2.7 now has Supported evidence and a 64.03 overall score, but its successor, MiniMax M3, is the stronger open-weight budget default at 69.8. M2.7 can still be the cheaper workload-specific choice.

200K context is another differentiator. At $0.30/M input, feeding large codebases into M2.7 is dramatically cheaper than any alternative with comparable SWE-bench scores.

Head-to-head: which budget model wins?

Coding

Table 4
Model SWE-bench Pro LiveCodeBench Price (in/out)
MiniMax M2.7 56.22 $0.30/$1.20
GPT-5.4 mini 54.4 $0.75/$4.50
GPT-5.4 nano 52.4 $0.20/$1.25
Claude Haiku 4.5 46 36 $1.00/$5.00
Gemini 3 Flash 44 36 $0.50/$3.00
DeepSeek V3 37.6 $0.27/$1.10

MiniMax M2.7 leads. For budget coding workloads (code review, bug fixing, refactors) it's the best value option in the tier. GPT-5.4 mini is close behind with the added benefit of being a reasoning model.

Agentic tasks

Table 5
Model Terminal-Bench 2.0 OSWorld-Verified Price (in/out)
GPT-5.4 mini 60 72.2 $0.75/$4.50
MiniMax M2.7 57 $0.30/$1.20
Gemini 3 Flash 56 53 $0.50/$3.00
Claude Haiku 4.5 41 57 $1.00/$5.00
GPT-5.4 nano 46.3 39 $0.20/$1.25
GPT-5 nano 38 30 $0.05/$0.40

GPT-5.4 mini dominates agentic benchmarks in this tier. OSWorld-Verified 72.2 is a standout: closer to full GPT-5.4 (85) than any other budget model gets to its flagship sibling. If you're building an agent on a budget, mini is the pick.

Knowledge

Table 6
Model GPQA HLE Price (in/out)
GPT-5.4 mini 88 41.5 $0.75/$4.50
GPT-5.4 nano 82.8 37.7 $0.20/$1.25
GPT-5 nano 71.2 $0.05/$0.40
Gemini 3 Flash 69 6 $0.50/$3.00
Claude Haiku 4.5 67 11 $1.00/$5.00
DeepSeek V3 59.1 $0.27/$1.10

GPT-5.4 mini and nano dominate knowledge benchmarks in the budget tier. HLE scores of 41.5 and 37.7 are particularly impressive; Claude Haiku 4.5 scores 11 and Gemini 3 Flash scores 6 on the same benchmark.

Multimodal

Table 7
Model MMMU-Pro Price (in/out)
Claude Haiku 4.5 82 $1.00/$5.00
Gemini 3 Flash 80 $0.50/$3.00
GPT-5.4 mini 76.6 $0.75/$4.50
GPT-5.4 nano 66.1 $0.20/$1.25
GPT-5 nano 58 $0.05/$0.40

Claude Haiku 4.5 and Gemini 3 Flash lead the budget tier on multimodal. MiniMax M2.7 has no MMMU-Pro score, another gap in its benchmark coverage.

When to use what

High-volume classification and tagging: GPT-5 nano ($0.05/$0.40) or GPT-5.4 nano ($0.20/$1.25). If you're processing millions of tokens daily on simple tasks, nano-tier pricing is hard to argue with. GPT-5.4 nano is substantially better on quality if the 4x price increase fits your budget.

Budget coding assistant: MiniMax M2.7 ($0.30/$1.20). Highest SWE-bench Pro in the tier (56.22) at the second-lowest price. The 200K context window handles large codebases well. The caveat: limited benchmark coverage outside coding, so evaluate on your specific tasks.

Budget AI agent: GPT-5.4 mini ($0.75/$4.50). OSWorld-Verified 72.2 and Terminal-Bench 60 are the best agentic scores in the budget tier by a wide margin. The reasoning capability helps with multi-step agent workflows.

Long-context workloads: Gemini 3 Flash ($0.50/$3.00) with 1M context, or GPT-5.4 mini ($0.75/$4.50) with 400K. Context length still needs a retrieval-quality test; the largest advertised window is not automatically the best long-document system.

Best budget all-rounder: Gemini 3.5 Flash. It leads the current budget set at 75.02 overall with Supported evidence. MiniMax M3 is the open-weight alternative at 69.8.

Cheapest reasoning model: GPT-5.4 nano ($0.20/$1.25). The only reasoning model under $0.50/M input with broad benchmark coverage. GPQA 82.8 and HLE 37.7 show real reasoning capability at an ultra-budget price.

The data gap problem

MiniMax M2.7's overall score still does not tell the full story. BenchAlign v5 estimates the model from the evidence it has, preserves uncertainty where it does not, and separates Supported from Estimated rows instead of assigning a zero for an unreported benchmark.

On the direct benchmarks that exist, M2.7 remains competitive with GPT-5.4 mini. SWE-bench Pro 56.22 and Terminal-Bench 57 are useful signals. They should inform a coding trial, not stand in for a general-purpose verdict.

This is a recurring problem in AI benchmarking. As we covered in the benchmark reliability explainer, benchmark coverage and provenance matter as much as the scores themselves. A model with 10 strong scores and 20 unknowns is a riskier choice than a model with 25 moderate scores.

The practical takeaway: If your workload is coding, M2.7's published rows justify trying it. For a broad open-weight default, start with MiniMax M3 and measure both on your own acceptance tests.

What this means for the market

Three takeaways from this week's releases:

1. The reasoning gap is closing at the bottom. GPT-5.4 mini and nano bring reasoning-class capability to the budget tier. A year ago, reasoning models started at $2.50/M input. Now you can get HLE 37.7 for $0.20/M input.

2. Chinese models keep punching above on coding. MiniMax M2.7 posting the highest SWE-bench Pro score in the budget tier (above both GPT-5.4 mini and nano) continues the trend of Chinese labs producing strong coding models at aggressive price points.

3. Budget doesn't mean weak anymore. GPT-5.4 mini's OSWorld-Verified 72.2 would have been a frontier-class score 12 months ago. The models that cost $0.30–$0.75/M input today are materially better than the $15/M models of early 2025.

Check the live leaderboard for the latest scores as more benchmarks roll in for these models. Prices and capabilities shift fast; what's budget today is obsolete tomorrow.

Reader questions

Frequently asked questions

01What is the best budget LLM in 2026?

Gemini 3.5 Flash leads models priced at or below $1.50 per million input tokens with a 75.02 BenchAlign score and Supported evidence. MiniMax M3 is the strongest open-weight budget option at 69.8, while GPT-5.4 nano is the strongest ultra-budget OpenAI row at 66.72.

02How does GPT-5.4 mini compare to GPT-5.4?

Full GPT-5.4 ranks above GPT-5.4 mini on the current public leaderboard. Mini costs $0.75/$4.50 versus $2.50/$15, so the right choice depends on whether the full model's measured advantage survives your workload-specific evaluation.

03Is GPT-5.4 nano good enough for production?

GPT-5.4 nano scores 66.72 overall with Supported evidence and costs $0.20/$1.25. It is a credible high-volume production option for classification, extraction, summarization, and simple Q&A; use task-specific evaluation before assigning it difficult coding or autonomous work.

04Is MiniMax M2.7 available outside China?

MiniMax M2.7 is available via API with international access at $0.30/$1.20 per million tokens and a 200K context window. Its successor, MiniMax M3, has replaced it as the stronger open-weight budget default on the current overall ranking.

05What is the cheapest LLM API in 2026?

GPT-5 nano is the cheapest major LLM API at $0.05 per million input tokens and $0.40 per million output tokens. Seed 1.6 Flash ($0.08/$0.30) and Gemini 3.1 Flash-Lite ($0.10/$0.40) are close behind. These ultra-budget models are best suited for high-volume, lower-stakes tasks where cost matters more than peak quality.

06Should I use Gemini 3.5 Flash or GPT-5.4 mini?

Gemini 3.5 Flash ranks above GPT-5.4 mini on the current public leaderboard and is the better broad default. Mini is cheaper on input, so it can still win after a workload-specific evaluation.

Share or save

Share on XShare on LinkedIn

Keep reading

All research

New models drop every week. Join 2,000+ readers for one email a week on what moved, why, and what still needs proof.