LLM API pricing in 2026 spans more than 600x: GPT-5 nano is the cheapest major API at $0.05 per million input tokens, while GPT-5.5 Pro and GPT-5.4 Pro top out at $30/$180. Claude Opus 4.8 costs $5/$25 and Claude Fable 5 $10/$50. Gemini 3.5 Flash currently leads the budget-capability set at 75.02 Supported for $1.50/$9.
Pricing varies by more than 600x across major LLM APIs, from $0.05 to $30 per million input tokens. The right model for your workload depends on the task, volume, and how much quality you're trading for cost. This guide covers current pricing for every major model and breaks down the math for the most common use cases.
All prices are per million tokens. We keep the full sortable comparison on the LLM pricing page; this guide is about how to read it. For where prices are heading, see the Token Price Index.
Full price table (July 2026)
The 25 highest-scoring priced models, cheapest first. We regenerate this table from the live pricing catalog on every site build.
| Model | Creator | Input | Output | Context | Overall Score |
|---|---|---|---|---|---|
| GPT-5.4 nano | OpenAI | $0.2 | $1.25 | 400K | 67 |
| MiniMax M3 | MiniMax | $0.3 | $1.2 | 1M | 70 |
| GPT-5.6 Luna | OpenAI | $1 | $6 | 1M | 67 |
| GLM-5 | Z.AI | $1 | $3.2 | 200K | 66 |
| GLM-5-Turbo | Z.AI | $1.2 | $4 | 200K | 67 |
| Grok 4.3 | xAI | $1.25 | $2.5 | 1M | 65 |
| GLM-5.1 | Z.AI | $1.4 | $4.4 | 203K | 68 |
| GPT-5.3 Codex | OpenAI | $1.75 | $14 | 400K | 67 |
| Inkling | Thinking Machines Lab | $1.87 | $4.68 | 1M | 68 |
| Grok 4.5 | xAI | $2 | $6 | 500K | 77 |
| Gemini 3 Pro | $2 | $12 | 2M | 68 | |
| Claude Sonnet 5 | Anthropic | $2 | $10 | 1M | 65 |
| GPT-5.4 | OpenAI | $2.5 | $15 | 1.05M | 74 |
| GPT-5.6 Terra | OpenAI | $2.5 | $15 | 1M | 73 |
| Kimi K3 | Moonshot AI | $3 | $15 | 1.05M | 81 |
| Claude Sonnet 4.6 | Anthropic | $3 | $15 | 200K | 65 |
| GPT-5.6 Sol | OpenAI | $5 | $30 | 1M | 82 |
| Claude Opus 4.8 | Anthropic | $5 | $25 | 1M | 78 |
| GPT-5.5 | OpenAI | $5 | $30 | 1M | 74 |
| Claude Opus 4.7 | Anthropic | $5 | $25 | 1M | 72 |
| Claude Opus 4.6 | Anthropic | $5 | $25 | 1M | 69 |
| Claude Opus 4.7 (Adaptive) | Anthropic | $5 | $25 | 1M | 66 |
| Claude Mythos 5 | Anthropic | $10 | $50 | 1M+ | 84 |
| Claude Fable 5 | Anthropic | $10 | $50 | 1M+ | 84 |
| GPT-5.2 Pro | OpenAI | $25 | $150 | 400K | 67 |
Benchmark scores from the leaderboard. Prices per million tokens.
The table sorts into three cost tiers.
Under $0.50/M input: Nano and flash models. GPT-5 nano, Gemini 3.1 Flash-Lite, DeepSeek V3, Grok 3 Mini. Best for high-volume, lower-stakes tasks: classification, summarization, simple Q&A. Quality varies significantly.
$1-3/M input: The production sweet spot. Gemini 3.5 Flash ($1.50), GPT-5.1 ($1.25), GPT-5.4 ($2.50), Claude Sonnet 4.6 ($3.00). Strong capability at rates many teams can sustain.
$5-30/M input: Flagship tier. Claude Opus 4.6 ($5), GPT-5.2 Pro ($25), GPT-5.4 Pro ($30). Reserved for tasks where the extra capability is worth the price: legal analysis, complex research, high-stakes decisions.
Cost by use case
Chat and Q&A (1M tokens/month budget)
At $2.50/M input, GPT-5.4 gives you ~400K input tokens per month per $1 of input budget. For a typical chat application averaging 500 input tokens per message, that's 800 conversations per dollar. At that scale, GPT-5.4 and Claude Sonnet 4.6 ($3.00) are both reasonable choices.
If you're handling 10M+ tokens/month, the difference between $2.50 and $15.00/M input tokens becomes $125K/year at that volume. That's where the flagship vs mid-tier decision really matters.
Coding assistance
For a coding assistant or IDE integration:
- High-volume autocomplete: Gemini 3.5 Flash ($1.50/$9) or another low-latency model that passes your completion tests.
- Code review and refactoring: Start with the live coding ranking, then compare accepted patches and repair loops.
- Agentic coding: GPT-5.6 Sol ($5/$30) is third agentic at 75.49 Supported; Claude Fable 5 ($10/$50) is second at 76.84. The small capability gap and large price gap deserve a workflow trial.
Document processing (per document cost)
Assuming a 10-page document ≈ 4,000 tokens input, 500 tokens output:
| Model | Cost per doc |
|---|---|
| Gemini 3.1 Flash-Lite | $0.0018 |
| DeepSeek V3 | $0.0014 |
| Gemini 3.1 Pro | $0.0140 |
| GPT-5.4 | $0.0175 |
| Claude Sonnet 4.6 | $0.0195 |
| Claude Opus 4.6 | $0.098 |
For document pipelines processing thousands of documents per day, model selection has a direct P&L impact. Gemini 3.1 Pro at about $0.014/doc vs Claude Opus at about $0.045/doc is still a meaningful cost difference.
The value alternatives to frontier pricing
If you need GPT-5-class quality without GPT-5 pricing:
Gemini 3.5 Flash ($1.50/$9) scores 75.02 overall with Supported evidence, slightly above GPT-5.4's 74.18 while costing less. The two models have different capability profiles, so the score is a shortlist signal rather than proof of interchangeability.
Claude Sonnet 4.6 ($3/$15) scores 68 overall, lower than GPT-5.4, but still a viable option for writing, coding, and structured output if you prefer Anthropic's ecosystem.
GPT-5.2 ($1.75/$14) scores 77 overall, lower than GPT-5.4, but still a viable value row for teams that care about cost more than absolute frontier standing.
The starkest version of the trade is DeepSeek V3 at $0.27/$1.10 against GPT-5.4 at $2.50/$15:
- Input: GPT-5.4 is 9x more expensive
- Output: GPT-5.4 is 14x more expensive
For a pipeline generating 1M output tokens per day:
- DeepSeek V3: ~$1.10/day → $400/year
- GPT-5.4: ~$15/day → $5,475/year
The question is whether GPT-5.4's benchmark advantage justifies the 14x output cost premium. For general text generation, creative writing, and many coding tasks, probably not. For hard reasoning, agentic workflows, and tasks requiring frontier-level reliability, the benchmark gap is real.
DeepSeek R1 (the reasoning model at $0.55/$2.19) vs GPT-5.4 Pro ($30/$180) is an even starker comparison: ~80x cheaper on output tokens. The quality gap on hard reasoning is real but not 80x worth of quality.
What to use for your budget
Free tier / experiments: GPT-5 nano, Gemini 3.1 Flash-Lite Serious prototypes: DeepSeek V3, Gemini 3 Flash Small production apps (under $500/mo): Gemini 3.5 Flash, GPT-5.1, Claude Sonnet 4.6 Scale production (high volume): Gemini 3.5 Flash, MiniMax M3, or a routed mix Enterprise / high-stakes workflows: Claude Fable 5, GPT-5.6 Sol, Claude Opus 4.8
→ Use the Cost Calculator to estimate your monthly spend · Full pricing table
Reader questions
Frequently asked questions
01What is the cheapest LLM API in 2026?
GPT-5 nano is the cheapest major LLM API at $0.05 per million input tokens and $0.40 per million output tokens. Gemini 3.1 Flash-Lite now follows at $0.25/$1.50. For higher-quality tasks on a budget, DeepSeek V3 at $0.27/$1.10 offers strong performance at a fraction of Claude or GPT-5.4 pricing.
02How much does GPT-5.4 cost?
GPT-5.4 costs $2.50 per million input tokens and $15 per million output tokens. GPT-5.4 Pro (the flagship) costs $30 per million input tokens and $180 per million output tokens — more than 10x the cost of the standard GPT-5.4.
03How much does Claude Opus 4.6 cost?
Claude Opus 4.6 costs $5 per million input tokens and $25 per million output tokens. This still makes it materially more expensive than GPT-5.4 on output tokens. Claude Sonnet 4.6 offers a cheaper alternative at $3/$15 per million tokens.
04Is DeepSeek cheaper than GPT-5?
Yes, significantly. DeepSeek V3 costs $0.27/$1.10 per million tokens compared to GPT-5.4 at $2.50/$15. DeepSeek R1 (reasoning model) costs $0.55/$2.19 compared to GPT-5.4 Pro at $30/$180. For applications where DeepSeek's quality is sufficient, the cost savings are 5-100x depending on which GPT-5 tier you're comparing to.
05What is the best value LLM for production use in 2026?
Gemini 3.5 Flash is the strongest model priced at or below $1.50 per million input tokens, at 75.02 overall with Supported evidence and a $1.50/$9 rate. MiniMax M3 is the open-weight value leader at 69.8 and $0.30/$1.20. The best value still depends on workload success, output length, and review cost.
Continue with live BenchLM data
Share or save