Skip to main content

LLM Pricing 2026: Every Model, $0.11–$50 per 1M Tokens

How to actually compare LLM API pricing — what blended cost means, which discounts matter (caching, batch), and how to find the cheapest model for your workload. With current rates for GPT-5, Claude, Gemini, DeepSeek, and more.

Published
Last reviewed
Data as of
Reading time
8 min
External sources
0
Tags: pricing, comparison, cost, apiData and scoring methodology
In this article4 sections

LLM API pricing in 2026 spans more than 600x: GPT-5 nano is the cheapest major API at $0.05 per million input tokens, while GPT-5.5 Pro and GPT-5.4 Pro top out at $30/$180. Claude Opus 4.8 costs $5/$25 and Claude Fable 5 $10/$50. Gemini 3.5 Flash currently leads the budget-capability set at 75.02 Supported for $1.50/$9.

Pricing varies by more than 600x across major LLM APIs, from $0.05 to $30 per million input tokens. The right model for your workload depends on the task, volume, and how much quality you're trading for cost. This guide covers current pricing for every major model and breaks down the math for the most common use cases.

All prices are per million tokens. We keep the full sortable comparison on the LLM pricing page; this guide is about how to read it. For where prices are heading, see the Token Price Index.

Full price table (July 2026)

The 25 highest-scoring priced models, cheapest first. We regenerate this table from the live pricing catalog on every site build.

Table 1
Model Creator Input Output Context Overall Score
GPT-5.4 nano OpenAI $0.2 $1.25 400K 67
MiniMax M3 MiniMax $0.3 $1.2 1M 70
GPT-5.6 Luna OpenAI $1 $6 1M 67
GLM-5 Z.AI $1 $3.2 200K 66
GLM-5-Turbo Z.AI $1.2 $4 200K 67
Grok 4.3 xAI $1.25 $2.5 1M 65
GLM-5.1 Z.AI $1.4 $4.4 203K 68
GPT-5.3 Codex OpenAI $1.75 $14 400K 67
Inkling Thinking Machines Lab $1.87 $4.68 1M 68
Grok 4.5 xAI $2 $6 500K 77
Gemini 3 Pro Google $2 $12 2M 68
Claude Sonnet 5 Anthropic $2 $10 1M 65
GPT-5.4 OpenAI $2.5 $15 1.05M 74
GPT-5.6 Terra OpenAI $2.5 $15 1M 73
Kimi K3 Moonshot AI $3 $15 1.05M 81
Claude Sonnet 4.6 Anthropic $3 $15 200K 65
GPT-5.6 Sol OpenAI $5 $30 1M 82
Claude Opus 4.8 Anthropic $5 $25 1M 78
GPT-5.5 OpenAI $5 $30 1M 74
Claude Opus 4.7 Anthropic $5 $25 1M 72
Claude Opus 4.6 Anthropic $5 $25 1M 69
Claude Opus 4.7 (Adaptive) Anthropic $5 $25 1M 66
Claude Mythos 5 Anthropic $10 $50 1M+ 84
Claude Fable 5 Anthropic $10 $50 1M+ 84
GPT-5.2 Pro OpenAI $25 $150 400K 67

Benchmark scores from the leaderboard. Prices per million tokens.

The table sorts into three cost tiers.

Under $0.50/M input: Nano and flash models. GPT-5 nano, Gemini 3.1 Flash-Lite, DeepSeek V3, Grok 3 Mini. Best for high-volume, lower-stakes tasks: classification, summarization, simple Q&A. Quality varies significantly.

$1-3/M input: The production sweet spot. Gemini 3.5 Flash ($1.50), GPT-5.1 ($1.25), GPT-5.4 ($2.50), Claude Sonnet 4.6 ($3.00). Strong capability at rates many teams can sustain.

$5-30/M input: Flagship tier. Claude Opus 4.6 ($5), GPT-5.2 Pro ($25), GPT-5.4 Pro ($30). Reserved for tasks where the extra capability is worth the price: legal analysis, complex research, high-stakes decisions.

Cost by use case

Chat and Q&A (1M tokens/month budget)

At $2.50/M input, GPT-5.4 gives you ~400K input tokens per month per $1 of input budget. For a typical chat application averaging 500 input tokens per message, that's 800 conversations per dollar. At that scale, GPT-5.4 and Claude Sonnet 4.6 ($3.00) are both reasonable choices.

If you're handling 10M+ tokens/month, the difference between $2.50 and $15.00/M input tokens becomes $125K/year at that volume. That's where the flagship vs mid-tier decision really matters.

Coding assistance

For a coding assistant or IDE integration:

  • High-volume autocomplete: Gemini 3.5 Flash ($1.50/$9) or another low-latency model that passes your completion tests.
  • Code review and refactoring: Start with the live coding ranking, then compare accepted patches and repair loops.
  • Agentic coding: GPT-5.6 Sol ($5/$30) is third agentic at 75.49 Supported; Claude Fable 5 ($10/$50) is second at 76.84. The small capability gap and large price gap deserve a workflow trial.

Document processing (per document cost)

Assuming a 10-page document ≈ 4,000 tokens input, 500 tokens output:

Table 2
Model Cost per doc
Gemini 3.1 Flash-Lite $0.0018
DeepSeek V3 $0.0014
Gemini 3.1 Pro $0.0140
GPT-5.4 $0.0175
Claude Sonnet 4.6 $0.0195
Claude Opus 4.6 $0.098

For document pipelines processing thousands of documents per day, model selection has a direct P&L impact. Gemini 3.1 Pro at about $0.014/doc vs Claude Opus at about $0.045/doc is still a meaningful cost difference.

The value alternatives to frontier pricing

If you need GPT-5-class quality without GPT-5 pricing:

Gemini 3.5 Flash ($1.50/$9) scores 75.02 overall with Supported evidence, slightly above GPT-5.4's 74.18 while costing less. The two models have different capability profiles, so the score is a shortlist signal rather than proof of interchangeability.

Claude Sonnet 4.6 ($3/$15) scores 68 overall, lower than GPT-5.4, but still a viable option for writing, coding, and structured output if you prefer Anthropic's ecosystem.

GPT-5.2 ($1.75/$14) scores 77 overall, lower than GPT-5.4, but still a viable value row for teams that care about cost more than absolute frontier standing.

The starkest version of the trade is DeepSeek V3 at $0.27/$1.10 against GPT-5.4 at $2.50/$15:

  • Input: GPT-5.4 is 9x more expensive
  • Output: GPT-5.4 is 14x more expensive

For a pipeline generating 1M output tokens per day:

  • DeepSeek V3: ~$1.10/day → $400/year
  • GPT-5.4: ~$15/day → $5,475/year

The question is whether GPT-5.4's benchmark advantage justifies the 14x output cost premium. For general text generation, creative writing, and many coding tasks, probably not. For hard reasoning, agentic workflows, and tasks requiring frontier-level reliability, the benchmark gap is real.

DeepSeek R1 (the reasoning model at $0.55/$2.19) vs GPT-5.4 Pro ($30/$180) is an even starker comparison: ~80x cheaper on output tokens. The quality gap on hard reasoning is real but not 80x worth of quality.

What to use for your budget

Free tier / experiments: GPT-5 nano, Gemini 3.1 Flash-Lite Serious prototypes: DeepSeek V3, Gemini 3 Flash Small production apps (under $500/mo): Gemini 3.5 Flash, GPT-5.1, Claude Sonnet 4.6 Scale production (high volume): Gemini 3.5 Flash, MiniMax M3, or a routed mix Enterprise / high-stakes workflows: Claude Fable 5, GPT-5.6 Sol, Claude Opus 4.8

Use the Cost Calculator to estimate your monthly spend · Full pricing table


Reader questions

Frequently asked questions

01What is the cheapest LLM API in 2026?

GPT-5 nano is the cheapest major LLM API at $0.05 per million input tokens and $0.40 per million output tokens. Gemini 3.1 Flash-Lite now follows at $0.25/$1.50. For higher-quality tasks on a budget, DeepSeek V3 at $0.27/$1.10 offers strong performance at a fraction of Claude or GPT-5.4 pricing.

02How much does GPT-5.4 cost?

GPT-5.4 costs $2.50 per million input tokens and $15 per million output tokens. GPT-5.4 Pro (the flagship) costs $30 per million input tokens and $180 per million output tokens — more than 10x the cost of the standard GPT-5.4.

03How much does Claude Opus 4.6 cost?

Claude Opus 4.6 costs $5 per million input tokens and $25 per million output tokens. This still makes it materially more expensive than GPT-5.4 on output tokens. Claude Sonnet 4.6 offers a cheaper alternative at $3/$15 per million tokens.

04Is DeepSeek cheaper than GPT-5?

Yes, significantly. DeepSeek V3 costs $0.27/$1.10 per million tokens compared to GPT-5.4 at $2.50/$15. DeepSeek R1 (reasoning model) costs $0.55/$2.19 compared to GPT-5.4 Pro at $30/$180. For applications where DeepSeek's quality is sufficient, the cost savings are 5-100x depending on which GPT-5 tier you're comparing to.

05What is the best value LLM for production use in 2026?

Gemini 3.5 Flash is the strongest model priced at or below $1.50 per million input tokens, at 75.02 overall with Supported evidence and a $1.50/$9 rate. MiniMax M3 is the open-weight value leader at 69.8 and $0.30/$1.20. The best value still depends on workload success, output length, and review cost.

Share or save

Share on XShare on LinkedIn

Keep reading

All research

Model pricing changes frequently. Join 2,000+ readers for one email a week on what moved, why, and what still needs proof.