Skip to main content

Provider pricing hub

Kimi API Pricing (July 2026)

Last synced

Kimi K3 costs $3.00 per million cache-miss input tokens and $15.00 per million output tokens, with a $0.30 cache-hit input rate and a 1M-token context window. The Kimi API costs $0.95/$4.00 for Kimi K2.6 and the coding-tuned Kimi K2.7 Code, and $0.60/$3.00 for Kimi K2.5.

Moonshot's Kimi K3 report replaces the placeholder with an exact premium rate: $3.00 input, $0.30 cache-hit input, and $15.00 output per million tokens. Kimi K3 also raises the context ceiling to 1M. The previous flagship tier remains at $0.95/$4.00 (K2.6 and its coding sibling K2.7 Code) and a value tier at $0.60/$3.00 (K2.5). The older rows keep 256K context. The table below lists every Kimi model in the registry; exact rates remain tied to Moonshot's official platform pages and Kimi K3 technical report.

Price is only half the story: Kimi K3 is now the highest-ranked Kimi row on the overall leaderboard, and the family features prominently in the Chinese model rankings and the price-vs-performance view.

Every Kimi API price per 1M tokens

ModelInput $/MCached input $/MOutput $/MContext
Kimi K3$3$0.3$151.05M
Kimi 2.6$0.95$4256K
Kimi K2.7 Code$0.95$4256K
Kimi K2.5$0.6$3256K
Kimi K2$0.6$2.5128K

Moonshot publishes cache-hit input at $0.30/M for Kimi K3, $0.19/M for K2.7 Code, and $0.15/M for Kimi K2. Moonshot does not publish a Batch API discount tier, and K2.5's thinking and non-thinking modes share one price and are collapsed into one row.

Estimate your monthly Kimi API bill

ModelEst. monthly cost
Kimi K3$60
Kimi 2.6$17.5
Kimi K2.7 Code$17.5
Kimi K2.5$12
Kimi K2$11

Estimates use Kimi's standard per-token rates from the table above. For task-level presets and cross-provider comparison, use the full AI cost calculator.

How Moonshot prices the Kimi API

Kimi K3 is the new premium API tier at $3.00 input / $15.00 output per million tokens, with cache-hit input at $0.30. The previous flagship rate of $0.95/$4.00 still covers both K2.6 (the general model) and kimi-k2.7-code (the coding endpoint), so choosing the specialist costs nothing extra. K2.5 sits at $0.60/$3.00, and Moonshot explicitly documents that its thinking and non-thinking modes run under the same model at the same price — no separate reasoning surcharge to model.

The discount Moonshot publishes is cache-hit input: $0.30 per million tokens on Kimi K3, $0.19 on K2.7 Code, and $0.15 on the older Kimi K2. For coding agents that resend large repository context on every turn, that cache lane is where most of the real savings live. There is no Batch API tier — if you need asynchronous half-price processing, Kimi doesn't offer it yet.

Kimi vs DeepSeek and OpenAI pricing

Against its closest rival, the price ordering is now one-directional: at $0.435/$0.87, DeepSeek V4 Pro is about 54% cheaper than K2.6 ($0.95/$4.00) on input and 78% cheaper on output. That does not settle quality, latency, or deployment fit. For pure volume work, DeepSeek V4 Flash ($0.14/$0.28) undercuts every Kimi tier by a wider margin.

Kimi K3's $15 output rate matches GPT-5.6 Terra and is half GPT-5.6 Sol's $30 rate. K2.6 remains the cheaper Kimi endpoint at $4 output. See the full DeepSeek API pricing and OpenAI API pricing tables, or line models up head-to-head on the compare hub.

Which Kimi model to use

Kimi K3 is the capability pick. It scores 76 on the provisional overall leaderboard, 69.8 in coding, and 94.6 in agentic work. K2.6 remains close on coding at 69.5 while costing less. If your workload is coding agents specifically, K2.7 Codeis the same $0.95/$4.00 with a published $0.19 cache-hit input rate, though it doesn't yet have a ranked BenchLM score of its own.

K2.5 ($0.60/$3.00) is the value tier — 63 overall as the non-thinking row and 69 in reasoning mode, still with 256K context. The legacy Kimi K2 ($0.60/$2.50, 128K) remains on the price sheet mostly for existing integrations. And because the flagship weights are published, teams with GPUs can self-host instead of paying per token.

Kimi API pricing FAQ

How much does the Kimi API cost?

Kimi K3 costs $3.00 input and $15.00 output per million tokens, with cache-hit input at $0.30. Kimi K2.6 and Kimi K2.7 Code both cost $0.95/$4.00, while Kimi K2.5 costs $0.60/$3.00. Moonshot publishes no separate Batch API discount tier.

Is the Kimi K2 API free?

No — Moonshot's hosted Kimi API is pay-per-token. The genuinely free path is self-hosting: K2.6 is an open-weight release, so teams with their own GPUs can run it without per-token fees. It currently tops the open-weight rows on BenchLM's coding leaderboard; see the open-source model rankings for alternatives.

What is Moonshot AI's API pricing per million tokens?

Kimi K3 is $3 input and $15 output per million tokens. K2.6 and K2.7 Code cost $0.95/$4, and K2.5 costs $0.60/$3. Moonshot lists cache-hit input at $0.30 for Kimi K3 and $0.19 for K2.7 Code. There is no published Batch API discount tier.

How does Kimi pricing compare to DeepSeek?

DeepSeek V4 Pro ($0.435/$0.87) is about 54% cheaper than Kimi K2.6 ($0.95/$4.00) on input and 78% cheaper on output. Token price therefore favors DeepSeek at current direct rates, while model quality, latency, support, and deployment requirements can still favor Kimi. For high-volume pipelines DeepSeek V4 Flash ($0.14/$0.28) is cheaper again — see the DeepSeek pricing hub for the full table.

Keep comparing