| Provider | Input / 1M | Output / 1M | Monthly |
|---|
Calculating Kimi API Costs
Kimi models from Moonshot AI bill per token, with one twist that matters more than any other AI provider: the same model is sold at meaningfully different prices depending on where you buy it. Enter your request size and monthly volume above and the calculator prices your workload at Moonshot’s official rates, then shows what the identical usage costs on other providers below the result.
A worked example: a coding agent on Kimi K2.7 Code sending 100,000 input tokens and generating 10,000 output tokens per run costs $0.135 per run at official rates with no cache hits. The same run on the new Kimi K3 flagship costs $0.45 – $0.30 for the input and $0.15 for the output. At 10,000 runs a month that is $1,350 on K2.7 Code against $4,500 on K3.
Kimi K3 and K2 Pricing
Official Moonshot international platform rates, checked July 17, 2026. Kimi K3, released July 16, is the new flagship with a 1,048,576-token context window; the K2 models carry 262,144 tokens.
| Model | Input (cache hit) | Input (cache miss) | Output | Context |
|---|---|---|---|---|
| kimi-k3 | $0.30 | $3.00 | $15.00 | 1M |
| kimi-k2.7-code | $0.19 | $0.95 | $4.00 | 256K |
| kimi-k2.7-code-highspeed | $0.38 | $1.90 | $8.00 | 256K |
| kimi-k2.6 | $0.16 | $0.95 | $4.00 | 256K |
| kimi-k2.5 | $0.10 | $0.60 | $3.00 | 256K |
Per 1 million tokens, USD, taxes excluded and calculated at checkout. HighSpeed is the same K2.7 Code model served faster at double the price; there is no HighSpeed variant for K3.
K3 is Moonshot’s flagship for long-horizon coding, knowledge work, and reasoning; K2.7 Code stays on as the dedicated coding model at less than a third of K3’s input price, with K2.6 as the general-purpose model and K2.5 as the budget option. K3’s pricing is flat across the full 1M-token window – there is no tiering by context length, unlike GPT-5.6 and Gemini 3.1 Pro, which both charge roughly double once a prompt passes 200K tokens.
Two launch details worth knowing. Moonshot is running a top-up rebate from July 15 to August 12, 2026: a single top-up of $20 or more earns a voucher of 10% to 30% depending on the amount, valid for 90 days. And Moonshot also operates a separate Chinese platform (platform.kimi.com) with its own RMB pricing and billing – the two platforms are not interchangeable even though the models are.
Reasoning Tokens on K3
K3 always reasons before it answers – thinking mode cannot be turned off, and the reasoning_effort setting currently supports only its maximum level. Those reasoning tokens bill as output at $15.00 per million, on top of the visible answer. For short prompts the reasoning can be several times longer than the answer itself, so when you enter an output figure in the calculator above, budget for the full generation, not just the text you expect to read. Moonshot’s default cap is 131,072 output tokens per request, configurable up to the full 1M. The K2 models let you switch thinking off; K3 does not.
Provider Price Differences
Because the K2 models are open-weight, third-party hosts run them and set their own prices.

K3 is a different story for now. It is a 2.8-trillion-parameter open-weight release, but the weights do not ship until July 27, 2026, so no third party hosts it yet: DeepInfra does not list the model, Together shows it as coming soon, and OpenRouter passes K3 requests straight through to Moonshot at the official $3.00 / $15.00 – with capacity limits producing frequent 429 errors in the launch week. Until independent hosts spin up, there is no cheaper place to buy K3 tokens.
For the K2 models the spread is real. Rates observed July 17, 2026, per million tokens:
| Provider | K3 in / out | K2.7 Code in / out | K2.6 in / out | K2.5 in / out |
|---|---|---|---|---|
| Moonshot (official) | $3.00 / $15.00 | $0.95 / $4.00 | $0.95 / $4.00 | $0.60 / $3.00 |
| OpenRouter | $3.00 / $15.00 | $0.72 / $3.50 | $0.66 / $3.41 | $0.375 / $2.025 |
| DeepInfra | not hosted | $0.74 / $3.50 | $0.75 / $3.50 | $0.45 / $2.25 |
| Together | coming soon | $0.95 / $4.00 | $1.20 / $4.50 | not listed |

Two caveats before switching providers on price alone. OpenRouter routes requests across underlying hosts, so its listed price reflects current routing and can change without notice. And third-party hosts run the open-weight release, which can lag the official platform on model updates, context handling, or caching behavior – cheap tokens from a host with no cache support can cost more in practice than official tokens at an 80% hit rate.
Cache Hit Pricing
Moonshot caches repeated context automatically on its platform. Cached input on K3 bills at $0.30 per million instead of $3.00 – a 10x reduction, twice as steep as K2.7 Code’s 5x ($0.19 against $0.95). It is the single biggest lever on an agent bill, where the repeated project context usually dwarfs the new instructions in each request. Multi-turn agent sessions typically land high hit rates; independent one-shot requests land near zero. The pattern is the same one that drives DeepSeek’s cache-hit pricing, though DeepSeek’s discount is steeper at 50x.
Kimi Prices Compared With Other APIs
K3 launched at exactly Claude Sonnet 5’s list price: $3.00 in, $15.00 out. Sonnet 5 is on introductory pricing of $2.00 / $10.00 through August 31, 2026, so until then K3 actually costs more per token than Anthropic’s mid-tier – after that they match. Against the other flagships, K3 undercuts GPT-5.6 Sol ($5.00 / $30.00) and Claude Opus 4.8 ($5.00 / $25.00), sits above Gemini 3.1 Pro ($2.00 / $12.00 under 200K tokens), and costs roughly 7 times DeepSeek V4 Pro’s input rate ($0.435 / $0.87). Rates below are provider list prices checked July 17, 2026:

Two structural differences matter beyond the sticker numbers. K3’s price is flat to 1M tokens of context, where GPT-5.6 Sol doubles ($10.00 / $45.00) and Gemini 3.1 Pro doubles ($4.00 / $18.00) past 200K – for long-context work the gap widens in K3’s favor. And K3’s always-on reasoning means its billed output runs higher than a non-reasoning model producing the same visible answer. For a Claude-side workload comparison with caching TTLs and batch discounts modeled properly, run the same numbers through the Claude API token cost calculator.
Keeping the Bill Down
Match the model to the job before optimizing anything else: K2.7 Code handles agentic coding at under a third of K3’s input price, K2.5 covers classification and reformatting, and K3 is for long-horizon reasoning and genuinely long context. Keep the stable part of your prompt byte-identical between requests so Moonshot’s cache catches it – at $0.30 against $3.00, an 80% hit rate cuts K3 input cost by 72%. Skip HighSpeed unless latency genuinely costs you money: it doubles every K2.7 Code token. If you are provider-shopping for K2 models, re-check prices at the source before committing – third-party rates in this market have moved monthly. And if you are waiting on cheaper K3, the weights land July 27; check back whether independent hosts have undercut Moonshot after that.