Kimi K2.6 Pricing Hides a Precision Trade
Kimi K2.6 lists at $0.95 in and $4.00 out, but 21 hosts serve it from $0.549 to $1.20 at four different quantizations. Which route you get matters.
Table of Contents
Kimi K2.6 lists at $0.95 per million input tokens and $4.00 output with a 262,144 token context. Almost nobody pays that, because twenty-one different companies serve the model and the cheapest of them charges $0.549 in and $2.313 out, roughly 42% below list on both sides. The catch is not hidden in the price column. It is in the one next to it: those hosts run the same weights at four different numeric precisions, and the cheapest full-precision route costs less than eight of the routes charging full list. All figures were read from OpenRouter's public endpoints API on 17 August 2026.
Key Takeaways:
- Kimi K2.6 is a multimodal Moonshot AI model for long-horizon coding and multi-agent orchestration, with a 262,144 token context and text plus image input.
- Twenty-one endpoints serve it. Input runs $0.549 to $1.200 per million, a 2.18x spread; output runs $2.313 to $4.600, a 1.99x spread.
- Precision varies by host: fp4, int4, fp8, and bf16 all appear. Crusoe is the only bf16 route at $0.70 / $3.50, which undercuts the eight hosts charging list price on int4 or fp4.
- Moonshot AI's own endpoint charges list, $0.95 / $4.00, and serves int4.
- Unlike Grok 4.6, none of the 21 endpoints carries a context-based price override. One rate applies to the whole window.
What Kimi K2.6 is and what it lists at
Kimi K2.6 is Moonshot AI's model aimed at long-horizon coding, coding-driven UI generation, and multi-agent orchestration, handling end-to-end tasks across languages including Python, Rust, and Go. It takes text and images and returns text. It has been on OpenRouter since 20 April 2026, which by the standards of this year makes it an established model rather than a new one. Moonshot's own developer platform is the first-party reference for the family.
The context window is 262,144 tokens. That is a quarter of what its larger sibling offers and it is the first thing to check against your workload, because it is the constraint you cannot buy your way around on this model.
Twenty-one hosts and a 2.2x spread on the same model
Open weights mean anyone with the hardware can serve the model, and twenty-one companies currently do. Here is the cheap end and the expensive end of that list.
| Provider | Input / 1M | Output / 1M | Precision | Context |
|---|---|---|---|---|
| Decart | $0.549 | $2.313 | fp4 | 262,144 |
| Chutes | $0.580 | $3.400 | int4 | 262,144 |
| StreamLake | $0.598 | $2.520 | fp8 | 256,000 |
| CoreWeave | $0.650 | $3.410 | fp4 | 262,144 |
| Crusoe | $0.700 | $3.500 | bf16 | 262,144 |
| DigitalOcean | $0.760 | $3.200 | unknown | 262,144 |
| Moonshot AI | $0.950 | $4.000 | int4 | 262,144 |
| Sail Research | $1.000 | $4.000 | int4 | 262,144 |
| Together | $1.200 | $4.500 | unknown | 262,144 |

On input that is a 2.18x range for identical weights. On output it is 1.99x. Over a month of steady agent traffic that difference is not a rounding error, and unlike the context-based price cliff on Grok 4.6, none of it depends on how long your prompt happens to be. There are no overrides anywhere in the K2.6 endpoint list, so one rate covers the whole window.
The precision column almost nobody reads
Sort that table by price and you learn something uncomfortable.
The four cheapest routes are fp4, int4, fp8, and fp4. Those are compressed copies of the model, quantized down from the original weights to fit more throughput onto less hardware. Quantization is a legitimate engineering trade and often an invisible one, but it is a trade, and at four bits it is not always invisible on the tasks people buy this model for, which are long coding runs where a small drop in instruction-following compounds across forty tool calls.
Now find the full-precision option. There is exactly one, Crusoe at bf16, and it charges $0.70 / $3.50. That is cheaper than Moonshot AI's own endpoint at $0.95 / $4.00, which serves int4. It is cheaper than Baidu, Cloudflare, Fireworks, BaseTen, and AtlasCloud, all at list price on int4 or fp4. It is cheaper than Together at $1.20.

So the ranking that matters is not the one the price column produces. You can pay more for fewer bits, and on this model most of the list-price routes do exactly that. Nothing about that is deceptive; every host publishes its quantization and OpenRouter surfaces it. It is just that the sorting control everyone reaches for is price, and price and precision are not correlated here.
One more number puts the list price in perspective. OpenRouter reports the weighted average customers actually pay for K2.6 as $0.3354 input and $3.539 output, which is below even the cheapest endpoint's headline rate, because caching and provider discounts sit underneath the posted numbers. The $0.95 on the model page is close to a ceiling, not an estimate.

Six of the twenty-one report their quantization as unknown, which is its own answer. If precision matters to your workload, treat unknown as "do not route here" rather than as "probably fine".
Why a Kimi K2.6 price increase alert is usually not one
This model generates false alarms, and the mechanism is worth understanding because it applies to every open-weight model on a routing platform.
A monitor that diffs prices day over day can read two different numbers for K2.6 depending on what it looks at. The model-level figure is $0.95 / $4.00. Individual endpoints range from $0.549 up to $1.200. If yesterday's reading captured one endpoint and today's captured the model-level list, the diff shows a jump of about 46% on input that never happened to anyone.
We hit exactly this last week. A pricing alert reported K2.6 moving from $0.65 / $3.41 to $0.95 / $4.00. Both numbers were real, and neither was a change: $0.65 / $3.41 is CoreWeave's endpoint and $0.95 / $4.00 is the list rate that eight hosts including Moonshot's own charge. Nothing had moved in either place. The comparison had.
The rule that falls out of it is simple. Compare an endpoint against the same endpoint, or the list against the list. Never one against the other. If you are alerting on prices, key your monitor on the provider slug, not on the model ID.
Kimi K2.6 against Kimi K3, and when the cheap sibling wins
The two models are further apart than one decimal place suggests.
| Kimi K2.6 | Kimi K3 | |
|---|---|---|
| List input / output | $0.95 / $4.00 | $3.00 / $15.00 |
| Cheapest route | $0.549 / $2.313 | $2.60 / $13.00 |
| Context | 262,144 | 1,048,576 |
| Input types | text, image | text, image, video |
| Endpoints | 21 | 13 |
On list prices K3 costs 3.2x more on input and 3.75x more on output. At the cheap end of each list the gap is wider still. K3 buys you four times the context window, video input, and a larger model, and whether that is worth roughly four times the money is entirely a question about your workload rather than about the models.
The honest split: if your task fits in 262,144 tokens and does not involve video, K2.6 is the default and K3 is the upgrade you justify rather than assume. If your agent regularly runs past a quarter of a million tokens, no price on K2.6 helps you, because you will be truncating. Our side-by-side on K3 versus K2.6 goes deeper on the capability difference, and what Kimi K3 is covers the larger model on its own terms.
Picking a Kimi K2.6 route on purpose
Three decisions, in the order that saves the most money per minute spent.
Decide your precision floor first, before you look at any price. If the answer is "full precision or nothing", your list is one entry long and you are done. If four-bit is acceptable for your workload, verify that on your own evals rather than on a vendor's, and then the cheap end of the list opens up.
Pin the provider, do not sort by price. Sorting is a rule that re-runs every request and will happily move you between precisions mid-workload. Pinning is a decision you made once. OpenRouter's provider.only and provider.order fields do this in one JSON block, documented under provider routing.
Set a spend ceiling anyway. provider.max_price turns an unpredictable bill into a visible failure, which is the cheaper of the two problems.
All three of those live in configuration, and configuration that lives on a laptop drifts the moment you pick the work up somewhere else. That is the flat, unglamorous reason to run this class of job on a hosted machine such as MoClaw, where the pins and the caps sit in one environment that is still there tomorrow.
What the endpoint list cannot tell you
Everything above is metadata. It tells you what each host claims to serve and what they charge for it. It does not tell you whether the int4 copy at Chutes actually performs worse than the bf16 copy at Crusoe on your tasks, and I have not measured that.
That gap is worth naming clearly, because the intuitive conclusion from this article would be "always buy bf16", and the intuitive conclusion is not supported by anything here. Quantized serving is often indistinguishable in practice. The claim I am making is narrower: you are currently choosing precision by accident, the choice is not priced consistently, and if it matters to you then you should be making it deliberately.
Uptime figures on the endpoint list move by the half hour, so treat any single reading of them as a snapshot rather than a track record.
FAQ
How much does Kimi K2.6 cost?
The list price is $0.95 per million input tokens and $4.00 per million output. Across the 21 hosts serving it on OpenRouter as of 17 August 2026, actual rates run from $0.549 / $2.313 at the cheap end to $1.200 / $4.500 at the expensive end.
What is Kimi K2.6's context window?
262,144 tokens on almost every host. Two of them, StreamLake and Venice, serve 256,000 instead, and BaseTen reports 262,000. The differences are small but they are real if you are running right at the ceiling.
Is Kimi K2.6 cheaper than Kimi K3?
Substantially. K3 lists at $3.00 / $15.00 against K2.6's $0.95 / $4.00, so roughly 3.2x on input and 3.75x on output. K3 buys a 1,048,576 token context and video input for that.
Which provider should I use for Kimi K2.6?
It depends on whether quantization matters to you. Crusoe is the only bf16 route at $0.70 / $3.50. If four-bit serving is acceptable, Decart at $0.549 / $2.313 is the cheapest, and Chutes, CoreWeave, and Inceptron sit just above it.
Did Kimi K2.6 get more expensive recently?
Not as of 17 August 2026. Reports of a jump from around $0.65 to $0.95 are comparing an individual provider's rate against the model-level list rate. Both figures are current and neither represents a change.
Buy a route, not a model name
The useful mental shift with any open-weight model is that "Kimi K2.6" is not a thing you can purchase. It is a set of weights that twenty-one companies will run for you at their own price, at their own precision, on their own hardware. What you actually buy is one of those twenty-one arrangements.
So spend the five minutes. Open the endpoint list, decide what precision your work needs, pick the cheapest host that clears that bar, and pin it. On this model that process is worth up to 2.18x on input, and it will occasionally tell you that the expensive option and the good option are not the same option.
Continue Reading
More GuideThe MoClaw editorial team writes about workflow automation, AI agents, and the tools we build. Default byline for industry overviews, listicles, and collaborative pieces.
Ready to put this into practice?
MoClaw runs browser tasks, research, and schedules automatically. Try it free.
References: Kimi K2.6 on OpenRouter · Kimi K3 on OpenRouter · Moonshot AI developer platform · OpenRouter provider routing documentation · Moonshot AI on Hugging Face