Grok 4.6 Doubles Its Price at 200K Tokens

9 min read · · MoClaw Editorial
Grok 4.6 Doubles Its Price at 200K Tokens

Grok 4.6 lists a 500,000 token context at $2.00 in and $6.00 out, but every one of its four endpoints doubles that rate once a prompt passes 200,000.

Table of Contents

Share this

Grok 4.6 lists at $2.00 per million input tokens and $6.00 output, with a 500,000 token context window. What the headline rate does not tell you is that every one of its four OpenRouter endpoints carries an override: once a single prompt passes 200,000 tokens, the whole request bills at $4.00 and $12.00 instead. The model was listed on OpenRouter on 12 August 2026 under the underlying ID x-ai/grok-4.6-20260810, and all figures here were read from OpenRouter's public endpoints API on 17 August 2026.

Key Takeaways:

  • Grok 4.6 offers a 500,000 token context, but the advertised $2.00 / $6.00 rate only applies below 200,000 prompt tokens. Above that it is $4.00 / $12.00.
  • The cliff applies to cached input too: $0.50 per million cached read becomes $1.00 once you cross.
  • There are four endpoints and xAI runs all of them: standard, priority, zero data retention, and zero data retention with priority. Priority costs double at every level, so a long prompt on priority bills at $8.00 / $24.00.
  • A 201,000 token prompt costs about twice what a 199,000 token prompt costs. One percent more context, roughly one hundred percent more money.
  • Kimi K2.6 across 21 endpoints, Kimi K3 across 13, and Qwen 3.8 Max on its single endpoint carry no price overrides at all. This structure is Grok's, not the industry's.

What Grok 4.6 is, and what the 500K window lists at

Grok 4.6 is xAI's current flagship, described on its model page as their smartest model with frontier performance on coding, knowledge work, and STEM. It takes text, images, and files as input and returns text, and xAI's own model documentation is the reference for the family. The context window is 500,000 tokens, which puts it below the million-token tier that Kimi K3 and Qwen 3.8 Max occupy but well above most of what you would have been using a year ago.

One cosmetic detail worth noting because it confuses search results: OpenRouter displays the provider as SpaceXAI, while the provider_name field its API returns for the same endpoints is still xAI. If you are grepping logs or dashboards for a vendor string, expect both.

OpenRouter's provider rows for Grok 4.6 on 17 August 2026, showing the standard and zero-data-retention endpoints at $2.00 input and $6.00 output with cache read at $0.50, listed under the provider name SpaceXAI
OpenRouter's provider rows for Grok 4.6 on 17 August 2026, showing the standard and zero-data-retention endpoints at $2.00 input and $6.00 output with cache read at $0.50, listed under the provider name SpaceXAI

The part that matters for anyone budgeting a workload is not the window size. It is the shape of the bill underneath it.

A cost cliff you cross at 3am is one you find out about at 9.
Long agent runs grow their own context and nobody is watching the token counter at hour six. MoClaw is a hosted cloud AI computer that keeps the run and its logs in one place you can check from a phone, running alongside the setup you already have.
Watch the overnight run without staying up for it…Try MoClaw →

The price cliff at 200,000 prompt tokens

Pull the endpoint list for Grok 4.6 and each entry carries a pricing block with an overrides array. Every one of them says the same thing: at a minimum prompt size of 200,000 tokens, input goes to $4.00 per million and output to $12.00.

Run the arithmetic on two requests that differ by one percent.

199,000 token prompt 201,000 token prompt
Input rate $2.00 / M $4.00 / M
Input cost $0.398 $0.804
Output rate (2,000 tokens) $6.00 / M $12.00 / M
Output cost $0.012 $0.024
Total $0.41 $0.83

Two bars comparing one Grok 4.6 request at 199,000 prompt tokens costing $0.41 against the same request at 201,000 tokens costing $0.83, above a table of all four endpoint variants and their rates above and below the threshold
Two bars comparing one Grok 4.6 request at 199,000 prompt tokens costing $0.41 against the same request at 201,000 tokens costing $0.83, above a table of all four endpoint variants and their rates above and below the threshold

Two thousand extra tokens of context, roughly one paragraph of a file listing, and the request costs about twice as much. OpenRouter exposes the same block through provider routing, which is also where you cap what a request is allowed to spend. Nothing about the model changed. Nothing about your prompt got harder. You crossed a line in a pricing table.

There is a second-hand trace of this in OpenRouter's own numbers. Its pricing panel reports the weighted average price customers actually pay for Grok 4.6 as $0.8892 input and $6.317 output. Input landing far below the $2.00 list is what caching does, since cache reads bill at $0.50. Output landing above the $6.00 list is harder to explain any other way than some share of traffic billing at the over-threshold rate, which is what you would expect if real prompts routinely cross 200,000 tokens.

OpenRouter's pricing panel for Grok 4.6 on 17 August 2026, reporting a weighted average input price of $0.8892 and output price of $6.317 per million tokens against a listed $2.00 and $6.00
OpenRouter's pricing panel for Grok 4.6 on 17 August 2026, reporting a weighted average input price of $0.8892 and output price of $6.317 per million tokens against a listed $2.00 and $6.00

The cached-read rate moves with it. Cached input reads at $0.50 per million below the line and $1.00 above, which quietly undoes part of the reason you set up caching in the first place: the workloads with enough repeated context to benefit from a cache are exactly the workloads whose prompts are large.


Four endpoints, all of them xAI

The API returns four endpoint variants for Grok 4.6 and no third party serves any of them, because the weights are not public. The site lists the two standard rows; the priority forms of each show up in the endpoints payload. What differs between them is not the host but the terms.

Endpoint tag Below 200K Above 200K
xai $2.00 / $6.00 $4.00 / $12.00
xai/zdr $2.00 / $6.00 $4.00 / $12.00
xai/priority $4.00 / $12.00 $8.00 / $24.00
xai/zdr/priority $4.00 / $12.00 $8.00 / $24.00

The zdr variants are zero data retention, priced identically to the standard tier, which is a genuinely good deal compared with vendors who charge a premium for the same promise. Priority doubles the rate at both ends of the cliff, so a long prompt on the priority tier bills at four times the number in the headline.

All four report a 500,000 token context. The window does not shrink when the price rises, which is worth saying plainly because the natural assumption is that a cheaper tier must be a smaller one. It is not. You are paying more for the same window past a threshold.

Web search, if you enable it, is billed separately at $0.005.


Why agent workloads cross 200K without anyone noticing

A chat session rarely gets near 200,000 tokens. An agent does, and it gets there by accretion rather than by any single decision.

You hand it a repository and it reads twelve files. It runs the test suite and the output goes into context. It reads three more files based on what the failure said, keeps the diff it just wrote, and re-reads one file to check its own edit. Forty tool calls in, nobody typed anything long, and the prompt is 230,000 tokens because every step carried the previous steps with it.

That is the shape we covered in the context of agent workflows for the previous generation, and the accumulation behaviour has not changed. What changed is that Grok 4.6 now prices the second half of that curve differently from the first.

The practical consequence is that per-request cost stops being predictable from your prompt template. Two runs of the same agent on two different tickets can differ by 2x on rate alone, depending only on how many files the agent decided it needed to open.


Keeping a long Grok 4.6 run on the cheap side of the cliff

Three things are worth doing, and none of them requires giving up the model.

Measure where your prompts actually land. Log prompt token counts per request for a week before you tune anything. Most teams discover their distribution is bimodal: a large cluster well under 100,000 and a thin tail over 200,000 that quietly accounts for a disproportionate share of spend.

Compact deliberately rather than accidentally. If your harness summarises tool output when context gets large, set the trigger below 200,000 rather than at whatever the default is. The cheapest tokens are the ones you never resend.

Split the work when the split is natural. Two 150,000 token requests bill entirely at the low rate. One 300,000 token request bills entirely at the high one. Where the task genuinely decomposes, the decomposition is now worth money as well as being better engineering.

The monitoring half of this is the part people skip, because it needs something that is awake while the run is. MoClaw is a hosted cloud AI computer built for that: the run keeps going after your laptop closes, and the token counts land somewhere you can look at later instead of in a terminal you already quit.


How this compares to models with no cliff at all

I checked whether this is standard practice. It is not.

Kimi K2.6 is served by 21 endpoints and not one of them carries a price override. Kimi K3 across 13 endpoints, the same. Qwen 3.8 Max on its single Alibaba endpoint, the same. Those models bill one rate for one token regardless of how many tokens came before it.

That does not make Grok 4.6 expensive in absolute terms. Below 200,000 tokens it is $2.00 / $6.00, which undercuts Kimi K3 at $3.00 / $15.00 by a wide margin on output. The point is narrower and more annoying: Grok 4.6 is cheap or mid-priced depending on a variable your agent controls and you do not. If you are comparing it against alternatives on a spreadsheet, one rate in one cell will give you the wrong answer for half your traffic. For the earlier generation's rates we kept a running note in our Grok API pricing writeup, and the comparison against Anthropic's line is in Grok 4.5 vs Claude.


What is still unverified about Grok 4.6

Everything above is a pricing structure, not a quality judgment. I have not benchmarked the model, and the benchmark claims that exist come from xAI, published alongside the release rather than reproduced independently.

Two other things I could not settle. Whether the override applies to the whole request or only to the tokens above the threshold is not spelled out on the endpoint payload; the arithmetic in this article assumes the whole request, which is how OpenRouter's override structure reads, and it is what you should budget for until someone publishes a real invoice either way. And whether xAI intends this as a permanent structure or a launch-period setting, nobody outside xAI knows.

If you run Grok 4.6 at any volume, the cheapest experiment available is to send one 199,000 token request and one 201,000 token request and compare the two line items on your bill. That settles the question for your account in about a minute.


FAQ

How much does Grok 4.6 cost?

$2.00 per million input tokens and $6.00 per million output on the standard endpoint, as long as the prompt stays under 200,000 tokens. Above that it is $4.00 and $12.00. The priority endpoints are double at both tiers, so $4.00 / $12.00 below the threshold and $8.00 / $24.00 above it.

What is Grok 4.6's context window?

500,000 tokens on all four endpoints. The window does not change between the cheaper and more expensive tiers; only the rate does.

Who hosts Grok 4.6?

Only xAI. All four OpenRouter endpoints are xAI's own, in standard, priority, zero data retention, and zero data retention with priority variants. The model is closed, so no third party can serve it.

Does the zero data retention option cost extra?

No. The xai/zdr endpoint is priced identically to the standard one at $2.00 / $6.00, with the same 200,000 token override. The premium tier is priority, not privacy.

When was Grok 4.6 released?

It appeared on OpenRouter on 12 August 2026, with the underlying model ID x-ai/grok-4.6-20260810.


Read the override before you budget the window

The number in the pricing headline is the number for small prompts. That has always been sort of true and it is now literally true, with a threshold you can name.

So when you evaluate Grok 4.6, evaluate it twice. Once at $2.00 / $6.00 for the requests your agent keeps short, and once at $4.00 / $12.00 for the ones it does not. Then go look at your own logs and find out which of those two models you are actually buying, because on a long-running agent it is usually both.

Continue Reading

M
MoClaw Editorial MoClaw editorial team

The MoClaw editorial team writes about workflow automation, AI agents, and the tools we build. Default byline for industry overviews, listicles, and collaborative pieces.

Ready to put this into practice?

MoClaw runs browser tasks, research, and schedules automatically. Try it free.

grok 4.6 pricing grok 4.6 context window grok 4.6 api grok 4.6 xai grok 4.6 openrouter grok 4.6 release date

References: Grok 4.6 on OpenRouter · xAI model documentation · OpenRouter provider routing documentation · Kimi K2.6 on OpenRouter · Kimi K3 on OpenRouter · Qwen3.8 Max on OpenRouter