Opus 5 vs Kimi K3: A Tie, Then a Choice
Opus 5 vs Kimi K3: held-out testing puts them in a statistical tie, so the real choice is price and deployment shape. Full pricing, license and specs.
Table of Contents
Opus 5 vs Kimi K3 has a boring answer on capability and an interesting one on everything else. On a held-out interactive reasoning suite the two models finish in a statistical tie, so the real decision is about what you pay per token and what you are actually allowed to run.
The tie is measurable. On Witness, a private suite of ARC-AGI-3-style interactive puzzle games, Claude Opus 5 scored 43.4 plus or minus 3.2 and Kimi K3 scored 42.8 plus or minus 1.9 (Guanghan Ning). Those error bars overlap comfortably. As of July 29, 2026 that is the closest thing to a like-for-like capability read anyone has published on these two models.
Key Takeaways:
- On the held-out Witness suite, Opus 5 (43.4 ± 3.2) and Kimi K3 (42.8 ± 1.9) are statistically tied. Kimi K3 posted the tightest error bar of the three models tested.
- Per token, Opus 5 costs exactly 1.67x Kimi K3 on both sides of the meter: $5/$25 per million against $3/$15.
- No one has published a like-for-like per-task cost comparison. Any single "Opus 5 costs X times more per task" figure is invented.
- Kimi K3's weights are downloadable, but the checkpoint is 1.56 TB and Moonshot's own serving floor is 8x GB300. The open-weight right is real and, for most teams, theoretical.
- Opus 5 carries a documented runtime condition: cyber-flagged requests fall back to Opus 4.8, and the check reads everything in context, not just your prompt.
Opus 5 vs Kimi K3 on the Benchmarks Everyone Quotes
Public leaderboards hand you a different winner depending on which tab you open.
Anthropic's launch post reports Opus 5 at 30.2% on ARC-AGI-3 and describes that as "three times as high as the next-best model" (Anthropic). The Decoder put GPT-5.6 Sol at 7.8% on the same evaluation, which makes the raw ratio closer to four (The Decoder). Meanwhile on the independent Artificial Analysis Intelligence Index v4.1, Kimi K3 scored 57.1, sitting roughly level with Claude Opus 4.8 and behind Claude Fable 5 at 59.9 (Artificial Analysis).
Neither figure answers the question a buyer is asking. ARC-AGI-3 is public, so it can be trained toward. An aggregate index is an average, so a two-point gap can conceal a thirty-point swing on the one task you care about.
What the public boards prove: both models sit in the same tier, and each vendor can find a chart where it leads. What they leave unsolved: whether either score predicts behavior on a problem the model has not effectively already seen.
What the Held-Out Witness Suite Changed
Witness is the useful test precisely because it is not public. Guanghan Ning ran Opus 5, Kimi K3 and Fable 5 against a private set of ARC-AGI-3-style interactive games and published composites for all three: Opus 5 at 43.4 ± 3.2, Kimi K3 at 42.8 ± 1.9, Fable 5 at 43.8 ± 9.7 (Ning).
Read the middle column again. The open-weight model did not just keep pace with a closed frontier release, it did so with the tightest error bar of the three. Fable 5's ± 9.7 spans a range wider than the entire gap between first and last place. On a suite designed to reward novel problem solving, the three models are indistinguishable, and the most consistent of them is the one you can download.

We covered what this did to Opus 5's own headline number in our write-up of the Witness results. The point for this comparison is narrower and, if anything, more useful: once you hold the problems out, the capability argument between Opus 5 and Kimi K3 stops being an argument.
What Witness proved: on held-out interactive reasoning, there is no measurable capability gap between these two. What it left unsolved: it is one suite from one evaluator, and a tie on interactive puzzles is not a tie on long-horizon agent work, where nobody has run the equivalent test.
Opus 5 vs Kimi K3 Pricing and Specs
| Claude Opus 5 | Kimi K3 | |
|---|---|---|
| Released | July 24, 2026 | July 16, 2026 (weights July 27) |
| Weights | Closed | Open weight, Kimi K3 License |
| API price per 1M tokens | $5 in / $25 out | $3 in / $15 out |
| Context window | 1M tokens | 1M tokens |
| Witness held-out composite | 43.4 ± 3.2 | 42.8 ± 1.9 |
| Self-hosting | Not available | Permitted; 1.56 TB checkpoint |
| Consumer plan | Claude Pro and Max | Kimi app, ¥199/mo entry tier |
The per-token ratio is unusually clean. Opus 5 costs 1.67x Kimi K3 on input and exactly 1.67x on output, so there is no crossover point where one becomes cheaper for input-heavy or output-heavy work. Anthropic's fast mode runs roughly 2.5x faster at 2x the base rate, which pushes Opus 5 to about 3.3x K3's price when you buy latency (Anthropic). Kimi K3 quotes $0.30 per million for cached input, which matters if your agent replays a large stable prompt (Moonshot).
What the price sheet proves: K3 is meaningfully cheaper per token and the discount is uniform. What it leaves unsolved: per-token price is not your bill.
The Per-Task Number Nobody Has Published
This is where most comparisons quietly start guessing, so here is the honest state of it.
Moonshot puts Kimi K3 at roughly $0.94 per task in its own figures. Anthropic publishes no equivalent per-task number for Opus 5; the closest public data point is Artificial Analysis reporting Opus 5 at 20% lower cost per task than Fable 5 on its agentic knowledge work benchmark, which is a comparison to a different model on a different harness. Those two numbers were produced by different evaluators on different task sets. They cannot be divided into each other, and anyone handing you a single "Opus 5 costs X times more per task than Kimi K3" figure has manufactured it.
There is a structural reason the gap is not simply 1.67x. Opus 5 ships an effort dial and defaults to high, and Anthropic's own guidance is to use low and medium far more aggressively than most teams do. Effort changes how many tokens the model spends thinking, which means the same prompt can produce very different bills depending on a parameter that has no counterpart in the K3 API. A 67% price premium at high effort and a 67% premium at low effort are not the same expense.
What this section proves: the per-token gap is knowable and the per-task gap is not. What it leaves unsolved: until someone runs both models over an identical agent workload with effort held constant, the cost comparison stays an estimate you have to make on your own traffic.
Why Open Weights Does Not Mean You Will Run Kimi K3
The Kimi K3 License permits self-hosting cleanly. Download the checkpoint, run it on hardware you control, fine-tune it, keep every token inside your perimeter. Internal use carries no conditions at all (Kimi K3 License).
Then you look at the hardware. The published checkpoint runs 1.56 TB across 96 shards, and that is already the compressed release, trained quantization-aware so there is no smaller official version waiting. Eight H100s, the configuration a lot of launch-week blog posts suggested, give you 640 GB. The weights alone need more than twice that before a single token of KV cache for a million-token context. Moonshot's own vLLM recipe puts the floor at 8x GB300 (vLLM).

The second surprise arrived with the weights. Open checkpoints normally push hosted prices down, because providers compete on margin over the same free artifact. That did not happen. Together AI, Fireworks, Baseten, Modal, Nebius and DigitalOcean all went live the same day at exactly $3.00 input and $15.00 output, identical to Moonshot's own API, with no cheaper tier anywhere. Developers on r/LocalLLaMA connected it to the license's Model-as-a-Service clause, though Moonshot has confirmed nothing and the license text sets no prices. We walked through the actual terms in our breakdown of the Kimi K3 License.

What the license proves: the right to self-host is genuine and unconditioned for internal use. What it leaves unsolved: for any team without a GPU cluster already on the balance sheet, that right is theoretical, and the open weights bought you no discount at all.
The Opus 5 Deployment Condition You Do Not Control
Opus 5 has its own asterisk, and it is not a capability one.
Anthropic runs classifiers over Opus 5 requests, and when one flags a high-risk offensive cybersecurity request the system falls back to Opus 4.8 by default. You still get an answer; you get it from a different model than the one you selected. The support documentation notes that the checks review everything the model reads, which includes memory, connector output, web results and files, not only the text you typed (Anthropic Support). Anthropic reports these classifiers intervene roughly 85% less often than the equivalents governing Fable 5.
Source-code vulnerability discovery, security triage and writing secure code all remain on Opus 5. Binary vulnerability scanning, penetration testing and exploit generation are the blocked categories. For a security team, that boundary is worth mapping before it surprises a pipeline. For everyone else it is a reminder that the closed model has runtime behavior you cannot inspect, in the same way the open model has hardware requirements you cannot dodge.
What this proves: both models constrain you, just in different currencies. What it leaves unsolved: there is no published figure for how often the fallback fires on ordinary engineering work, so the practical impact is anecdote for now.
FAQ
Is Kimi K3 better than Claude Opus 5?
Not measurably, on the one held-out test that covers both. Witness put Opus 5 at 43.4 ± 3.2 and Kimi K3 at 42.8 ± 1.9, which is a statistical tie. On public leaderboards each model leads somewhere, which is why the held-out result carries more weight than either.
Is Kimi K3 cheaper than Opus 5?
Per token, yes, by a uniform 1.67x: $3/$15 per million against $5/$25. Per task the honest answer is that no one has published a like-for-like comparison, and Opus 5's effort parameter makes the same prompt cost meaningfully different amounts.
Can I self-host Kimi K3 instead of paying for Opus 5?
The license allows it with no conditions for internal use. The checkpoint is 1.56 TB and Moonshot's serving floor is 8x GB300, so this is a real option for teams that already own a cluster and a non-option for everyone else.
Did Kimi K3's open weights make it cheaper than Opus 5?
No more than it already was. Every day-zero provider prices K3 at exactly $3.00 and $15.00 per million, matching Moonshot's own API, with no discounting since the July 27, 2026 weight release.
Which one should I use for coding agents?
Test both on your own tasks. The public coding benchmarks disagree with each other, the held-out reasoning test says they are level, and your repository is a better evaluator than either. Start with K3 if per-token cost dominates your bill, and with Opus 5 at low or medium effort if you want the effort dial as a cost lever.
Choosing Between Opus 5 and Kimi K3 by Deployment Shape
Once the capability question resolves to a tie, what is left is a choice between two shapes of the same constraint.
Claude Opus 5 is rented and closed. You pay 1.67x per token, you get an effort dial that is a genuine cost lever, and you accept a runtime condition where the model answering may not be the model you picked. Kimi K3 is open on paper and rented in practice. You pay less per token, the weights are yours to download and fine-tune, and the 1.56 TB checkpoint means almost nobody will exercise that right.
Neither of these is the "closed frontier versus open alternative" story the launch coverage told. Both are models you reach through someone else's API, differing in price, in the shape of their restrictions, and not, on the evidence available as of July 29, 2026, in what they can actually do.
That makes the practical move an empirical one rather than an editorial one. Point both at a task from your own backlog, hold the harness constant, and read your own bill. If you want to run them side by side without provisioning anything, MoClaw puts each on its own path: Kimi K3 as a managed agent on the Kimi tier, and Claude Opus 5 on Ultra (Ultra defaults to Opus 5). For the Claude side, our guide to getting access to Opus 5 covers the API, Claude Code, the plan tiers, and the Ultra shortcut inside MoClaw.
Continue Reading
More ComparisonThe MoClaw editorial team writes about workflow automation, AI agents, and the tools we build. Default byline for industry overviews, listicles, and collaborative pieces.
Use Kimi K3 on MoClaw, without the setup
Run Kimi K3 as an always-on managed agent with memory, your tools, and scheduling. No API wiring, no plan gating, no self-hosting.
References: Anthropic: Introducing Claude Opus 5 · Guanghan Ning: Witness held-out results for Opus 5, Kimi K3 and Fable 5 · Kimi K3 License (Hugging Face) · vLLM recipe for Kimi K3 (serving floor) · Anthropic Support: Why Claude switched models in your conversation with Opus 5 · Kimi K3 on Artificial Analysis (Intelligence Index v4.1) · The Decoder: Opus 5 near Fable 5 at half the token price · Moonshot AI platform docs (Kimi K3 API pricing)