Opus 5's ARC-AGI-3 Jump Didn't Transfer
Opus 5's ~4x ARC-AGI-3 lead collapses to a statistical tie with Kimi K3 and Fable 5 on a held-out suite. What the Witness benchmark says about picking a model.
Table of Contents
Claude Opus 5's headline ARC-AGI-3 result, reported at 30.2% and framed as roughly four times the next-best model, does not survive a held-out test. On Witness, a private suite of ARC-AGI-3-style interactive puzzle games, Opus 5 scores 43.4 ± 3.2, a statistical tie with the open-weight Kimi K3 at 42.8 ± 1.9 and with Anthropic's own Fable 5 at 43.8 ± 9.7. The four-times gap flattens to no gap at all once the problems can't have been seen during training.
Key Takeaways:
- Anthropic reports Opus 5 at 30.2% on ARC-AGI-3 against GPT-5.6 Sol's 7.8%. The company's own wording is "three times as high as the next-best model."
- On the held-out Witness suite, Opus 5 (43.4 ± 3.2) ties Kimi K3 (42.8 ± 1.9) and Fable 5 (43.8 ± 9.7). An open model matches a closed flagship.
- Opus 5 still clears Opus 4.8 (34.8) on Witness. The gain is real, just small, and nowhere near what the public number implies.
- On the most novel puzzle, where rules have to be found by poking at the game, Opus 5 drops below Opus 4.8.
- Excellent on familiar templates, worse on genuinely new mechanics: that shape is what a leaked benchmark looks like. It's an industry problem, not one vendor's.
This is not an accusation that Anthropic cooked its numbers. It's a case study in why a single public benchmark, even a good one, can't decide your model.
What Anthropic Reported on ARC-AGI-3
ARC-AGI-3 exists to measure novel problem solving, the kind of reasoning you can't memorize your way through. In its launch post, Anthropic says Opus 5's score is "three times as high as the next-best model." Reporting from The Decoder puts the raw figure at 30.2% versus 7.8% for GPT-5.6 Sol, which works out closer to a fourfold gap, and it lands at half the token price of Fable 5, $5 and $25 per million tokens against Fable's $10 and $50.
Read cold, that's a big claim: near-frontier reasoning, a fraction of the price, and a benchmark built specifically to resist memorization. The number did most of the work in the launch-week coverage, and Decrypt's write-up leaned on it too. Whether it means what it looks like it means is a separate question, and it took someone with a private test set to ask it.
What's solid here: the ARC-AGI-3 score is real and Anthropic published it plainly. What's still open after the launch post: whether a high score on a public evaluation predicts anything about problems the model hasn't effectively already met.
What the Held-Out Witness Suite Found
Guanghan Ning, who runs the Witness evaluation, took Opus 5 and ran it on games built in the ARC-AGI-3 style but kept off the public internet. Same harness for every model, same compute budget, no home-field advantage. His summary of the result: "The leap doesn't transfer."
The composites tell the story without much need for interpretation. Opus 5 comes in at 43.4 ± 3.2. Kimi K3, which anyone can download and run, sits at 42.8 ± 1.9. Fable 5, Anthropic's most capable model, lands at 43.8 ± 9.7. Those error bars overlap so heavily that calling any one of them the winner is statistical noise. Opus 5 does beat Opus 4.8's 34.8, so there's a genuine improvement generation over generation; it's roughly nine points, not the twentyfold jump the ARC-AGI-3 framing suggested.

Sit with the Kimi K3 comparison for a second, because it's the uncomfortable part. On the public benchmark, Opus 5 looks like it's playing a different sport. On the held-out one, an open-weight model you can run on your own hardware is indistinguishable from it. If you're weighing where Opus 5 actually stands against K3, that tie is the whole conversation, and it's the one we dig into in Claude Opus 5 vs Kimi K3.
The Decomposition That Matters
A tie on the aggregate would be interesting on its own. The traces are what make it damning. Ning reports that on the familiar puzzle types, the ones structurally close to what's floating around publicly, Opus 5 produced the same optimal solution run after run, across repeated seeds, with almost no exploratory moves. It didn't work the problem; it recognized it and went straight to the answer.
Then the mechanics change. On the most novel game, the one with unusual rule combinations that a model has to discover by interacting with the environment, Opus 5 regresses below Opus 4.8. The newer, pricier, higher-scoring model does worse than its predecessor precisely where memorization stops helping.
Ning's own read, and I want to be careful to attribute this to him rather than smuggle it in as fact, is that the behavior looks like pattern-matching to seen structures rather than reasoning from scratch. He hasn't alleged that Anthropic did anything improper. He's describing what the traces show and letting the shape speak. (One detail I couldn't confirm against the original thread, which sits behind a paywall: some coverage describes the repeated-seed solutions as byte-for-byte identical across five of five runs. Treat that specific precision as reported, not verified, until you can read the source directly.)
Why Strong-on-Templates, Weak-on-Novelty Reads Like a Training Signature
Here's the mechanism, stated plainly. When a benchmark is public, its problem types leak into the training corpus, sometimes as the actual items, more often as synthetic data shaped just like them. A model trained on that distribution gets very good at the template and no better at the underlying skill. You see it as a top score on the named benchmark and a flat or worse score the moment the mechanics are genuinely new. That gap between the two is the fingerprint.
Opus 5 on ARC-AGI-3 versus Witness fits the fingerprint about as cleanly as anyone has managed to show. It isn't proof of contamination; you can't prove it from the outside without the training data. But it's the first case I've seen where the public number and the held-out number diverge this sharply, with trace-level evidence for why. The community reaction ran the same direction: one researcher replying to Ning put it as "a real problem with all these benchmarks," and the fuller Latent Space roundup of launch-week reactions circles the same worry from several angles.
The restrained version of the takeaway, which is the correct one: this is a measurement failure, not a fraud. ARC-AGI-3 was supposed to be memorization-resistant and turned out not to be resistant enough. Every lab shipping against public leaderboards inherits the same problem.
How to Pick a Model When One Benchmark Can Lie
Stop shortlisting models off a launch-day leaderboard. That's the practical lesson, and it costs you almost nothing to act on. Build a tiny held-out evaluation from your own work: twenty to fifty real tasks the model has provably never seen, scored the way you actually care about, run against every candidate under the same budget. It's an afternoon of setup and it catches exactly the failure Witness caught.
A quick sketch of how this bites in practice. Suppose you're choosing a reasoning model for an agent that has to handle inputs you can't fully anticipate. You pick on the ARC-AGI-3 headline, ship Opus 5, and it sails through everything that resembles a known pattern, then quietly underperforms on the genuinely weird inputs that were the reason you wanted strong reasoning in the first place. Your own eval would have shown the flat spot before you committed. The leaderboard hid it.
If you do want Opus 5 in the mix, the sensible setup is to run it next to a cheaper or open alternative and let your eval referee, rather than trusting either one's marketing. You can run Opus 5 alongside Kimi K3 through a single managed layer and compare them on your own tasks instead of someone else's puzzles. And if you're still at the how-do-I-even-turn-this-on stage, start with how to access Claude Opus 5.
FAQ
What did Opus 5 score on ARC-AGI-3?
30.2%, per The Decoder, against 7.8% for GPT-5.6 Sol. Anthropic describes it as three times the next-best model; the raw ratio is closer to four times. Both framings come from the same public result.
What is the Witness benchmark?
A held-out suite of ARC-AGI-3-style interactive puzzle games, kept off the public internet by researcher Guanghan Ning, used to test whether a public benchmark score reflects general skill or memorized templates.
Did Opus 5 beat Kimi K3 on Witness?
No. Opus 5 scored 43.4 ± 3.2 and Kimi K3 scored 42.8 ± 1.9, a statistical tie. Fable 5 was also tied at 43.8 ± 9.7.
Does this mean Anthropic faked its benchmark?
No. The ARC-AGI-3 score is real and was reported honestly. The problem is that the public benchmark appears not to predict held-out performance, which is a measurement issue that affects the whole industry.
Is Opus 5 still an upgrade over Opus 4.8?
On Witness, yes, by roughly nine points (43.4 vs 34.8), though it regresses below 4.8 on the most novel puzzle. The improvement is real but far smaller than the ARC-AGI-3 number implies.
One Benchmark Can't Tell You Which Model to Ship
The Opus 5 ARC-AGI-3 story is going to get retold as "Anthropic overstated it," and that retelling misses the point. The number is accurate. The benchmark just doesn't measure what its 30.2% seemed to promise, and it took a private test set to expose the gap. That's the durable lesson, worth more than any single model's ranking: the only benchmark that can tell you which model to ship is one built from work the model has never seen, which usually means one you build yourself. Everything on the public leaderboard is a starting hypothesis, not a decision.
Continue Reading
More ResearchThe MoClaw editorial team writes about workflow automation, AI agents, and the tools we build. Default byline for industry overviews, listicles, and collaborative pieces.
Turn insights into action.
MoClaw automates the recurring work your analysis points to. No engineering required.
References: Anthropic: Introducing Claude Opus 5 · The Decoder: Opus 5 near Fable 5 at half the token price · Guanghan Ning: Witness held-out results for Opus 5 · Latent Space: Claude Opus 5 launch reactions · Decrypt: Opus 5 outscores Fable 5 on most benchmarks at half price