Opus 5 vs Fable 5: Two Different Rankings

Comparison · 8 min read · Published: · Updated:

Opus 5 vs Fable 5: Fable leads the aggregate index, Opus 5 leads agentic work at half the price. Real benchmark numbers and which to pick for your workload.

MoClaw Editorial · MoClaw editorial team
Opus 5 vs Fable 5: Two Different Rankings
Table of Contents

Share this

Opus 5 vs Fable 5 produces two different winners, and both results are correct. On aggregate capability indices Claude Fable 5 still edges ahead; on agentic work measured per dollar Claude Opus 5 leads by a wide margin, at half the price. Anthropic has not claimed Opus 5 is its most intelligent model, and that is the tell.

The gap between those two readings is bigger than the gap between the models. Epoch AI's Capability Index put Opus 5 at 159 against Fable 5's 161, while Artificial Analysis ranked Opus 5 first on agentic knowledge work, beating Fable 5 "by nearly 150 Elo while cutting cost per task by 20%" (Yellow, July 25, 2026). A two-point deficit and a 150-Elo lead describe the same pair of models.

Key Takeaways:

  • Opus 5 costs $5 per million input tokens and $25 output, half of Fable 5's rate (Decrypt).
  • Fable 5 leads on the aggregate index (Epoch ECI 161 vs 159) and the two tied on software engineering.
  • Opus 5 leads on agentic evaluations: Frontier-Bench v0.1 at 43.3% vs 33.7%, GDPval-AA v2 at 1,861 vs 1,747.
  • On a held-out interactive reasoning suite they are statistically tied, but Fable 5's error bar is three times wider.
  • Anthropic states Opus 5 trails its Mythos 5 model on offensive cybersecurity work, so neither of these is the answer there.

Opus 5 vs Fable 5 at Half the Price

Start with the number that frames everything else. Opus 5 is priced at $5 per million input tokens and $25 per million output, which Decrypt reports is half what Fable 5 costs, putting Fable 5 at $10 and $50 (Decrypt, July 24, 2026).

Anthropic did not pair that price cut with a claim of superiority. The launch positioning describes Opus 5 as the new default on Claude Max and the strongest model available on Claude Pro, not as the most intelligent model the company ships (Anthropic). Fable 5 remains the frontier product. Opus 5 is the one you are now routed to by default.

A vendor releasing a cheaper model and declining to call it the best is unusual enough to take at face value. It is also the single most useful fact in this comparison, because it tells you the two models were built to win different arguments.

What the pricing proves: the 2x gap is real and Anthropic is not positioning it as a discount on the same thing. What it leaves unsolved: half the price per token is not half the price per task, and the two diverge in both directions.


Where Fable 5 Still Wins: The Aggregate Index

Epoch AI's Capability Index scored Fable 5 at 161 and Opus 5 at 159, and the two tied on the software engineering component (Yellow).

Two points is not much, and the community reaction was that it was too little. The launch discussion on Latent Space captured a running argument that Opus 5 was underrated by aggregate scoring and that the field needs harder public benchmarks to separate models at this level (Latent Space). That reaction is itself the interesting datum: a two-point spread that practitioners felt was wrong by a wide margin.

An aggregate index is an average across task families. When two models are close on the average and far apart on a specific family, the average is the least informative number available. It is the right tool for tracking a field over years and the wrong tool for choosing a model on a Tuesday.

What Epoch proves: on breadth of capability, Fable 5 remains marginally ahead. What it leaves unsolved: nothing in a two-point aggregate spread tells you which model does your work better.


Where Opus 5 Wins: Agentic Work Per Dollar

Change the measurement to completed agentic work and the ordering inverts, decisively rather than marginally.

On Frontier-Bench v0.1, Opus 5 scored 43.3% against Fable 5's 33.7%. On GDPval-AA v2 it reached 1,861 against Fable 5's 1,747. On AutomationBench, Opus 5 completed workflows at 100% where previous models failed outright (Decrypt). Artificial Analysis put it first on agentic knowledge work, ahead of Fable 5 by close to 150 Elo while costing 20% less per task (Yellow).

That last combination is the one to sit with. A lead of 150 Elo alongside a 20% lower cost per task means Opus 5 is not trading quality for price on this class of work. It is ahead on both axes simultaneously, against a model that beats it on the aggregate index.

Opus 5 versus Fable 5 across three benchmarks: Fable 5 leads the Epoch capability index 161 to 159, while Opus 5 leads Frontier-Bench v0.1 and GDPval-AA v2
Opus 5 versus Fable 5 across three benchmarks: Fable 5 leads the Epoch capability index 161 to 159, while Opus 5 leads Frontier-Bench v0.1 and GDPval-AA v2

What agentic benchmarks prove: for multi-step work with tools, Opus 5 is ahead by margins that dwarf the index gap. What they leave unsolved: agentic benchmarks are young, harness-dependent, and mostly built by parties with an interest in the result.


The Held-Out Test Where They Tie

There is one evaluation covering both models that neither vendor could have trained toward, and it produces the flattest result of all.

On Witness, a private suite of ARC-AGI-3-style interactive puzzle games, Opus 5 scored 43.4 plus or minus 3.2 and Fable 5 scored 43.8 plus or minus 9.7 (Guanghan Ning). The central values are within half a point of each other. The error bars are not remotely comparable.

Fable 5's spread of plus or minus 9.7 is roughly three times Opus 5's. On a suite built to reward novel problem solving, the more expensive model is not more capable, it is less predictable. For a one-shot query that variance is invisible. For an agent taking forty steps where each one compounds, it is the whole ballgame. We broke down what this suite did to Opus 5's own headline claims in our analysis of the Witness results.

Held-out Witness composites: Claude Opus 5 at 43.4 plus or minus 3.2 against Claude Fable 5 at 43.8 plus or minus 9.7, an error bar roughly three times wider
Held-out Witness composites: Claude Opus 5 at 43.4 plus or minus 3.2 against Claude Fable 5 at 43.8 plus or minus 9.7, an error bar roughly three times wider

What Witness proves: on held-out interactive reasoning the two are indistinguishable on average, and Opus 5 is markedly more consistent. What it leaves unsolved: one private suite from one evaluator is a data point, not a verdict, and no one has run the equivalent test on long-horizon agent work.


What the People Running Both in Production Say

Vendor-adjacent testimony is worth reading precisely because it is attributable, so here it is with names attached.

Scott Wu, CEO of Cognition, said Opus 5 "approaches Fable-level performance at half the cost inside Devin." Wade Foster of Zapier said the model topped his company's automation leaderboard without spending more tokens than its predecessors (Yellow). VentureBeat's launch coverage carried the same two accounts alongside Anthropic's own framing of Opus 5 as a model for coding agents and enterprise workflows (VentureBeat).

The dissent is more useful than the praise. Claire Vo described Opus 5 as "brilliant but annoying," citing a neurotic streak (Yellow). That is a qualitative failure mode no benchmark in this article measures: a model that arrives at good answers while being tiresome to work alongside. If you have found Opus 5 verbose or over-eager to re-verify its own work, you are describing something a named practitioner reported independently, not a problem with your prompt.

What the testimony proves: teams running both in production report Fable-adjacent quality from Opus 5 at materially lower cost. What it leaves unsolved: every one of these accounts comes from a party with a commercial relationship to the outcome, and none is a controlled comparison.


Where Neither Is the Answer: Offensive Cybersecurity

Anthropic states that Opus 5 trails its Mythos 5 model on offensive cybersecurity work (Yellow). That is a rare piece of vendor self-limitation and it should be read literally rather than as modesty.

The practical consequence is not only which model scores higher. Opus 5 runs classifiers that route flagged offensive-security requests to Opus 4.8, which means a security workflow on Opus 5 can be answered by a different model than the one selected. We covered the mechanism, the boundaries and the setting that governs it in why Claude switches models mid-conversation.

What this proves: for offensive security specifically, the Opus 5 vs Fable 5 comparison is the wrong question. What it leaves unsolved: Mythos 5's availability is far narrower than either model discussed here, so knowing it leads does not mean you can use it.


FAQ

Is Opus 5 better than Fable 5?

It depends entirely on the measurement. Fable 5 leads the Epoch aggregate index 161 to 159. Opus 5 leads agentic evaluations by wide margins, including Frontier-Bench v0.1 at 43.3% versus 33.7%, and does it at half the token price.

How much cheaper is Opus 5 than Fable 5?

Half, per token. Opus 5 is $5 per million input and $25 output; Fable 5 is $10 and $50. Artificial Analysis separately measured Opus 5 at 20% lower cost per completed task on agentic knowledge work.

Does Anthropic say Opus 5 is its most intelligent model?

No. Anthropic positions Opus 5 as the default on Claude Max and the strongest model on Claude Pro, without claiming the top intelligence slot. Fable 5 remains the frontier product.

Which is better for long-running agents?

The evidence points to Opus 5, on two grounds: it leads the agentic benchmarks, and on held-out interactive reasoning its results are roughly three times more consistent than Fable 5's. Consistency compounds across steps in a way that average scores do not capture.

Why do the benchmarks disagree with each other?

Because they measure different things. An aggregate capability index averages across task families; an agentic benchmark measures completed multi-step work with tools and a cost budget. Two models can legitimately swap places between the two.

Should I switch from Fable 5 to Opus 5?

Run your own workload before deciding. The case for switching is strongest for tool-using agents and cost-sensitive volume; the case for staying is strongest for one-shot work at the edge of what any model can do, where the aggregate index still favors Fable 5.


Choosing Between Opus 5 and Fable 5 by Workload, Not Index Score

The honest summary is that "aggregate intelligence" and "work completed per dollar" have become two separate rankings, and this pair of models is the clearest demonstration so far.

If your work is one-shot and hard, the two-point index gap is at least pointing in a direction, and Fable 5 remains the more ambitious choice. If your work is agentic, repeated, or budgeted, Opus 5 wins on the benchmarks built to measure that, wins on cost per task, and is the steadier of the two when the problems are genuinely novel. Anthropic charging half as much for it while declining to call it the smartest model is not a contradiction. It is the company describing exactly this split.

The same pattern is showing up across the field, not just inside Anthropic's lineup: we found capability parity and a deployment-shaped difference comparing Opus 5 against Kimi K3 as well. If you want to try Opus 5 on a real workload without provisioning anything, MoClaw runs it as a managed agent.

Continue Reading

M
MoClaw Editorial MoClaw editorial team

The MoClaw editorial team writes about workflow automation, AI agents, and the tools we build. Default byline for industry overviews, listicles, and collaborative pieces.

Choosing between tools? Let MoClaw run the work.

Always-on AI assistant on its own cloud computer. No switching required, no setup.

claude opus 5 vs fable 5 opus 5 vs fable 5 pricing is opus 5 better than fable 5 claude opus 5 epoch eci opus 5 agentic benchmark claude fable 5 cost per task anthropic most intelligent model

References: Anthropic: Introducing Claude Opus 5 · Yellow: Experts split over Claude Opus 5 after the first independent tests · Decrypt: Opus 5 outscores Fable 5 on most benchmarks at half price · Guanghan Ning: Witness held-out results for Opus 5, Kimi K3 and Fable 5 · VentureBeat: Anthropic launches Claude Opus 5 · Latent Space: Claude Opus 5 launch reactions