Back to Insights
AI Software

Claude Opus 5: Anthropic Undercuts Its Own Flagship

Claude Opus 5 ships at $5/$25, half of Fable 5, with near-Fable scores, looser cyber classifiers and Anthropic's lowest misalignment audit score yet.

S5 Labs Team July 24, 2026

Anthropic released Claude Opus 5 today at $5 per million input tokens and $25 per million output — identical to Opus 4.8, and exactly half what Claude Fable 5 costs. The company’s framing is blunt: “the frontier intelligence of Claude Fable 5 at half the price.” It went live the same day on the API, Claude.ai, Claude Code, and Claude Cowork, and it is now the default on Max and the top model on Pro.

Shipping a better model was the expected part of today. The price is what’s worth stopping on. Six weeks ago the Fable tier existed to justify a $10/$50 premium; most of what that premium bought now sits in the cheaper product, and Anthropic put that in its own opening line rather than leaving anyone to work it out.

Two-panel chart: Claude Opus 5 ships at 5 input and 25 output per million tokens, half of Fable 5's 10/50, with fast mode at 2× base price landing exactly on Fable 5's rate; and the safety dials moving in opposite directions — misalignment audit score dropping to 2.3, the lowest of recent models, while cyber classifiers intervene roughly 85% less often than Fable 5's.

The Pricing Move Is the Announcement

Anthropic has now held the Opus line at $5/$25 through 4.6, 4.7, 4.8 and 5. Holding a price while capability climbs is the standard frontier-lab move, and on its own it wouldn’t be news. What makes this release different is that the tier above it barely has room left.

Fable 5 launched in June as the Mythos-class model made safe for general use, at double Opus pricing. On CursorBench 3.2, Anthropic says Opus 5 at maximum effort lands within 0.5% of Fable 5’s peak score. On OSWorld 2.0 it beats Fable 5’s best result outright, at just over a third of the cost. If those hold up, the remaining case for paying Fable rates is narrow: work where the last half-percent is worth 2× the bill, or the specific dual-use capabilities Fable retains.

The fine print makes the collapse literal. Opus 5 ships a fast mode — roughly 2.5× default speed at twice the base price. Twice $5/$25 is $10/$50. Anthropic’s new fast tier is priced to the dollar at what Fable 5 costs, and buys speed instead of a capability step. That is a company repricing its own lineup around throughput rather than intelligence, which is a reasonable read of where the constraint has moved for anyone running agents in production.

ModelInput / MtokOutput / Mtok
Claude Opus 5$5$25
Claude Opus 5, fast mode$10$50
Claude Opus 4.8$5$25
Claude Fable 5$10$50

Satya Nadella’s argument that enterprises are paying for the same intelligence twice gets sharper here. When the premium tier and the fast tier of the cheaper model cost the same, the thing being sold stops being intelligence and starts being latency.

What the Benchmarks Do and Don’t Say

Anthropic claims Opus 5 tops Frontier-Bench and GDPval-AA, more than doubles Opus 4.8 on Frontier-Bench v0.1 at lower cost, scores three times the next-best model on ARC-AGI 3, and passes Zapier’s AutomationBench at about 1.5× the next-best rate for the same cost per task. On life sciences, it reports 10.2 points over Opus 4.8 on organic chemistry and 7.7 points on protein functionality prediction.

Almost every one of those is a relative claim — “three times the next-best,” “within 0.5%,” “1.5×” — rather than an absolute score you could set beside another lab’s published figure. This is the same reporting pattern we flagged at the Sonnet 5 launch, and it has now hardened into house style. Relative claims are not false, but they are unfalsifiable without the baseline, and the baseline is chosen by the company making the claim.

The benchmark set itself has rotated. SWE-bench, the number that anchored every Anthropic coding announcement through the 4.x line, does not appear. In its place: Frontier-Bench v0.1, CursorBench 3.2, Zapier AutomationBench, OSWorld 2.0, GDPval-AA. Some of that rotation is legitimate — SWE-bench saturated, and agentic work needs harder tests. But “v0.1” is a benchmark with no track record, no independent replication, and no adversarial history. A model topping a benchmark released this quarter tells you less than a model topping one that has survived two years of people trying to game it. Given that OpenAI models were caught training against a leaked eval set two days ago, the provenance of the scoreboard is not a pedantic concern.

The capability anecdote Anthropic leads with is more persuasive than the numbers. Given a drawing of a machine part and no ability to view images, Opus 5 wrote its own computer vision pipeline to reconstruct the 3D model — a task the company says competing models failed after five attempts. That is the behavior the announcement is actually selling — a model that verifies its own work and keeps iterating rather than confidently returning something wrong. For agent workloads, that difference compounds across a long run in a way a benchmark delta doesn’t capture.

The Safety Dials Moved in Opposite Directions

Anthropic reports that Opus 5 scores 2.3 on overall misaligned behavior in its automated behavioral audit, the lowest of any recent model, and that it adheres to Claude’s Constitution better than Opus 4.8, Sonnet 5, or Fable 5, with the lowest observed rates of deceptive behavior.

In the same document, the company says Opus 5’s cyber classifiers are proportionally less restrictive than Fable 5’s and should intervene roughly 85% less often.

Both of those can be true at once, and the reasoning is defensible: Fable 5’s classifiers were bolted on after a jailbreak produced working exploit-demonstration code, and they were tuned conservatively under time pressure. Opus 5 is materially weaker at exploit development — Anthropic concedes it trails Mythos 5 badly there even though it finds vulnerabilities at a similar rate — so a lighter filter is proportionate to a lower ceiling. Blocked requests on Claude.ai and Claude Code still fall back to Opus 4.8, the same routing pattern Fable 5 uses, and a Cyber Verification Program opens restricted security work to vetted enterprises and researchers.

Still, an 85% reduction in intervention frequency is a large step to take on the strength of an internal audit, announced the same week as the model it governs. The classifier loosening will be tested by adversaries long before it’s evaluated by anyone independent. What makes it easier to accept than it would have been three months ago is that the fallback architecture already exists and has been running in production since July 1. The safeguard here is a dial being turned, not a new system being trusted for the first time.

What Changes If You’re Building On This

For anyone already running Opus 4.8 in production, this is close to a free upgrade: same price, same model family, better verification behavior, and a lighter cyber filter that should stop tripping on legitimate security work. Test it, but the migration risk is low.

Fable 5 customers have the harder question, which is whether they can name the specific capability the extra $5 and $25 are buying. “Frontier tier” no longer answers it, because Opus 5 is the frontier tier for most work now. Biology depth and exploit-development capability are real reasons to stay. Habit is not.

For anyone comparing across labs, the advice is the boring one: hold the relative benchmark claims loosely and run your own evals on your own tasks. It is also the only thing that survives a week where a flagship gets cheaper by half, the scoreboard gets replaced, and the safety posture loosens, all in one announcement and all measured by the company shipping it. Two things will settle whether today’s numbers meant anything — whether the loosened classifiers hold under adversarial pressure, and whether Fable 5 still has a business at twice the price.

Sources

Want to discuss this topic?

We'd love to hear about your specific challenges and how we might help.