Anthropic shipped Claude Opus 5 on July 24, 2026, roughly two months after Opus 4.8, which is the same upgrade cadence the company has kept for the entire Opus line. What’s unusual isn’t the timing. It’s the pricing decision: Opus 5 costs exactly what Opus 4.8 cost, $5 per million input tokens and $25 per million output tokens, and Anthropic is framing the entire release around getting meaningfully more model for that same dollar rather than charging more for more capability. That’s a different kind of upgrade than most model releases, and it changes how you should think about whether to switch.

Now, if you’re willing to overlook many of the Reddit posts regarding the Opus 5 and genuinely curious about learning how much it improves upon the Opus 4.8, you’re at the right place.
Opus 5 vs Opus 4.8 – What has changed?

Anthropic’s own positioning for Opus 5 is that it comes close to the frontier intelligence of its top-tier model, Claude Fable 5, at half the price of that model, while staying at Opus 4.8’s existing price point. It’s now the default model on Claude Max and the strongest model available on Claude Pro. On two of Anthropic’s headline evaluations, Frontier-Bench and GDPval-AA, Opus 5 is the new state of the art among Anthropic’s models. However, it remains behind Mythos 5 specifically on cybersecurity tasks, which is a deliberate safety choice rather than a capability gap (more on that below).
The short version: this isn’t a “smarter but pricier” release. It’s a “same price, meaningfully more capable, and far more efficient” release, aimed squarely at the workloads teams run every single day rather than at chasing a benchmark crown.
Spec sheet, side by side
| Claude Opus 4.8 | Claude Opus 5 | |
| Released | May 28, 2026 | July 24, 2026 |
| Input price | $5.00 / million tokens | $5.00 / million tokens |
| Output price | $25.00 / million tokens | $25.00 / million tokens |
| Context window | 1,000,000 tokens | 1,000,000 tokens |
| Fast Mode | $10 / $50 at 2.5x speed | $10 / $50 at 2.5x speed |
| Effort control | High / extra / max | Low through max, with finer-grained token-vs-intelligence tradeoffs |
| Default tier placement | Was default on Claude Max | New default on Claude Max, strongest on Claude Pro |
| Cyber safeguard strictness | Baseline for its generation | Classifiers intervene roughly 85% less often than on Fable 5, looser than 4.8’s baseline |
| Alignment audit score | Higher misaligned-behavior score than Opus 5 | Lowest misaligned-behavior score (2.3) of any recent Anthropic model |
Nothing moved on price, context window, or Fast Mode’s speed multiplier. Everything moved on how much useful work the model does inside those same constraints, which is a more interesting kind of upgrade to evaluate, because it means the honest question isn’t “is Opus 5 better,” it’s “how much better, on what, and does that translate to fewer tokens for the same job.”

Here’s how I’d explain the difference to someone who doesn’t want to read a benchmark table. Opus 4.8 is a car with three gears: high, extra, and max effort. You pick a gear, and it’s a reasonable engine either way. Opus 5 is the same car with a proper gearbox, a wider spread of effort settings from low up through max, tuned so that even the lowest gear now clears obstacles the old car needed second or third gear to get over.
That reframes the whole “which effort level should I use” question. With Opus 4.8, going cheap meant giving something up. With Opus 5, several of Anthropic’s own benchmark charts show the low-effort setting beating other models’ best results on the same task, meaning the cheap gear isn’t a compromise anymore so much as a different point on a curve that’s shifted upward across its entire length.
Opus 5 vs Opus 4.8 – Benchmark comparison
| Benchmark | What it tests | The Opus 5 vs Opus 4.8 result |
| Frontier-Bench v0.1 | Real-world software engineering tasks on an agent harness | Opus 5 more than doubles Opus 4.8’s performance, at a lower cost per task |
| CursorBench 3.2 | In-editor coding agent performance | At max effort, Opus 5 lands within 0.5% of Fable 5’s peak score, at half Fable 5’s cost per task |
| ARC-AGI 3 | Novel problem-solving the model hasn’t seen trained patterns for | Opus 5 scores roughly three times higher than the next-best model |
| Zapier AutomationBench | End-to-end business task completion | Roughly 1.5x the pass rate of the next-best model at the same cost; even Opus 5’s lowest effort setting beats every other model |
| OSWorld 2.0 | Computer-use benchmark: operating a real desktop environment | Opus 5 outperforms every model at any cost, and beats Fable 5’s best result at about a third of Fable 5’s cost |
| Life sciences (internal suite) | Structural biology, organic chemistry, bioinformatics | Better than Opus 4.8 on every single evaluation in the suite |
Translating a few of these into what they mean day to day:
Frontier-Bench more than doubling is the number that matters if your team runs Opus on real engineering tickets rather than toy coding problems. This benchmark specifically measures messy, real-world software tasks on an agent harness, not clean leetcode-style problems, so a 2x-plus jump there is closer to “the agent finishes twice as many real tasks correctly” than “the model got a better score on a quiz.”
CursorBench within 0.5% of Fable 5 at half the cost is the number for anyone who’s been paying Fable-tier prices to get Fable-tier coding quality inside an editor. If that half-percent gap is inside your tolerance for a given task, and for most day-to-day coding work it will be, you’re now getting nearly the same output for half the spend.
ARC-AGI 3 at three times the next-best model is worth flagging specifically because ARC-AGI is designed to resist memorization. It’s testing whether the model can genuinely reason through a problem shape it hasn’t seen before, rather than pattern-match to something similar in its training data. A 3x gap on a benchmark built to resist gaming is a stronger signal than a 3x gap on most other evals.
OSWorld 2.0 beating Fable 5’s ceiling at a third of the cost matters for anyone building computer-use agents, browser automation, or anything that drives a real desktop environment rather than just writing text. This is one of the harder categories of agentic work to get right, and Opus 5 topping it while undercutting the previous best result on price is a genuinely unusual combination.
Life sciences improvements, quantified: on organic chemistry tasks like inferring a molecule’s structure from spectroscopy data, Opus 5 scores 10.2 percentage points higher than Opus 4.8 on Anthropic’s internal benchmark. On protein-function prediction, the gain is 7.7 percentage points. If your team runs Opus on genuinely technical scientific work rather than general reasoning, those are the two numbers to weigh most heavily, because they’re domain-specific gains rather than general capability creep.
Token efficiency: the part that shows up on your invoice
The benchmark wins are one story. The efficiency story is arguably the more important one for anyone running Opus at real production volume, because it directly changes your monthly bill even at an unchanged price per token.
| Source | Reported efficiency gain over Opus 4.8 |
| Harvey (legal AI) | Similar performance to Opus 4.8’s max-reasoning mode while generating 26% fewer tokens on average |
| Unnamed trading firm | Strongest Opus yet on their trading benchmark, using roughly a seventh of the reasoning tokens and under half the latency |
| Fundamental Research Lab | 9 percentage points higher accuracy with a third fewer turns and tool calls, and 60% less time, on hard financial-modeling tasks |
| Box | 8% overall improvement, with 11% in data analysis workflows and 17% in due-diligence workflows |
| Lovable | Up 22% over Opus 4.7 on the hardest agentic coding tasks, with noticeably less run-to-run variance |
That last point, on variance, is worth pulling out separately, because it’s a different kind of improvement than raw capability. A model that’s slightly better on average but wildly inconsistent is harder to build a reliable product on than a model that’s consistently good. Multiple early-access partners specifically called out steadiness and predictability as a bigger practical win than the headline benchmark numbers, because production systems need to be trustworthy on the tenth run just as much as the first.
The trading-firm result, a seventh of the reasoning tokens at under half the latency for a better result, is the clearest single data point that this release is about efficiency more than raw peak intelligence. At an unchanged price per token, using a seventh as many tokens for a comparable or better task outcome is functionally a massive price cut, even though the sticker price didn’t move at all.
Alignment and safety: the part that isn’t about speed
Anthropic’s automated behavioral audit, which the company runs against every major model before release, scored Opus 5 at 2.3 on overall misaligned behavior, the lowest score of any recent Anthropic model, ahead of Opus 4.8, Sonnet 5, and even Fable 5. Anthropic describes it as adhering to Claude’s Constitution more closely, showing the lowest rates of deceptive behavior of its recent models, and being the hardest to trick into misuse. It’s also described as the safest model yet at avoiding reckless actions with hard-to-reverse consequences, which matters more with every release as these models get handed longer-running, less-supervised tasks.
On the cybersecurity side, Opus 5 deliberately does not chase the frontier. It comes close to Mythos 5 at finding vulnerabilities in code, per Anthropic’s OSS-Fuzz evaluation, but stays substantially behind Mythos 5 at turning those findings into working exploits, which is the more dangerous half of that capability. Anthropic built this asymmetry on purpose: Opus 5’s cyber classifiers are proportionally looser than Fable 5’s, expected to intervene around 85% less often, allowing legitimate vulnerability-finding work in source code while still blocking binary-based vulnerability scanning, penetration testing, and exploit generation. Flagged requests fall back to Opus 4.8 by default in the Claude apps, with an equivalent fallback option available on the API.
On biology, Opus 5 carries safeguards similar to Opus 4.8’s, and Anthropic now calls it its most capable generally available model for scientific research, while noting it still has real limitations on long, autonomous research tasks specifically, which is where Anthropic believes the actual biology-related risk concentrates. One practical side effect: biology-related requests that were previously blocked on Fable 5 now route to Opus 5 rather than Opus 4.8, so if your workflow touches that territory, you’ll likely notice the fallback destination changed even if you never explicitly switched models.
Two things developers should actually change
Mid-conversation tool changes, now in beta on the Claude Platform, let you swap which tools are available partway through a conversation without invalidating the prompt cache. If you’ve been avoiding tool changes mid-session because it meant eating a full cache miss, that constraint is gone, and it’s worth revisiting any workflow you previously architected around it.
Automatic API fallbacks, also new in beta, let you configure requests flagged by Opus 5’s or Fable 5’s safety classifiers to route automatically to another model instead of returning a refusal. If you’re running the kind of cost-optimized fallback chain built around retrying on hard failures, this is a second, different fallback layer worth wiring in alongside it, specifically for classifier-triggered refusals rather than outages or rate limits.

Model string is claude-opus-5 for anyone updating a routing config, and there are no data retention requirements for general access, consistent with prior Opus releases.
Is there any reason to stay on Opus 4.8?
Given identical pricing, this is a short list. The clearest one: if you’ve built extensive prompt tuning or effort-level calibration specifically around Opus 4.8’s behavior at high or max effort, and you’re getting reliable results today, there’s a real migration cost to re-validating that tuning against a model whose gearbox now has more, differently-spaced gears. That’s a legitimate reason to test before you switch rather than flip the model string in production on day one, especially for a high-stakes workflow where consistency matters more than a marginal capability gain.
Beyond that, there isn’t a strong case for staying. Same price, better or equal performance on every benchmark Anthropic published, lower token consumption reported across a wide range of independent early-access partners, and a better alignment audit score. This is closer to a straightforward upgrade than most model releases, which usually ask you to trade cost against capability. Opus 5 mostly just asks you to re-test your prompts and then update a model string.
The Bottom Line
Opus 4.8 was already a strong, workhorse-tier model. Opus 5 keeps its price, roughly doubles its performance on real-world software engineering tasks, gets within half a percent of the company’s frontier model on in-editor coding at half that model’s cost, and does much of this while burning meaningfully fewer tokens per task, according to numbers reported directly from early-access customers running it against real production workloads. The interesting story here isn’t a new capability ceiling. It’s that the floor, the everyday, budget-priced tier of Anthropic’s lineup, just moved up dramatically without moving the price tag at all.