Turn ideas into shipped features at lightspeed with Friday.
Get Started
Claude Opus 5 vs Claude Opus 4.8
The internet hasn't been kind to Claude Opus 5, so I'll be.

Claude Opus 5 vs Claude Opus 4.8: Direct Coding Comparison

Article Contents

Anthropic shipped Claude Opus 5 on July 24, 2026, roughly two months after Opus 4.8, which is the same upgrade cadence the company has kept for the entire Opus line. What’s unusual isn’t the timing. It’s the pricing decision: Opus 5 costs exactly what Opus 4.8 cost, $5 per million input tokens and $25 per million output tokens, and Anthropic is framing the entire release around getting meaningfully more model for that same dollar rather than charging more for more capability. That’s a different kind of upgrade than most model releases, and it changes how you should think about whether to switch. 

Now, if you’re willing to overlook many of the Reddit posts regarding the Opus 5 and genuinely curious about learning how much it improves upon the Opus 4.8, you’re at the right place.

Opus 5 vs Opus 4.8 – What has changed?

Anthropic’s own positioning for Opus 5 is that it comes close to the frontier intelligence of its top-tier model, Claude Fable 5, at half the price of that model, while staying at Opus 4.8’s existing price point. It’s now the default model on Claude Max and the strongest model available on Claude Pro. On two of Anthropic’s headline evaluations, Frontier-Bench and GDPval-AA, Opus 5 is the new state of the art among Anthropic’s models. However, it remains behind Mythos 5 specifically on cybersecurity tasks, which is a deliberate safety choice rather than a capability gap (more on that below).

The short version: this isn’t a “smarter but pricier” release. It’s a “same price, meaningfully more capable, and far more efficient” release, aimed squarely at the workloads teams run every single day rather than at chasing a benchmark crown.

Spec sheet, side by side

Claude Opus 4.8Claude Opus 5
ReleasedMay 28, 2026July 24, 2026
Input price$5.00 / million tokens$5.00 / million tokens
Output price$25.00 / million tokens$25.00 / million tokens
Context window1,000,000 tokens1,000,000 tokens
Fast Mode$10 / $50 at 2.5x speed$10 / $50 at 2.5x speed
Effort controlHigh / extra / maxLow through max, with finer-grained token-vs-intelligence tradeoffs
Default tier placementWas default on Claude MaxNew default on Claude Max, strongest on Claude Pro
Cyber safeguard strictnessBaseline for its generationClassifiers intervene roughly 85% less often than on Fable 5, looser than 4.8’s baseline
Alignment audit scoreHigher misaligned-behavior score than Opus 5Lowest misaligned-behavior score (2.3) of any recent Anthropic model

Nothing moved on price, context window, or Fast Mode’s speed multiplier. Everything moved on how much useful work the model does inside those same constraints, which is a more interesting kind of upgrade to evaluate, because it means the honest question isn’t “is Opus 5 better,” it’s “how much better, on what, and does that translate to fewer tokens for the same job.”


Here’s how I’d explain the difference to someone who doesn’t want to read a benchmark table. Opus 4.8 is a car with three gears: high, extra, and max effort. You pick a gear, and it’s a reasonable engine either way. Opus 5 is the same car with a proper gearbox, a wider spread of effort settings from low up through max, tuned so that even the lowest gear now clears obstacles the old car needed second or third gear to get over.

That reframes the whole “which effort level should I use” question. With Opus 4.8, going cheap meant giving something up. With Opus 5, several of Anthropic’s own benchmark charts show the low-effort setting beating other models’ best results on the same task, meaning the cheap gear isn’t a compromise anymore so much as a different point on a curve that’s shifted upward across its entire length.

Opus 5 vs Opus 4.8 – Benchmark comparison

BenchmarkWhat it testsThe Opus 5 vs Opus 4.8 result
Frontier-Bench v0.1Real-world software engineering tasks on an agent harnessOpus 5 more than doubles Opus 4.8’s performance, at a lower cost per task
CursorBench 3.2In-editor coding agent performanceAt max effort, Opus 5 lands within 0.5% of Fable 5’s peak score, at half Fable 5’s cost per task
ARC-AGI 3Novel problem-solving the model hasn’t seen trained patterns forOpus 5 scores roughly three times higher than the next-best model
Zapier AutomationBenchEnd-to-end business task completionRoughly 1.5x the pass rate of the next-best model at the same cost; even Opus 5’s lowest effort setting beats every other model
OSWorld 2.0Computer-use benchmark: operating a real desktop environmentOpus 5 outperforms every model at any cost, and beats Fable 5’s best result at about a third of Fable 5’s cost
Life sciences (internal suite)Structural biology, organic chemistry, bioinformaticsBetter than Opus 4.8 on every single evaluation in the suite

Translating a few of these into what they mean day to day:

Frontier-Bench more than doubling is the number that matters if your team runs Opus on real engineering tickets rather than toy coding problems. This benchmark specifically measures messy, real-world software tasks on an agent harness, not clean leetcode-style problems, so a 2x-plus jump there is closer to “the agent finishes twice as many real tasks correctly” than “the model got a better score on a quiz.”

CursorBench within 0.5% of Fable 5 at half the cost is the number for anyone who’s been paying Fable-tier prices to get Fable-tier coding quality inside an editor. If that half-percent gap is inside your tolerance for a given task, and for most day-to-day coding work it will be, you’re now getting nearly the same output for half the spend.

ARC-AGI 3 at three times the next-best model is worth flagging specifically because ARC-AGI is designed to resist memorization. It’s testing whether the model can genuinely reason through a problem shape it hasn’t seen before, rather than pattern-match to something similar in its training data. A 3x gap on a benchmark built to resist gaming is a stronger signal than a 3x gap on most other evals.

OSWorld 2.0 beating Fable 5’s ceiling at a third of the cost matters for anyone building computer-use agents, browser automation, or anything that drives a real desktop environment rather than just writing text. This is one of the harder categories of agentic work to get right, and Opus 5 topping it while undercutting the previous best result on price is a genuinely unusual combination.

Life sciences improvements, quantified: on organic chemistry tasks like inferring a molecule’s structure from spectroscopy data, Opus 5 scores 10.2 percentage points higher than Opus 4.8 on Anthropic’s internal benchmark. On protein-function prediction, the gain is 7.7 percentage points. If your team runs Opus on genuinely technical scientific work rather than general reasoning, those are the two numbers to weigh most heavily, because they’re domain-specific gains rather than general capability creep.

Token efficiency: the part that shows up on your invoice

The benchmark wins are one story. The efficiency story is arguably the more important one for anyone running Opus at real production volume, because it directly changes your monthly bill even at an unchanged price per token.

SourceReported efficiency gain over Opus 4.8
Harvey (legal AI)Similar performance to Opus 4.8’s max-reasoning mode while generating 26% fewer tokens on average
Unnamed trading firmStrongest Opus yet on their trading benchmark, using roughly a seventh of the reasoning tokens and under half the latency
Fundamental Research Lab9 percentage points higher accuracy with a third fewer turns and tool calls, and 60% less time, on hard financial-modeling tasks
Box8% overall improvement, with 11% in data analysis workflows and 17% in due-diligence workflows
LovableUp 22% over Opus 4.7 on the hardest agentic coding tasks, with noticeably less run-to-run variance

That last point, on variance, is worth pulling out separately, because it’s a different kind of improvement than raw capability. A model that’s slightly better on average but wildly inconsistent is harder to build a reliable product on than a model that’s consistently good. Multiple early-access partners specifically called out steadiness and predictability as a bigger practical win than the headline benchmark numbers, because production systems need to be trustworthy on the tenth run just as much as the first.

The trading-firm result, a seventh of the reasoning tokens at under half the latency for a better result, is the clearest single data point that this release is about efficiency more than raw peak intelligence. At an unchanged price per token, using a seventh as many tokens for a comparable or better task outcome is functionally a massive price cut, even though the sticker price didn’t move at all.

Alignment and safety: the part that isn’t about speed

Anthropic’s automated behavioral audit, which the company runs against every major model before release, scored Opus 5 at 2.3 on overall misaligned behavior, the lowest score of any recent Anthropic model, ahead of Opus 4.8, Sonnet 5, and even Fable 5. Anthropic describes it as adhering to Claude’s Constitution more closely, showing the lowest rates of deceptive behavior of its recent models, and being the hardest to trick into misuse. It’s also described as the safest model yet at avoiding reckless actions with hard-to-reverse consequences, which matters more with every release as these models get handed longer-running, less-supervised tasks.

On the cybersecurity side, Opus 5 deliberately does not chase the frontier. It comes close to Mythos 5 at finding vulnerabilities in code, per Anthropic’s OSS-Fuzz evaluation, but stays substantially behind Mythos 5 at turning those findings into working exploits, which is the more dangerous half of that capability. Anthropic built this asymmetry on purpose: Opus 5’s cyber classifiers are proportionally looser than Fable 5’s, expected to intervene around 85% less often, allowing legitimate vulnerability-finding work in source code while still blocking binary-based vulnerability scanning, penetration testing, and exploit generation. Flagged requests fall back to Opus 4.8 by default in the Claude apps, with an equivalent fallback option available on the API.

On biology, Opus 5 carries safeguards similar to Opus 4.8’s, and Anthropic now calls it its most capable generally available model for scientific research, while noting it still has real limitations on long, autonomous research tasks specifically, which is where Anthropic believes the actual biology-related risk concentrates. One practical side effect: biology-related requests that were previously blocked on Fable 5 now route to Opus 5 rather than Opus 4.8, so if your workflow touches that territory, you’ll likely notice the fallback destination changed even if you never explicitly switched models.

Two things developers should actually change

Mid-conversation tool changes, now in beta on the Claude Platform, let you swap which tools are available partway through a conversation without invalidating the prompt cache. If you’ve been avoiding tool changes mid-session because it meant eating a full cache miss, that constraint is gone, and it’s worth revisiting any workflow you previously architected around it.

Automatic API fallbacks, also new in beta, let you configure requests flagged by Opus 5’s or Fable 5’s safety classifiers to route automatically to another model instead of returning a refusal. If you’re running the kind of cost-optimized fallback chain built around retrying on hard failures, this is a second, different fallback layer worth wiring in alongside it, specifically for classifier-triggered refusals rather than outages or rate limits.

Read: How to Route LLM Requests: Building a Cost-Optimized Fallback Chain

Model string is claude-opus-5 for anyone updating a routing config, and there are no data retention requirements for general access, consistent with prior Opus releases.

Is there any reason to stay on Opus 4.8?

Given identical pricing, this is a short list. The clearest one: if you’ve built extensive prompt tuning or effort-level calibration specifically around Opus 4.8’s behavior at high or max effort, and you’re getting reliable results today, there’s a real migration cost to re-validating that tuning against a model whose gearbox now has more, differently-spaced gears. That’s a legitimate reason to test before you switch rather than flip the model string in production on day one, especially for a high-stakes workflow where consistency matters more than a marginal capability gain.

Beyond that, there isn’t a strong case for staying. Same price, better or equal performance on every benchmark Anthropic published, lower token consumption reported across a wide range of independent early-access partners, and a better alignment audit score. This is closer to a straightforward upgrade than most model releases, which usually ask you to trade cost against capability. Opus 5 mostly just asks you to re-test your prompts and then update a model string.

The Bottom Line

Opus 4.8 was already a strong, workhorse-tier model. Opus 5 keeps its price, roughly doubles its performance on real-world software engineering tasks, gets within half a percent of the company’s frontier model on in-editor coding at half that model’s cost, and does much of this while burning meaningfully fewer tokens per task, according to numbers reported directly from early-access customers running it against real production workloads. The interesting story here isn’t a new capability ceiling. It’s that the floor, the everyday, budget-priced tier of Anthropic’s lineup, just moved up dramatically without moving the price tag at all.

Friday

The AI workspace that turns prompts into results.

Plan, research, and ship faster with AI that understands your work.

From PRD to production before the week is over. Build with Friday AI

Available on:

tryfriday.ai
product_team_goals:
time_to_market: "shipped_in_hours"
dev_alignment: "prds_to_clean_code"
overhead: "zero_waste_meetings"
sprint_status: features_deployed_successfully...

The Autonomous Product Team for Founders Product Managers Leaders

Friday AI turns your PC into an autonomous product team. Combines the best models from Claude, OpenAI, Gemini, and ElevenLabs into one seamless operator that doesn’t just think, it executes.