Anthropic launches Claude Sonnet 5: 40% cheaper, matches Opus 4.8 on some tasks
- Anthropic released Claude Sonnet 5, calling it the most capable Sonnet-series model yet for agentic work (autonomously executing tasks).
- Limited-time pricing is $2/$10 per million input/output tokens (through August 31, 2026), rising afterward to $3/$15; for comparison, the flagship Opus 4.8 is priced at $5/$25.
- Available today across all plans — Free, Pro, Max, Team, Enterprise — plus Claude Code and the Claude developer platform. It's the default model for Free and Pro plans.
- Safety evaluations show its overall rate of problematic behavior is lower than the previous-generation Sonnet 4.6, but its ability to develop software exploits and other cyberattack capabilities is notably weaker than Opus 4.8 — real-time cybersecurity protections are enabled by default.
- It uses a new tokenizer, so the same text may be split into more tokens (roughly 1.0 to 1.35x). The limited-time pricing already factors this in, so this upgrade works out to be roughly cost-neutral overall.
The cheap one just caught up to the expensive one
Anthropic recently released Claude Sonnet 5, calling it the most capable Sonnet-series model to date, especially strong at agentic work (meaning the model can break down tasks on its own, use tools like a browser and terminal, run through multiple steps in a row, and check its own work along the way).
Why it matters: per million output tokens, Sonnet 5 costs $15 standard, versus $25 for Opus 4.8 — exactly 60%. The limited-time price is even lower, at $2/$10 input/output. And on two benchmarks — BrowseComp (agentic search) and OSWorld-Verified (computer use) — Sonnet 5 ties Opus 4.8 once you turn the effort tier up. "Cheaper" and "reaches flagship level" have landed on the same Sonnet for the first time.
Sonnet has always been first to ship models that actually get work done
For a lot of developers, the whole "AI that just does the work" wave started with Sonnet: Claude Sonnet 3.5, 3.6, and 3.7 were the first models to really turn heads with coding and tool use. Lately, though, the most visible capability gains have come from the pricier Opus line, and this Sonnet track fell behind. What Sonnet 5 is meant to do is close that gap back up.
agentic ability
vs. Opus
gap back up
Compared to the previous-generation Sonnet 4.6, Anthropic says Sonnet 5 makes clear gains across the key areas tied to agentic performance: reasoning, tool use, coding, and knowledge work.
How much intelligence the same dollar buys now
Anthropic released two cost-versus-performance curves comparing Sonnet 5, Sonnet 4.6, and Opus 4.8 across different effort tiers — the x-axis is cost per task, the y-axis is benchmark score. The takeaway: Sonnet 5 (orange line) beats Sonnet 4.6 (grey line) across the board, covers a far wider cost range than Opus 4.8 (yellow line), delivers a clear efficiency boost at mid tiers, and at its highest tier ties Opus 4.8 on some tasks.
Two benchmark-methodology corrections (updated June 30)
Anthropic revised this launch post on June 30: the original BrowseComp chart used a simpler method that understated Sonnet 5's performance, and has now been redrawn using the standard method from the system card (10-million-token budget + compression + programmatic tool calls). Two other legacy scores were also corrected due to updated scoring methods: Humanity's Last Exam's Sonnet 4.6 score is now 34.6% (no tools) / 46.8% (with tools); OSWorld-Verified's Sonnet 4.6 score is now 78.5%. These differ from the figures in the Sonnet 4.6 launch blog post simply because the scoring method changed.
Spend more effort, think one step further: how one model pulls off both cheap and top-tier
The reason Sonnet 5 can span such a wide price range comes down to a mechanism called effort (a compute/reasoning-intensity tier): the same model, but you get to choose how hard it "thinks." Lower tiers are cheaper and faster but may be less thorough; higher tiers spend more compute on repeated reasoning and self-checking — more accurate answers, but pricier and slower.
It's like ordering the same dish at the same restaurant — you can have the chef make it the usual way, or pay extra to have them put in more care and taste it themselves before it goes out to make sure it's right. It's the same dish either way; what changes is how much care went in and how many times it got checked. The effort tier is dialing that "level of care."
may not be thorough enough
clear efficiency gain
more accurate and stable
ties Opus 4.8 on some tasks
In the past, getting more capability meant switching to a bigger, pricier model. Now you don't switch models at all — you just turn a dial: the low tier is a cheap, fast entry option; the high tier (xhigh) spends more compute on repeated reasoning and self-checking, tying the flagship Opus 4.8 on some tasks. One single Sonnet 5 now spans the entire range from entry-level to near-flagship in one go, instead of hitting a ceiling early like Sonnet 4.6 did. Where to strike the balance between cost and performance is now up to you, project by project.
Early user feedback: no nagging required — it checks its own work
Anthropic says feedback from early-access partners has been remarkably consistent: Sonnet 5 is noticeably more capable at working autonomously than previous generations. Here's what testers described in objective terms.
- On complex tasks, it keeps going until the job is done — earlier Sonnet generations often stopped partway through.
- It checks its own output for correctness without being explicitly asked to.
- The pricing for this kind of autonomous work is still quite appealing.
Safer overall — but cyberattack capability was deliberately held back
Pre-deployment safety evaluations show Sonnet 5 is safer overall than Sonnet 4.6: better at refusing malicious requests, more resistant to prompt injection (where an attacker hides malicious instructions inside a webpage or email the model is processing, trying to hijack the model into following the attacker's instructions instead of the user's), and less prone to hallucination and sycophancy. In an automated behavioral audit covering a range of problematic behaviors, it scored lower overall (meaning safer), though still higher than the stronger Opus 4.8 and Claude Mythos Preview.
Cybersecurity is the one area deliberately held back. Anthropic says it did not specifically train Sonnet 5 on cybersecurity tasks: it can handle routine, benign network-related tasks, but on evaluations for developing potentially harmful software exploits, it performs notably worse than Opus 4.8 and Mythos 5.
A concrete test: can it hack its way through Firefox
"Weak at cyberattacks" sounds abstract, so Anthropic gave a concrete number: having each model develop an exploit for a vulnerability in the Firefox browser. This evaluation was jointly developed by Anthropic and Mozilla, and all vulnerabilities involved have already been fixed in Firefox 148.
Neither Sonnet model produced a single complete, working exploit (both at 0.0%). Sonnet 5 only edged out Sonnet 4.6 slightly on partial success rate, which Anthropic attributes mostly to spillover from general intelligence gains rather than targeted training. For comparison, both Opus 4.8 and Mythos 5 have far stronger cyberattack capabilities than either Sonnet model.
Because Sonnet 5 is slightly stronger than the previous generation on this kind of task, Anthropic has enabled real-time cybersecurity protections by default — able to detect and block dangerous network-related use in real time, at the same tier as Claude Opus 4.7 and 4.8. Anthropic judges Sonnet 5's overall cybersecurity risk to be low, so these protections are less restrictive than the ones on Fable 5 (which blocks a far broader range of cybersecurity-related tasks).
Looks like a price cut, but the ruler changed too
Sonnet 5 uses a new tokenizer. Before a model can process text, it first has to be split into individual tokens for computation and billing. With the new tokenizer, the same piece of text may get split into more tokens — roughly 1.0 to 1.35x, depending on content type. So while the per-token price has dropped, the same passage now consumes more tokens, meaning the actual per-unit cost hasn't dropped by as much as it looks.
Per million tokens, price has dropped from Sonnet 4.6's level to the limited-time $2/$10.
The limited-time pricing was set specifically to offset the tokenizer change, making the switch from Sonnet 4.6 to Sonnet 5 roughly cost-neutral overall.
Anthropic states outright: the limited-time pricing was set so that this upgrade works out to be close to cost-neutral. That's why "price cut" needs air quotes — you have to do the math with token counts included. This tokenizer adjustment follows the same approach used for Claude Opus 4.7.
Available now: rollout, pricing, and how to choose
Sonnet 5 is available today across every plan: it's the default model for Free and Pro plans, and Max, Team, and Enterprise users can access it too; it's also live in Claude Code and the Claude developer platform, where developers can call it via the Claude API as claude-sonnet-5.
| Model | Input / Output (per million tokens) |
|---|---|
| Sonnet 5 (limited-time, through 2026-08-31) | $2 / $10 |
| Sonnet 5 (standard price afterward) | $3 / $15 |
| Opus 4.8 (for comparison) | $5 / $25 |
$2 / $10
Limited-time
$3 / $15
Standard price
To accommodate the higher token consumption that comes with higher effort tiers, Anthropic has already raised rate limits across Chat, Cowork, Claude Code, and the Claude developer platform — so you can pick the tier that fits each project.
How to choose
| Who you are | Recommendation |
|---|---|
| Developer | If your budget's the same and you want stronger agentic coding and tool use: go high tier. If you want to save money: dial effort down and get near-flagship results at lower cost — find your own balance between cost and performance. |
| Enterprise / team | Rate limits for Chat, Cowork, Claude Code, and the developer platform have already been raised to accommodate the higher token consumption at higher tiers. |
| Security-related work | Default cybersecurity protections match Opus 4.7/4.8. For work needing fewer restrictions on cybersecurity research or offense/defense work, Anthropic recommends switching to Opus 4.8 instead of Sonnet 5. |
Sonnet 5 closes the gap: it performs close to Opus 4.8, but at a lower price. Anthropic, "Introducing Claude Sonnet 5"
Want more power? No need to switch to a pricier model — just turn one dial, and the cheaper Sonnet 5 catches up to flagship Opus
Anthropic launches Claude Sonnet 5: standard price is only 60% of flagship Opus 4.8's, and official benchmarks show it tying Opus at higher effort — here's the whole story in one page with a chart.
↓ Read this page in full · one chart moves
Claude is Anthropic's AI assistant, and Sonnet is its "mid-tier" product line. The new version is called Sonnet 5, and Anthropic says it's the most capable Sonnet yet at working autonomously (agentic — meaning the model can break tasks down on its own, use tools like a browser and terminal, run through several steps in a row, and check its own work along the way).
✘ But it couldn't reach the capability ceiling of its own flagship, Opus
Lately, the pricier Opus line has been gaining capability the fastest. Wanting the strongest model used to mean paying up for a bigger one — the cheaper previous-gen Sonnet 4.6 hit its ceiling early.
Sonnet 5 closes that gap: it's priced well below the flagship, yet official benchmarks show it tying Opus 4.8 on some tasks at higher tiers. Early user feedback is consistent too — it keeps working on complex tasks all the way through instead of stopping halfway, and it checks its own work without being asked.
Your only options were the pricier Opus 4.8 ($25 per million output), or living with Sonnet 4.6's ceiling.
Standard price is $15 per million output — turn the effort dial to max and it ties Opus 4.8 on some tasks; want to save money, just dial it down.
How can the same model be both cheap and flagship-capable? The key is a dial called effort.
effort (the compute tier) lets you choose how hard it "thinks": at a low tier, it's cheap and fast but may not be thorough enough; at a high tier, it spends more compute on repeated reasoning and self-checking — more accurate, but pricier and slower. Turn the dial from the lowest tier, low, all the way to the highest, xhigh, and the same Sonnet 5 spans everything from a cheap tier to near-flagship.
Anthropic gives figures as "dollars per million tokens," which doesn't mean much to most people. It's more intuitive translated into what the same task actually costs — take a mid-sized task producing about 1 million output tokens and compare what each of the three options costs.
the strongest one...
hits its ceiling early.
Used to be one way,
caught the expensive one!
this much of flagship,
cheap AND top-tier?
call it "think harder."
- × cyberattack capability deliberately held back
- × new tokenizer, gotta count tokens too
- × every number here is Anthropic's own self-report
just turn the dial up.
