Product Launch · XiaoHu Explains

Anthropic launches Claude Sonnet 5: 40% cheaper, matches Opus 4.8 on some tasks

Official benchmarks show that at higher effort tiers, it matches Opus 4.8 on some tasks — yet standard pricing is only 60% of Opus's price.
TL;DR
  • Anthropic released Claude Sonnet 5, calling it the most capable Sonnet-series model yet for agentic work (autonomously executing tasks).
  • Limited-time pricing is $2/$10 per million input/output tokens (through August 31, 2026), rising afterward to $3/$15; for comparison, the flagship Opus 4.8 is priced at $5/$25.
  • Available today across all plans — Free, Pro, Max, Team, Enterprise — plus Claude Code and the Claude developer platform. It's the default model for Free and Pro plans.
  • Safety evaluations show its overall rate of problematic behavior is lower than the previous-generation Sonnet 4.6, but its ability to develop software exploits and other cyberattack capabilities is notably weaker than Opus 4.8 — real-time cybersecurity protections are enabled by default.
  • It uses a new tokenizer, so the same text may be split into more tokens (roughly 1.0 to 1.35x). The limited-time pricing already factors this in, so this upgrade works out to be roughly cost-neutral overall.
Stance note: all data in this article comes from Anthropic's official launch page and its own benchmarks and system card — this is vendor self-reporting. The performance comparisons, safety scores, and success rates below are Anthropic's official figures or self-reported data, faithfully relayed below without further fact-checking.
1 The One-Line Version

The cheap one just caught up to the expensive one

Anthropic recently released Claude Sonnet 5, calling it the most capable Sonnet-series model to date, especially strong at agentic work (meaning the model can break down tasks on its own, use tools like a browser and terminal, run through multiple steps in a row, and check its own work along the way).

Here's the counterintuitive part: Sonnet 5's standard price is only 60% of the flagship Opus 4.8's — but official benchmarks show that when you turn the effort tier up, it can match Opus 4.8 on some tasks.
📌

Why it matters: per million output tokens, Sonnet 5 costs $15 standard, versus $25 for Opus 4.8 — exactly 60%. The limited-time price is even lower, at $2/$10 input/output. And on two benchmarks — BrowseComp (agentic search) and OSWorld-Verified (computer use) — Sonnet 5 ties Opus 4.8 once you turn the effort tier up. "Cheaper" and "reaches flagship level" have landed on the same Sonnet for the first time.

← Cheaper · entry-level effortPricier · flagship effort →
Sonnet 4.6
Sonnet 5
Opus 4.8
Illustrative: Opus 4.8 is a fixed high point. Sonnet 4.6 only covers a short, low stretch before hitting its ceiling early; Sonnet 5 stretches this band way out via its effort tiers, with its top end closing in on Opus 4.8. Same money, wider range of capability with Sonnet 5.
Claude Sonnet 5 benchmark comparison table
Official benchmark comparison: Sonnet 5 versus the previous-generation Sonnet 4.6, with the more all-around Opus 4.8 for reference. Full benchmarks in the Claude Sonnet 5 system card. Source: Anthropic's official site.
2 Backstory

Sonnet has always been first to ship models that actually get work done

For a lot of developers, the whole "AI that just does the work" wave started with Sonnet: Claude Sonnet 3.5, 3.6, and 3.7 were the first models to really turn heads with coding and tool use. Lately, though, the most visible capability gains have come from the pricier Opus line, and this Sonnet track fell behind. What Sonnet 5 is meant to do is close that gap back up.

3.5
First to show
agentic ability
3.6
3.7
4.6
Gap widens
vs. Opus
5
Closes the
gap back up

Compared to the previous-generation Sonnet 4.6, Anthropic says Sonnet 5 makes clear gains across the key areas tied to agentic performance: reasoning, tool use, coding, and knowledge work.

3 Value For Money

How much intelligence the same dollar buys now

Anthropic released two cost-versus-performance curves comparing Sonnet 5, Sonnet 4.6, and Opus 4.8 across different effort tiers — the x-axis is cost per task, the y-axis is benchmark score. The takeaway: Sonnet 5 (orange line) beats Sonnet 4.6 (grey line) across the board, covers a far wider cost range than Opus 4.8 (yellow line), delivers a clear efficiency boost at mid tiers, and at its highest tier ties Opus 4.8 on some tasks.

Sonnet 5 Sonnet 4.6 Opus 4.8
Opus 4.8 level Score ↑ Cost (per task) →
Illustrative curve: as effort tier rises, Sonnet 5's score keeps climbing, closing in on Opus 4.8 at the top tier, and covering a far wider cost range than Sonnet 4.6. See the official chart below for exact figures.
Opus 4.8 level Score ↑ Cost (per task) →
Illustrative curve: same story for computer-use tasks — a clear efficiency boost at mid tiers, and the top tier ties Opus 4.8 on some tasks. See the official chart below for exact figures.
Cost-performance curves across effort tiers
Official cost-performance curves: the previous-generation Sonnet 4.6 clearly can't reach Opus 4.8; Sonnet 5 covers a far wider cost range and ties Opus 4.8 on some tasks at its top tier. Sonnet 5 in the chart is priced at the standard $3/$15 — at the limited-time $2/$10, actual cost is even lower. xhigh = the highest effort tier. Source: Anthropic's official site.
Two benchmark-methodology corrections (updated June 30)

Anthropic revised this launch post on June 30: the original BrowseComp chart used a simpler method that understated Sonnet 5's performance, and has now been redrawn using the standard method from the system card (10-million-token budget + compression + programmatic tool calls). Two other legacy scores were also corrected due to updated scoring methods: Humanity's Last Exam's Sonnet 4.6 score is now 34.6% (no tools) / 46.8% (with tools); OSWorld-Verified's Sonnet 4.6 score is now 78.5%. These differ from the figures in the Sonnet 4.6 launch blog post simply because the scoring method changed.

4 Core Mechanism

Spend more effort, think one step further: how one model pulls off both cheap and top-tier

The reason Sonnet 5 can span such a wide price range comes down to a mechanism called effort (a compute/reasoning-intensity tier): the same model, but you get to choose how hard it "thinks." Lower tiers are cheaper and faster but may be less thorough; higher tiers spend more compute on repeated reasoning and self-checking — more accurate answers, but pricier and slower.

An Analogy

It's like ordering the same dish at the same restaurant — you can have the chef make it the usual way, or pay extra to have them put in more care and taste it themselves before it goes out to make sure it's right. It's the same dish either way; what changes is how much care went in and how many times it got checked. The effort tier is dialing that "level of care."

low
Cheap and fast
may not be thorough enough
medium
Value sweet spot
clear efficiency gain
high
Reasons harder
more accurate and stable
xhigh
Top tier
ties Opus 4.8 on some tasks
Taller bar = thinks harder = more accurate, but also costs more compute (money and time). xhigh means extra high, the top tier.
Core Innovation

In the past, getting more capability meant switching to a bigger, pricier model. Now you don't switch models at all — you just turn a dial: the low tier is a cheap, fast entry option; the high tier (xhigh) spends more compute on repeated reasoning and self-checking, tying the flagship Opus 4.8 on some tasks. One single Sonnet 5 now spans the entire range from entry-level to near-flagship in one go, instead of hitting a ceiling early like Sonnet 4.6 did. Where to strike the balance between cost and performance is now up to you, project by project.

5 Early Feedback

Early user feedback: no nagging required — it checks its own work

Anthropic says feedback from early-access partners has been remarkably consistent: Sonnet 5 is noticeably more capable at working autonomously than previous generations. Here's what testers described in objective terms.

  • On complex tasks, it keeps going until the job is done — earlier Sonnet generations often stopped partway through.
  • It checks its own output for correctness without being explicitly asked to.
  • The pricing for this kind of autonomous work is still quite appealing.
6 Safety Evaluation

Safer overall — but cyberattack capability was deliberately held back

Pre-deployment safety evaluations show Sonnet 5 is safer overall than Sonnet 4.6: better at refusing malicious requests, more resistant to prompt injection (where an attacker hides malicious instructions inside a webpage or email the model is processing, trying to hijack the model into following the attacker's instructions instead of the user's), and less prone to hallucination and sycophancy. In an automated behavioral audit covering a range of problematic behaviors, it scored lower overall (meaning safer), though still higher than the stronger Opus 4.8 and Claude Mythos Preview.

Sonnet 4.6
Sonnet 5
Opus 4.8
Mythos Preview
Relative illustration (longer bar = higher rate of problematic behavior = less safe): Sonnet 5 is lower than Sonnet 4.6, but higher than Opus 4.8 and Mythos Preview. Bar lengths are relative ordering, not exact values — see the official chart below for specifics.
Rate of problematic behavior across Claude models
Rate of problematic behavior across models in Anthropic's automated behavioral audit: Sonnet 5 is overall lower than Sonnet 4.6 (safer), but higher than Mythos Preview and Opus 4.8. Full list in system card section 6.4. Source: Anthropic's official site.

Cybersecurity is the one area deliberately held back. Anthropic says it did not specifically train Sonnet 5 on cybersecurity tasks: it can handle routine, benign network-related tasks, but on evaluations for developing potentially harmful software exploits, it performs notably worse than Opus 4.8 and Mythos 5.

7 Test Case · Firefox

A concrete test: can it hack its way through Firefox

"Weak at cyberattacks" sounds abstract, so Anthropic gave a concrete number: having each model develop an exploit for a vulnerability in the Firefox browser. This evaluation was jointly developed by Anthropic and Mozilla, and all vulnerabilities involved have already been fixed in Firefox 148.

0.0%
Sonnet 5's success rate at fully developing a working exploit
0.0%
Sonnet 4.6's full success rate, tied with Sonnet 5

Neither Sonnet model produced a single complete, working exploit (both at 0.0%). Sonnet 5 only edged out Sonnet 4.6 slightly on partial success rate, which Anthropic attributes mostly to spillover from general intelligence gains rather than targeted training. For comparison, both Opus 4.8 and Mythos 5 have far stronger cyberattack capabilities than either Sonnet model.

Model scores on the Firefox 147 exploit-development evaluation
Firefox 147 exploit-development evaluation (jointly developed by Anthropic and Mozilla; the relevant vulnerabilities have been fixed in Firefox 148): for each model, the left bar is full-success rate for a working exploit, and the right bar is partial-success rate. Both Sonnet models scored 0.0% on full success, with Sonnet 5 slightly ahead of Sonnet 4.6 on partial success; both are far below Opus 4.8 and Mythos 5. See system card section 3.2.4 for details. Source: Anthropic's official site.

Because Sonnet 5 is slightly stronger than the previous generation on this kind of task, Anthropic has enabled real-time cybersecurity protections by default — able to detect and block dangerous network-related use in real time, at the same tier as Claude Opus 4.7 and 4.8. Anthropic judges Sonnet 5's overall cybersecurity risk to be low, so these protections are less restrictive than the ones on Fable 5 (which blocks a far broader range of cybersecurity-related tasks).

8 Pricing Fine Print

Looks like a price cut, but the ruler changed too

Sonnet 5 uses a new tokenizer. Before a model can process text, it first has to be split into individual tokens for computation and billing. With the new tokenizer, the same piece of text may get split into more tokens — roughly 1.0 to 1.35x, depending on content type. So while the per-token price has dropped, the same passage now consumes more tokens, meaning the actual per-unit cost hasn't dropped by as much as it looks.

Old tokenizer (Sonnet 4.6)splits into fewer tokens
thesamepassagefedintothe model
New tokenizer (Sonnet 5)splits into more tokens (roughly 1.0 to 1.35x)
thesamepassagefedintothemodel
Illustrative only: the split shown here is a demonstration, not the actual tokenization boundaries. Under the new tokenizer, the same passage gets chopped up finer, pushing up the token count.
Price alone

Per million tokens, price has dropped from Sonnet 4.6's level to the limited-time $2/$10.

Factoring in more tokens

The limited-time pricing was set specifically to offset the tokenizer change, making the switch from Sonnet 4.6 to Sonnet 5 roughly cost-neutral overall.

Anthropic states outright: the limited-time pricing was set so that this upgrade works out to be close to cost-neutral. That's why "price cut" needs air quotes — you have to do the math with token counts included. This tokenizer adjustment follows the same approach used for Claude Opus 4.7.

9 Available Now

Available now: rollout, pricing, and how to choose

Sonnet 5 is available today across every plan: it's the default model for Free and Pro plans, and Max, Team, and Enterprise users can access it too; it's also live in Claude Code and the Claude developer platform, where developers can call it via the Claude API as claude-sonnet-5.

ModelInput / Output (per million tokens)
Sonnet 5 (limited-time, through 2026-08-31)$2 / $10
Sonnet 5 (standard price afterward)$3 / $15
Opus 4.8 (for comparison)$5 / $25
Now
$2 / $10
Limited-time
From 2026-09-01
$3 / $15
Standard price

To accommodate the higher token consumption that comes with higher effort tiers, Anthropic has already raised rate limits across Chat, Cowork, Claude Code, and the Claude developer platform — so you can pick the tier that fits each project.

How to choose

Who you areRecommendation
DeveloperIf your budget's the same and you want stronger agentic coding and tool use: go high tier. If you want to save money: dial effort down and get near-flagship results at lower cost — find your own balance between cost and performance.
Enterprise / teamRate limits for Chat, Cowork, Claude Code, and the developer platform have already been raised to accommodate the higher token consumption at higher tiers.
Security-related workDefault cybersecurity protections match Opus 4.7/4.8. For work needing fewer restrictions on cybersecurity research or offense/defense work, Anthropic recommends switching to Opus 4.8 instead of Sonnet 5.
Sonnet 5 closes the gap: it performs close to Opus 4.8, but at a lower price. Anthropic, "Introducing Claude Sonnet 5"
This article is based on Anthropic's official post "Introducing Claude Sonnet 5" (including the June 30, 2026 correction) and the Claude Sonnet 5 system card. Benchmark scores, success rates, and pricing in this article are Anthropic's official figures and self-reported data; some charts are official originals, with illustrative charts/orderings labeled as such. Actual performance may vary in real-world use.
Translation complete — all visible text (headings, body copy, captions, chart labels, comic dialogue/SFX, alt text) rendered into English while CSS, SVG paths/coordinates, classes, IDs, and script logic were left untouched.