Claude Sonnet 5.5 is Anthropic’s new mid-tier model, released on 28 September 2026 as the second member of the Claude 5.5 family. Anthropic says it generates output more than 30% faster than Claude Sonnet 5 and costs up to 30% less for most work, while scoring close to the flagship Claude Opus 5.5 on several of the company’s own benchmarks.

The price list has not changed. Claude Sonnet 5.5 costs exactly what Sonnet 5 costs per token, so the “30% cheaper” in the headline is a claim about how many tokens and tool calls a task takes, not a lower sticker price. That distinction matters for anyone budgeting an API workload, and it is where this article spends most of its time.

Below, we set out what Anthropic announced, work through the cost arithmetic with a stated example, compare the benchmark table with Sonnet 5, Opus 5.5 and OpenAI’s GPT-6 Sol, and look at the new safety controls, including one that will change behaviour for some security teams. We finish with a practical checklist for teams deciding whether to switch.

What Anthropic Announced With Claude Sonnet 5.5

Claude Sonnet 5.5 - claude sonnet 5 5 30 faster 30 cheaper b three stacks of coins

Anthropic describes the release as “a clear upgrade over Claude Sonnet 5” that “runs 30%+ faster, and costs up to 30% less for most work.” The launch page positions it as a faster, lower-cost complement to Opus 5.5, which arrived the previous week.

The headline claims

Anthropic makes four claims for Claude Sonnet 5.5. Performance is up sharply, with 70.6% on Terminal-Bench 4.0 against Sonnet 5’s 10.3%. Writing is clearer, and early testers called it a better collaborator. Cost per task falls by up to 30% at unchanged token prices. Output generation is more than 30% faster, making it what Anthropic calls “our fastest Sonnet model to date.”

Where it sits in the Claude 5.5 family

Anthropic’s split is simple. Opus 5.5 is “built for complex work requiring careful judgment”, while Claude Sonnet 5.5 is “strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets.” A third model, Claude Haiku 5.5, is promised “in the coming weeks” for high-volume, cost-sensitive work, with no date or price yet. Our Claude 5 family guide covers how the earlier tiers were positioned.

Availability and model ID

The model is live across Anthropic’s own apps and the Claude Platform, and on Amazon Web Services, Google Cloud and Microsoft Azure. Developers call it as claude-sonnet-5-5. Like Opus 5.5 and Sonnet 5, it is available with zero data retention for eligible organisations. It is also the first Sonnet model to beat Pokémon Red working only from screenshots, a long-horizon test Anthropic uses to show sustained agentic play.

Why Claude Sonnet 5.5 Is Cheaper Without a Price Cut

claude sonnet 5 5 30 faster 30 cheaper c lightning bolt on a stand

The single most useful fact in the announcement is buried in the cost section: “Sonnet 5.5 is priced the same as Sonnet 5 at $2 per million input tokens, $10 per million output tokens, and $0.20 per million tokens for cache reads, but it typically needs far fewer tokens to do the same work.”

Same token prices as Sonnet 5

Cache writes cost $2.50 per million tokens. For comparison, Opus 5.5 costs $4 input, $20 output and $5 for cache writes, exactly double Claude Sonnet 5.5 on input and output. VentureBeat notes that Sonnet 5 launched in June at $2/$10 as introductory pricing, and that Anthropic made those rates permanent in August instead of moving to a planned $3/$15.

Where the 30% comes from

Anthropic’s own wording is “up to 30% less per task”, measured “in our testing”. The savings come from three behaviours early testers reported: fewer output tokens, fewer tool calls and fewer steps. In head-to-head runs, Anthropic says, the model “batched tool calls together more than Sonnet 5, leading to fewer steps and lower costs.” A task that finishes in fewer turns also re-reads less context, which cuts input tokens as well as output.

Illustrative taskInput tokensOutput tokensCost per taskCost per 10,000 tasks
Sonnet 5400,00060,000$1.40$14,000
Claude Sonnet 5.5, 30% fewer tokens280,00042,000$0.98$9,800
Claude Sonnet 5.5, 10% fewer tokens360,00054,000$1.26$12,600
Opus 5.5, same tokens as the 30% case280,00042,000$1.96$19,600

A worked cost example

The table uses one assumed task, an agentic coding job that consumed 400,000 uncached input tokens and 60,000 output tokens on Sonnet 5. At $2 and $10 per million, that is $0.80 plus $0.60, or $1.40. If Claude Sonnet 5.5 finishes the same job with 30% fewer tokens of each kind, the bill falls to $0.98, a saving of $4,200 on 10,000 tasks a month. If the real saving on your workload is only 10%, the monthly figure is $12,600.

The last row shows why the tier choice still matters more than the version bump. Opus 5.5 charges twice the Sonnet rate on every token, so the same 322,000-token job costs $1.96 there. Unless Opus 5.5 finishes a task in half the tokens, which Anthropic does not claim, Claude Sonnet 5.5 is the cheaper model for work it can do well.

How Much Faster Is Claude Sonnet 5.5?

claude sonnet 5 5 30 faster 30 cheaper d two meshing gear wheels

Anthropic says the model “generates outputs 30%+ faster than Sonnet 5.” It has not published a tokens-per-second figure, and VentureBeat said it had asked the company how the speed-up was measured.

Output speed versus task speed

Two different speeds are in play. Output speed is how quickly tokens stream back. Task speed is how long a whole job takes, and it also depends on the number of steps and tool calls. A model that streams 30% faster and also needs a third fewer tool calls can finish a multi-step job far sooner than 30% ahead. Those two effects compound, which is why several early customers reported bigger wall-clock gains than the headline figure.

What early customers measured

Anthropic published results from 13 early-access companies. The numbers below come straight from those quotes, converted to a percentage reduction where the quote gave a before-and-after pair.

Efficiency gains early testers reported for Claude Sonnet 5.5 against the model each compared it with (% reduction)
Balyasny, tokens per finance answer (497k to 121k) 75.7%
Base44, iterations per app build vs Opus 5 (7.7 to 3.6) 53.2%
Lovable, shell runs per coding task (roughly half) 50%
Lovable, tool calls per coding task (a third fewer) 33%
Zendesk, support ticket processing time 20%
Slack, output tokens on Slackbot evals 14%
Box, total tokens 12%

The spread is the useful part. Balyasny’s 2,441-task finance suite saw token use per answer fall by three quarters, while Box and Slack saw 12% and 14%. Those are very different workloads, and they bracket Anthropic’s “up to 30%” claim on both sides. Box also reported that the model was “2.4x faster” and “rechecks data in source documents, catching errors that Sonnet 5 failed to spot.”

TesterWorkloadReported result
Epic GamesGame system architectureTens of thousands of lines, multi-hour tasks, less prescriptive prompting
CodeRabbitCode reviewFewer output tokens; moving simple and moderate reviews over
UnityUnity Editor and coding tasksCompleted 90% of multi-step tasks
Base44118 real app buildsScored level with Opus 5 in 3.6 iterations versus 7.7
AtlassianRovo AgentsRuns up to 30% faster than on Sonnet 5
ZendeskSupport replies and escalationsFewer wrong decisions; tickets 20% faster

Claude Sonnet 5.5 Benchmarks Against Sonnet 5, Opus 5.5 and GPT-6 Sol

claude sonnet 5 5 30 faster 30 cheaper e three step winners podium

Anthropic’s launch table compares four models. A dash means the figure was not published for that model; GDPval-AA and AA-Briefcase are Elo-style scores run by Artificial Analysis, not percentages.

BenchmarkSonnet 5.5Sonnet 5Opus 5.5GPT-6 Sol
Terminal-Bench 4.070.6%10.3%66.4% (Xhigh)–
FrontierCode 1.152.1% Xhigh, 46.2% Max42.4%54.4%49.3%
CursorBench 4.055.5%34.1%57.8%–
GDPval-AA v2.11,8441,4491,8461,487
AA-Briefcase v1.11,8111,3591,8221,483
Humanity’s Last Exam, with tools64.5%54.9%67.7%–
OSWorld 2.180.1%57.0%81.8%–
Chartography, no tools61.6%15.6%64.4%53.6%

Coding

Coding shows the largest jump. Terminal-Bench 4.0, which tests multi-step professional work in a command line, rose from 10.3% to 70.6%, which puts Claude Sonnet 5.5 above Opus 5.5’s best reported score. On CursorBench, built from real Cursor sessions, it trails Opus 5.5 by 2.3 points. Anthropic also says that at High effort on FrontierCode it beats Sonnet 5 at the same setting by 10 points “at about one fifteenth of the cost per task.” Our comparison of AI coding assistants explains how these tools are usually deployed.

Gain from Sonnet 5 to Claude Sonnet 5.5, percentage points (from Anthropic’s launch table)
Terminal-Bench 4.0 (10.3% to 70.6%) +60.3
Chartography (15.6% to 61.6%) +46.0
OSWorld 2.1 (57.0% to 80.1%) +23.1
CursorBench 4.0 (34.1% to 55.5%) +21.4
Humanity’s Last Exam (54.9% to 64.5%) +9.6
FrontierCode 1.1 at Max (42.4% to 46.2%) +3.8

Knowledge work and computer use

On GDPval-AA, which covers tasks from 44 occupations in nine industries, Claude Sonnet 5.5 scored 1,844, two points below Opus 5.5 and 395 above Sonnet 5. In computer use (OSWorld 2.1) it trails Opus 5.5 by 1.7 points. Chart reading improved almost fourfold. Anthropic also describes an internal test in which the model turned a public company’s earnings materials into a 10-slide operating review that two experts judged “ready to send as is.”

The FrontierCode wrinkle at Max effort

One result runs the wrong way. Claude Sonnet 5.5 scores lower on FrontierCode at Max effort (46.2%) than at Xhigh (52.1%). Anthropic’s footnote explains that at Max the model more often ran Claude Code’s code-review skill, which splits review across many subagents, and in two cases this caused a timeout or edits beyond the task’s scope. FrontierCode penalises out-of-scope changes, so more effort produced a worse score.

What the benchmarks do not show

These are the vendor’s numbers, and two footnotes deserve attention. Artificial Analysis ran GDPval-AA and AA-Briefcase on a pre-release deployment with a structured-outputs bug that has since been fixed. And several GPT-6 Sol figures are missing because OpenAI did not report them. Anthropic itself says Opus 5.5 “remains clearly stronger at complex, open-ended work requiring sustained judgment.” Independent testing has yet to confirm the speed and cost claims.

Effort Levels Decide What Claude Sonnet 5.5 Really Costs

claude sonnet 5 5 30 faster 30 cheaper f clipboard with a checklist

Effort is a setting that trades speed and cost against thoroughness. At lower effort the model answers faster and uses fewer tokens. At higher effort it reasons for longer and checks its work more carefully.

Defaults differ by surface

Claude Code and the Claude apps default to Medium effort. The Claude Platform API defaults to High. So the same prompt can cost noticeably different amounts depending on where you send it, and the “30% cheaper” figure depends on which setting your workload uses. Pin the effort level explicitly in production rather than relying on a default.

Low and Medium effort beat Sonnet 5’s best

Anthropic’s cost charts make the strongest case for the release. On several benchmarks, Claude Sonnet 5.5 at Low or Medium effort beats Sonnet 5’s best score “for about a tenth of the cost per task.” On Terminal-Bench at Medium, it “far exceeds” Sonnet 5’s best for less than a tenth of the cost. On AA-Briefcase at Medium, it beats Sonnet 5’s best for about one ninth. For many teams the practical move is to drop effort by a level and re-test quality, not just to swap the model ID.

Safety Changes That Ship With Claude Sonnet 5.5

The release carries more safety machinery than any previous Sonnet model. Because it is more capable in cybersecurity, it inherits controls that until now applied only to Anthropic’s top tier.

Cyber safeguards and the Sonnet 5 fallback

Anthropic says the model’s cyber capabilities are comparable to Opus 5 and “a large improvement over Sonnet 5’s.” Routine bug finding and fixing still works, but “higher-risk cybersecurity tasks will visibly fall back to Sonnet 5.” Security teams doing offensive testing should expect some requests to be served by the older model and should read Anthropic’s real-time cyber safeguards article. A Cyber Verification Program will offer tiered access to vetted defenders.

Distillation classifiers and preserved thinking

Claude Sonnet 5.5 is the first Sonnet model to launch with classifiers that block reasoning extraction, aimed at distillation attacks in which “attackers use thousands of fake accounts to extract a model’s capabilities at industrial scale.” It also expands preserved thinking, so the model’s thinking cannot be decoupled from the account that created it. Most developers will not notice, but switching accounts mid-session in Claude Code is affected. We covered the wider fight in Anthropic’s plan to charge for rejected requests.

What the system card says

The 148-page system card concludes that Claude Sonnet 5.5 “does not cross any new RSP thresholds” and that misalignment risk is low. It reports the model as Anthropic’s most robust Sonnet yet against prompt injection. In Gray Swan’s Shade tests in coding environments, attacks succeeded on 3.01% of attempts, against 19.47% for Sonnet 5 with thinking. The card also flags weaker spots: some multi-turn regressions, including tracking and surveillance, and thinking that is “more illegible than many previous models.”

Claude Sonnet 5.5 Pricing Against Rivals

List prices per million tokens, from Anthropic’s pricing page and the launch coverage. Haiku 5.5 is not yet priced.

ModelInputOutputNotes
Claude Haiku 4.5$1$5Haiku 5.5 due “in the coming weeks”
Claude Sonnet 5.5$2$10Cache reads $0.20, cache writes $2.50
Claude Sonnet 5$2$10Introductory price made permanent in August
Claude Opus 5.5$4$20Cache writes $5
Claude Fable 5.1$10$50Anthropic’s most capable general model
OpenAI GPT-6 Sol$2$10Cached input $0.20, cache writes $2.50
Google Gemini 3.8 Flash$0.75$3.75Introductory to end-2026, then $1.50 and $7.50

Anthropic’s other tiers

Inside Anthropic’s range, the gap between tiers is now the main cost lever. Opus 5.5 is exactly twice the price of Claude Sonnet 5.5 per token, and Fable 5.1 is five times. If Sonnet 5.5 handles a workload at close to Opus quality, which Anthropic’s table suggests for many structured tasks, the tier saving dwarfs the version saving. Our LLM API pricing guide sets out the wider market.

OpenAI and Google

OpenAI’s GPT-6 Sol, released the week before, lists at the same $2 and $10 as Claude Sonnet 5.5, with the same $0.20 cached-input and $2.50 cache-write rates, according to VentureBeat. Google undercuts both on headline price with Gemini 3.8 Flash at an introductory $0.75 and $3.75, rising to $1.50 and $7.50 on 1 January 2027. With list prices this close, cost per completed task is the only comparison that means much.

Should You Switch to Claude Sonnet 5.5?

For most teams already on Sonnet 5 the answer is yes, with a test first. The price is the same, the benchmarks are higher, and the early reports point to fewer tokens. The risk is behavioural change, not cost.

When Claude Sonnet 5.5 is the right default

It suits bug fixing, code review, document, slide and spreadsheet generation, support triage and structured analysis, where the task is well defined. It also suits high-volume AI agents that make many tool calls, because fewer calls mean fewer failure points as well as lower bills. Slack reported better results on its evals “without changing any of our prompts.”

When to keep Opus 5.5

Keep Opus 5.5 for ambiguous, open-ended work where judgement over many steps matters, such as system design or research. Creator’s Kevin Ngo described a sensible split: let Opus 5.5 set the architecture and let Claude Sonnet 5.5 implement it. Our AI strategy work with clients usually lands on the same routing idea.

CheckWhat to do
Model IDChange to claude-sonnet-5-5 behind a flag, and keep Sonnet 5 as a rollback
Thinking offSwitch to the new between_tools setting before moving, per the migration guide
Effort levelSet it explicitly and test one level lower than today
Cost trackingMeasure tokens and tool calls per completed task, not per request
Security workloadsExpect visible fallbacks to Sonnet 5 on higher-risk cyber requests
Account switchingReview preserved thinking if conversations move between accounts

A short migration checklist

The table above covers the mechanics. Run your own evaluation set on both models at the same effort, then at one level lower on Claude Sonnet 5.5, and compare cost per completed task with a pass rate. Anthropic’s migration guide lists the breaking change for teams that run with thinking off. If your spend is mostly output tokens, expect the biggest saving there.

Claude Sonnet 5.5 FAQs

Is Claude Sonnet 5.5 actually cheaper than Sonnet 5?

Per token, no. It costs $2 per million input tokens and $10 per million output tokens, the same as Sonnet 5. Anthropic says it is up to 30% cheaper per task because it uses fewer tokens and tool calls.

How much faster is Claude Sonnet 5.5?

Anthropic says it generates output more than 30% faster than Sonnet 5, its fastest Sonnet to date. Whole tasks can finish sooner still because they take fewer steps.

Is Claude Sonnet 5.5 better than Opus 5.5?

Not overall. It beats Opus 5.5’s reported Terminal-Bench score and sits within a few points elsewhere, but Anthropic says Opus 5.5 remains clearly stronger at complex, open-ended work.

Where can I use Claude Sonnet 5.5?

In Claude’s apps, Claude Code, the Claude Platform API as claude-sonnet-5-5, and through Amazon Web Services, Google Cloud and Microsoft Azure.

When is Claude Haiku 5.5 coming?

Anthropic says “in the coming weeks” and has not given a date or a price.

References