Claude Sonnet 5.5 is Anthropic’s new mid-tier model, released on 28 September 2026 as the second member of the Claude 5.5 family. Anthropic says it generates output more than 30% faster than Claude Sonnet 5 and costs up to 30% less for most work, while scoring close to the flagship Claude Opus 5.5 on several of the company’s own benchmarks.
The price list has not changed. Claude Sonnet 5.5 costs exactly what Sonnet 5 costs per token, so the “30% cheaper” in the headline is a claim about how many tokens and tool calls a task takes, not a lower sticker price. That distinction matters for anyone budgeting an API workload, and it is where this article spends most of its time.
Below, we set out what Anthropic announced, work through the cost arithmetic with a stated example, compare the benchmark table with Sonnet 5, Opus 5.5 and OpenAI’s GPT-6 Sol, and look at the new safety controls, including one that will change behaviour for some security teams. We finish with a practical checklist for teams deciding whether to switch.
Table of contents
- What Anthropic Announced With Claude Sonnet 5.5
- Why Claude Sonnet 5.5 Is Cheaper Without a Price Cut
- How Much Faster Is Claude Sonnet 5.5?
- Claude Sonnet 5.5 Benchmarks Against Sonnet 5, Opus 5.5 and GPT-6 Sol
- Effort Levels Decide What Claude Sonnet 5.5 Really Costs
- Safety Changes That Ship With Claude Sonnet 5.5
- Claude Sonnet 5.5 Pricing Against Rivals
- Should You Switch to Claude Sonnet 5.5?
- Claude Sonnet 5.5 FAQs
- References
What Anthropic Announced With Claude Sonnet 5.5
Anthropic describes the release as “a clear upgrade over Claude Sonnet 5” that “runs 30%+ faster, and costs up to 30% less for most work.” The launch page positions it as a faster, lower-cost complement to Opus 5.5, which arrived the previous week.
The headline claims
Anthropic makes four claims for Claude Sonnet 5.5. Performance is up sharply, with 70.6% on Terminal-Bench 4.0 against Sonnet 5’s 10.3%. Writing is clearer, and early testers called it a better collaborator. Cost per task falls by up to 30% at unchanged token prices. Output generation is more than 30% faster, making it what Anthropic calls “our fastest Sonnet model to date.”
Where it sits in the Claude 5.5 family
Anthropic’s split is simple. Opus 5.5 is “built for complex work requiring careful judgment”, while Claude Sonnet 5.5 is “strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets.” A third model, Claude Haiku 5.5, is promised “in the coming weeks” for high-volume, cost-sensitive work, with no date or price yet. Our Claude 5 family guide covers how the earlier tiers were positioned.
Availability and model ID
The model is live across Anthropic’s own apps and the Claude Platform, and on Amazon Web Services, Google Cloud and Microsoft Azure. Developers call it as claude-sonnet-5-5. Like Opus 5.5 and Sonnet 5, it is available with zero data retention for eligible organisations. It is also the first Sonnet model to beat Pokémon Red working only from screenshots, a long-horizon test Anthropic uses to show sustained agentic play.
Why Claude Sonnet 5.5 Is Cheaper Without a Price Cut
The single most useful fact in the announcement is buried in the cost section: “Sonnet 5.5 is priced the same as Sonnet 5 at $2 per million input tokens, $10 per million output tokens, and $0.20 per million tokens for cache reads, but it typically needs far fewer tokens to do the same work.”
Same token prices as Sonnet 5
Cache writes cost $2.50 per million tokens. For comparison, Opus 5.5 costs $4 input, $20 output and $5 for cache writes, exactly double Claude Sonnet 5.5 on input and output. VentureBeat notes that Sonnet 5 launched in June at $2/$10 as introductory pricing, and that Anthropic made those rates permanent in August instead of moving to a planned $3/$15.
Where the 30% comes from
Anthropic’s own wording is “up to 30% less per task”, measured “in our testing”. The savings come from three behaviours early testers reported: fewer output tokens, fewer tool calls and fewer steps. In head-to-head runs, Anthropic says, the model “batched tool calls together more than Sonnet 5, leading to fewer steps and lower costs.” A task that finishes in fewer turns also re-reads less context, which cuts input tokens as well as output.
| Illustrative task | Input tokens | Output tokens | Cost per task | Cost per 10,000 tasks |
|---|---|---|---|---|
| Sonnet 5 | 400,000 | 60,000 | $1.40 | $14,000 |
| Claude Sonnet 5.5, 30% fewer tokens | 280,000 | 42,000 | $0.98 | $9,800 |
| Claude Sonnet 5.5, 10% fewer tokens | 360,000 | 54,000 | $1.26 | $12,600 |
| Opus 5.5, same tokens as the 30% case | 280,000 | 42,000 | $1.96 | $19,600 |
A worked cost example
The table uses one assumed task, an agentic coding job that consumed 400,000 uncached input tokens and 60,000 output tokens on Sonnet 5. At $2 and $10 per million, that is $0.80 plus $0.60, or $1.40. If Claude Sonnet 5.5 finishes the same job with 30% fewer tokens of each kind, the bill falls to $0.98, a saving of $4,200 on 10,000 tasks a month. If the real saving on your workload is only 10%, the monthly figure is $12,600.
The last row shows why the tier choice still matters more than the version bump. Opus 5.5 charges twice the Sonnet rate on every token, so the same 322,000-token job costs $1.96 there. Unless Opus 5.5 finishes a task in half the tokens, which Anthropic does not claim, Claude Sonnet 5.5 is the cheaper model for work it can do well.
How Much Faster Is Claude Sonnet 5.5?
Anthropic says the model “generates outputs 30%+ faster than Sonnet 5.” It has not published a tokens-per-second figure, and VentureBeat said it had asked the company how the speed-up was measured.
Output speed versus task speed
Two different speeds are in play. Output speed is how quickly tokens stream back. Task speed is how long a whole job takes, and it also depends on the number of steps and tool calls. A model that streams 30% faster and also needs a third fewer tool calls can finish a multi-step job far sooner than 30% ahead. Those two effects compound, which is why several early customers reported bigger wall-clock gains than the headline figure.
What early customers measured
Anthropic published results from 13 early-access companies. The numbers below come straight from those quotes, converted to a percentage reduction where the quote gave a before-and-after pair.
The spread is the useful part. Balyasny’s 2,441-task finance suite saw token use per answer fall by three quarters, while Box and Slack saw 12% and 14%. Those are very different workloads, and they bracket Anthropic’s “up to 30%” claim on both sides. Box also reported that the model was “2.4x faster” and “rechecks data in source documents, catching errors that Sonnet 5 failed to spot.”
| Tester | Workload | Reported result |
|---|---|---|
| Epic Games | Game system architecture | Tens of thousands of lines, multi-hour tasks, less prescriptive prompting |
| CodeRabbit | Code review | Fewer output tokens; moving simple and moderate reviews over |
| Unity | Unity Editor and coding tasks | Completed 90% of multi-step tasks |
| Base44 | 118 real app builds | Scored level with Opus 5 in 3.6 iterations versus 7.7 |
| Atlassian | Rovo Agents | Runs up to 30% faster than on Sonnet 5 |
| Zendesk | Support replies and escalations | Fewer wrong decisions; tickets 20% faster |
Claude Sonnet 5.5 Benchmarks Against Sonnet 5, Opus 5.5 and GPT-6 Sol
Anthropic’s launch table compares four models. A dash means the figure was not published for that model; GDPval-AA and AA-Briefcase are Elo-style scores run by Artificial Analysis, not percentages.
| Benchmark | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 10.3% | 66.4% (Xhigh) | – |
| FrontierCode 1.1 | 52.1% Xhigh, 46.2% Max | 42.4% | 54.4% | 49.3% |
| CursorBench 4.0 | 55.5% | 34.1% | 57.8% | – |
| GDPval-AA v2.1 | 1,844 | 1,449 | 1,846 | 1,487 |
| AA-Briefcase v1.1 | 1,811 | 1,359 | 1,822 | 1,483 |
| Humanity’s Last Exam, with tools | 64.5% | 54.9% | 67.7% | – |
| OSWorld 2.1 | 80.1% | 57.0% | 81.8% | – |
| Chartography, no tools | 61.6% | 15.6% | 64.4% | 53.6% |
Coding
Coding shows the largest jump. Terminal-Bench 4.0, which tests multi-step professional work in a command line, rose from 10.3% to 70.6%, which puts Claude Sonnet 5.5 above Opus 5.5’s best reported score. On CursorBench, built from real Cursor sessions, it trails Opus 5.5 by 2.3 points. Anthropic also says that at High effort on FrontierCode it beats Sonnet 5 at the same setting by 10 points “at about one fifteenth of the cost per task.” Our comparison of AI coding assistants explains how these tools are usually deployed.
Knowledge work and computer use
On GDPval-AA, which covers tasks from 44 occupations in nine industries, Claude Sonnet 5.5 scored 1,844, two points below Opus 5.5 and 395 above Sonnet 5. In computer use (OSWorld 2.1) it trails Opus 5.5 by 1.7 points. Chart reading improved almost fourfold. Anthropic also describes an internal test in which the model turned a public company’s earnings materials into a 10-slide operating review that two experts judged “ready to send as is.”
The FrontierCode wrinkle at Max effort
One result runs the wrong way. Claude Sonnet 5.5 scores lower on FrontierCode at Max effort (46.2%) than at Xhigh (52.1%). Anthropic’s footnote explains that at Max the model more often ran Claude Code’s code-review skill, which splits review across many subagents, and in two cases this caused a timeout or edits beyond the task’s scope. FrontierCode penalises out-of-scope changes, so more effort produced a worse score.
What the benchmarks do not show
These are the vendor’s numbers, and two footnotes deserve attention. Artificial Analysis ran GDPval-AA and AA-Briefcase on a pre-release deployment with a structured-outputs bug that has since been fixed. And several GPT-6 Sol figures are missing because OpenAI did not report them. Anthropic itself says Opus 5.5 “remains clearly stronger at complex, open-ended work requiring sustained judgment.” Independent testing has yet to confirm the speed and cost claims.
Effort Levels Decide What Claude Sonnet 5.5 Really Costs
Effort is a setting that trades speed and cost against thoroughness. At lower effort the model answers faster and uses fewer tokens. At higher effort it reasons for longer and checks its work more carefully.
Defaults differ by surface
Claude Code and the Claude apps default to Medium effort. The Claude Platform API defaults to High. So the same prompt can cost noticeably different amounts depending on where you send it, and the “30% cheaper” figure depends on which setting your workload uses. Pin the effort level explicitly in production rather than relying on a default.
Low and Medium effort beat Sonnet 5’s best
Anthropic’s cost charts make the strongest case for the release. On several benchmarks, Claude Sonnet 5.5 at Low or Medium effort beats Sonnet 5’s best score “for about a tenth of the cost per task.” On Terminal-Bench at Medium, it “far exceeds” Sonnet 5’s best for less than a tenth of the cost. On AA-Briefcase at Medium, it beats Sonnet 5’s best for about one ninth. For many teams the practical move is to drop effort by a level and re-test quality, not just to swap the model ID.
Safety Changes That Ship With Claude Sonnet 5.5
The release carries more safety machinery than any previous Sonnet model. Because it is more capable in cybersecurity, it inherits controls that until now applied only to Anthropic’s top tier.
Cyber safeguards and the Sonnet 5 fallback
Anthropic says the model’s cyber capabilities are comparable to Opus 5 and “a large improvement over Sonnet 5’s.” Routine bug finding and fixing still works, but “higher-risk cybersecurity tasks will visibly fall back to Sonnet 5.” Security teams doing offensive testing should expect some requests to be served by the older model and should read Anthropic’s real-time cyber safeguards article. A Cyber Verification Program will offer tiered access to vetted defenders.
Distillation classifiers and preserved thinking
Claude Sonnet 5.5 is the first Sonnet model to launch with classifiers that block reasoning extraction, aimed at distillation attacks in which “attackers use thousands of fake accounts to extract a model’s capabilities at industrial scale.” It also expands preserved thinking, so the model’s thinking cannot be decoupled from the account that created it. Most developers will not notice, but switching accounts mid-session in Claude Code is affected. We covered the wider fight in Anthropic’s plan to charge for rejected requests.
What the system card says
The 148-page system card concludes that Claude Sonnet 5.5 “does not cross any new RSP thresholds” and that misalignment risk is low. It reports the model as Anthropic’s most robust Sonnet yet against prompt injection. In Gray Swan’s Shade tests in coding environments, attacks succeeded on 3.01% of attempts, against 19.47% for Sonnet 5 with thinking. The card also flags weaker spots: some multi-turn regressions, including tracking and surveillance, and thinking that is “more illegible than many previous models.”
Claude Sonnet 5.5 Pricing Against Rivals
List prices per million tokens, from Anthropic’s pricing page and the launch coverage. Haiku 5.5 is not yet priced.
| Model | Input | Output | Notes |
|---|---|---|---|
| Claude Haiku 4.5 | $1 | $5 | Haiku 5.5 due “in the coming weeks” |
| Claude Sonnet 5.5 | $2 | $10 | Cache reads $0.20, cache writes $2.50 |
| Claude Sonnet 5 | $2 | $10 | Introductory price made permanent in August |
| Claude Opus 5.5 | $4 | $20 | Cache writes $5 |
| Claude Fable 5.1 | $10 | $50 | Anthropic’s most capable general model |
| OpenAI GPT-6 Sol | $2 | $10 | Cached input $0.20, cache writes $2.50 |
| Google Gemini 3.8 Flash | $0.75 | $3.75 | Introductory to end-2026, then $1.50 and $7.50 |
Anthropic’s other tiers
Inside Anthropic’s range, the gap between tiers is now the main cost lever. Opus 5.5 is exactly twice the price of Claude Sonnet 5.5 per token, and Fable 5.1 is five times. If Sonnet 5.5 handles a workload at close to Opus quality, which Anthropic’s table suggests for many structured tasks, the tier saving dwarfs the version saving. Our LLM API pricing guide sets out the wider market.
OpenAI and Google
OpenAI’s GPT-6 Sol, released the week before, lists at the same $2 and $10 as Claude Sonnet 5.5, with the same $0.20 cached-input and $2.50 cache-write rates, according to VentureBeat. Google undercuts both on headline price with Gemini 3.8 Flash at an introductory $0.75 and $3.75, rising to $1.50 and $7.50 on 1 January 2027. With list prices this close, cost per completed task is the only comparison that means much.
Should You Switch to Claude Sonnet 5.5?
For most teams already on Sonnet 5 the answer is yes, with a test first. The price is the same, the benchmarks are higher, and the early reports point to fewer tokens. The risk is behavioural change, not cost.
When Claude Sonnet 5.5 is the right default
It suits bug fixing, code review, document, slide and spreadsheet generation, support triage and structured analysis, where the task is well defined. It also suits high-volume AI agents that make many tool calls, because fewer calls mean fewer failure points as well as lower bills. Slack reported better results on its evals “without changing any of our prompts.”
When to keep Opus 5.5
Keep Opus 5.5 for ambiguous, open-ended work where judgement over many steps matters, such as system design or research. Creator’s Kevin Ngo described a sensible split: let Opus 5.5 set the architecture and let Claude Sonnet 5.5 implement it. Our AI strategy work with clients usually lands on the same routing idea.
| Check | What to do |
|---|---|
| Model ID | Change to claude-sonnet-5-5 behind a flag, and keep Sonnet 5 as a rollback |
| Thinking off | Switch to the new between_tools setting before moving, per the migration guide |
| Effort level | Set it explicitly and test one level lower than today |
| Cost tracking | Measure tokens and tool calls per completed task, not per request |
| Security workloads | Expect visible fallbacks to Sonnet 5 on higher-risk cyber requests |
| Account switching | Review preserved thinking if conversations move between accounts |
A short migration checklist
The table above covers the mechanics. Run your own evaluation set on both models at the same effort, then at one level lower on Claude Sonnet 5.5, and compare cost per completed task with a pass rate. Anthropic’s migration guide lists the breaking change for teams that run with thinking off. If your spend is mostly output tokens, expect the biggest saving there.
Claude Sonnet 5.5 FAQs
Is Claude Sonnet 5.5 actually cheaper than Sonnet 5?
Per token, no. It costs $2 per million input tokens and $10 per million output tokens, the same as Sonnet 5. Anthropic says it is up to 30% cheaper per task because it uses fewer tokens and tool calls.
How much faster is Claude Sonnet 5.5?
Anthropic says it generates output more than 30% faster than Sonnet 5, its fastest Sonnet to date. Whole tasks can finish sooner still because they take fewer steps.
Is Claude Sonnet 5.5 better than Opus 5.5?
Not overall. It beats Opus 5.5’s reported Terminal-Bench score and sits within a few points elsewhere, but Anthropic says Opus 5.5 remains clearly stronger at complex, open-ended work.
Where can I use Claude Sonnet 5.5?
In Claude’s apps, Claude Code, the Claude Platform API as claude-sonnet-5-5, and through Amazon Web Services, Google Cloud and Microsoft Azure.
When is Claude Haiku 5.5 coming?
Anthropic says “in the coming weeks” and has not given a date or a price.
References
Introducing Claude Sonnet 5.5 (Anthropic)
Claude Sonnet 5.5 System Card (Anthropic)
Sonnet 5.5 migration guide (Claude Platform docs)
Preserved thinking (Claude Platform docs)
Anthropic launches Claude Sonnet 5.5 with 30% cost reduction per task (VentureBeat)
Claude Sonnet 5.5 nearly matches Opus 5.5 on benchmarks (The Decoder)
Anthropic debuts Claude Sonnet 5.5 running 30% faster (SiliconANGLE)
Anthropic launches Claude Sonnet 5.5 with near-Opus performance (The New Stack)
Anthropic upgrades Claude with new Sonnet 5.5 model (9to5Mac)
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.