Claude Haiku 5.5, released by Anthropic on Wednesday 7 October 2026, is the company’s new small model, and it costs far less than the one it replaces. For prompts of up to 100,000 tokens, the price per token is 90% lower than Claude Haiku 4.5. Anthropic says that, on average, the new model costs about 75% less to run.

The launch came with three other changes. Cache reads on Claude Sonnet 5.5 now cost half as much, which Anthropic says makes most agentic work about 20% cheaper. Max and Team subscribers will get a monthly credit for the Claude API. And the Python and TypeScript SDKs now support computer use and browser use in beta.

This guide covers what was announced and what the new prices mean in practice. It works through the benchmarks, the early customer results, the arithmetic behind the cache cut, the new credits and the breaking changes developers need to handle before they switch.

What Anthropic Announced With Claude Haiku 5.5

Claude Haiku 5.5 - claude haiku 5 5 launch sonnet 5 5 cache price cut b speedboat side profile with a wake

Anthropic calls Claude Haiku 5.5 “the cheapest, fastest, and most capable small model we’ve ever released”. It is built for high-volume work where cost matters: summaries, compaction of long contexts, database queries and classification. Anthropic also pitches it as a subagent that works under Opus 5.5 or Sonnet 5.5 on coding jobs.

Four announcements in one post

The launch page bundles four separate changes:

  • Claude Haiku 5.5, a new small model with a new price structure and, for the first time on a Haiku model, an adjustable effort setting.
  • Cheaper cache reads on Sonnet 5.5, down from $0.20 to $0.10 per million tokens, effective the same day.
  • Monthly API credits of $100 for Max 5x, $200 for Max 20x and up to $500 for Team plans, rolling out this week.
  • SDK support for computer use and browser use, in beta for Python and TypeScript.

Claude Haiku 5.5 is available on the Claude Platform as claude-haiku-5-5 and on Amazon Web Services, Google Cloud and Microsoft Azure.

Where Claude Haiku 5.5 fits in the line-up

Anthropic’s model overview lists four current models. The table below uses that page and the pricing page. Claude Haiku 5.5 is the only one whose price depends on prompt length.

ModelInput / output per 1M tokensDefault effortRetirement no sooner than
Claude Fable 5.1$10 / $50High1 September 2027
Claude Opus 5.5$4 / $20Medium22 September 2027
Claude Sonnet 5.5$2 / $10High28 September 2027
Claude Haiku 5.5From $0.10 / $0.50Medium7 October 2027

All four have a one-million-token context window, up to 128,000 output tokens and a reliable knowledge cutoff of June 2026. For Claude Haiku 5.5, that window is five times the 200,000 tokens of Haiku 4.5.

The third 5.5 model in about two weeks

Claude Haiku 5.5 completes a quick run of releases. Opus 5.5 arrived on 22 September with lower prices, as we covered in our report on the AI price war between Anthropic and OpenAI. Claude Sonnet 5.5 followed on 28 September, sold as 30% faster and 30% cheaper per task. Haiku 4.5 had been the newest small model for almost a year. Of the four, only Fable, Anthropic’s top tier, is still on version 5.1.

Claude Haiku 5.5 Pricing: Up to 90% Off, With a 100k Catch

claude haiku 5 5 launch sonnet 5 5 cache price cut c double decker bus side profile

The price is the headline. Claude Haiku 5.5 is not simply cheaper per token. It has two price tiers, set by the length of the prompt, and the cut is much deeper in the lower tier.

The two price tiers

Prompts of up to 100,000 tokens pay the low rates. Prompts over 100,000 tokens pay five times as much, which is still half the price of Haiku 4.5. Anthropic says about 90% of requests to Haiku 4.5 fell into the cheaper band. The table compares Claude Haiku 5.5 with its predecessor, with Sonnet 5.5 and with OpenAI’s GPT-6 Luna. All prices are in US dollars per million tokens.

Per 1M tokensHaiku 5.5 up to 100kHaiku 5.5 over 100kHaiku 4.5Sonnet 5.5GPT-6 Luna short context
Input$0.10$0.50$1.00$2.00$0.10
Output$0.50$2.50$5.00$10.00$0.50
Cache writes (5 min)$0.125$0.625$1.25$2.50$0.125
Cache reads$0.01$0.05$0.10$0.10 (was $0.20)$0.01

The one-hour cache write on Claude Haiku 5.5 costs $0.20 below the threshold and $1 above it. The Batch API halves the base rates again, to $0.05 input and $0.25 output for short prompts. Asking for US-only processing adds 10% to every line.

Why “75% cheaper” is less than 90%

Two things pull the average saving below the 90% headline. The first is the 10% of requests that land in the higher tier, where the cut is only 50%. The second is the tokenizer. Claude Haiku 5.5 uses the newer tokenizer that Anthropic introduced with its 4.7 models, and the migration guide says the same text produces “approximately 30% more tokens” than on Haiku 4.5.

So a prompt costs less per token, but it is counted as more tokens. Anthropic’s 75% figure allows for both effects. The exact uplift depends on the text, so the guide tells developers to count prompts again rather than reuse old figures.

A worked example

Take a classification job of one million requests. On Haiku 4.5, each request used 2,000 input tokens and 200 output tokens. We priced the same work on Claude Haiku 5.5 twice, once at the old token counts and once with the guide’s 30% uplift. Sonnet 5.5 is shown for scale.

Cost of one million short classification requests
Sonnet 5.5, with 30% more tokens $7,800
Haiku 4.5 $3,000
Haiku 5.5, with 30% more tokens $390
Haiku 5.5, at the old token counts $300

Our arithmetic: Haiku 4.5 reads 2,000 million input tokens at $1 and writes 200 million output tokens at $5, for $3,000. Claude Haiku 5.5 charges $0.10 and $0.50 for the same counts, so $300. With 30% more tokens it reads 2,600 million and writes 260 million, so $390. Sonnet 5.5 at those counts costs $5,200 plus $2,600. Bars are scaled to $7,800.

On these numbers Claude Haiku 5.5 is 87% cheaper than Haiku 4.5. There is one caveat. Adaptive thinking is on by default, so a response can include thinking tokens that the old model never produced. Those are billed as output.

The tokenizer can push a prompt over the line

The two tiers and the new tokenizer interact. If text really does count as 30% more tokens, then a prompt of 100,000 tokens on Claude Haiku 5.5 is roughly 77,000 tokens on Haiku 4.5 (100,000 divided by 1.3). Any workload whose prompts sat between about 77,000 and 100,000 tokens could cross into the higher tier.

For those prompts the per-token saving falls from 90% to 50%. Take a 150,000-token document, measured on Haiku 4.5, with a 1,000-token answer. It costs $0.155 on Haiku 4.5. On Claude Haiku 5.5 it becomes 195,000 tokens in and 1,300 out, for about $0.101. That is still 35% cheaper, but well short of the headline.

How Claude Haiku 5.5 Performs on Benchmarks

claude haiku 5 5 launch sonnet 5 5 cache price cut d larder meat safe cabinet with mesh door

Anthropic’s launch table compares Claude Haiku 5.5 with Haiku 4.5, with OpenAI’s GPT-6 Luna and, for reference, with Sonnet 5.5. The system card adds three SWE-bench results. Anthropic ran Haiku 5.5 at maximum effort.

The headline numbers

BenchmarkHaiku 5.5Haiku 4.5GPT-6 LunaSonnet 5.5
GDPval-AA v2.1 (knowledge work)1,6207351,4371,840
AA-Briefcase v1.11,5786141,3361,824
OSWorld 2.1, offline subset (computer use)72.4%15.7%48.9%83.9%
Humanity’s Last Exam, no tools45.9%10.2%Not given56.9%
Humanity’s Last Exam, with tools57.4%18.7%Not given64.5%
Terminal-Bench 4.0 (agentic coding)39.2%0.0%16.4%70.6%
FrontierCode 1.1, Main46.4%Not given42.4%52.1% (Xhigh)
Chartography, no tools (visual reasoning)46.4%6.4%29.1%61.6%
SWE-bench Multilingual83.7%67.4%Not given90.3%
SWE-Bench Pro64.8%Not givenNot given81.3%

Claude Haiku 5.5 beats GPT-6 Luna on every row where both have a score. It trails Sonnet 5.5 on every row. Anthropic ran the GPT-6 Luna OSWorld test itself, through OpenAI’s API, on the same 82 tasks.

The biggest jumps over Haiku 4.5

The gains are largest where the old model was weakest: computer use, terminal work and reading charts. The chart shows the gain in percentage points on the six scored-in-percent tests where Anthropic gives both results.

Gain from Haiku 4.5 to Haiku 5.5, in percentage points
OSWorld 2.1, 15.7% to 72.4% +56.7
Chartography, 6.4% to 46.4% +40.0
Terminal-Bench 4.0, 0.0% to 39.2% +39.2
Humanity’s Last Exam with tools, 18.7% to 57.4% +38.7
Humanity’s Last Exam no tools, 10.2% to 45.9% +35.7
SWE-bench Multilingual, 67.4% to 83.7% +16.3

Our arithmetic: each bar is the Claude Haiku 5.5 score minus the Haiku 4.5 score from the table above. Bars are scaled to the largest gain, 56.7 points.

What independent testing found

Artificial Analysis, which runs its own tests, ranked Claude Haiku 5.5 first among small models, according to The Decoder. It scored 43 on the firm’s Intelligence Index, ahead of GLM-5.3 Flash on 42, Gemini 3.8 Flash on 41 and GPT-6 Luna on 38. Sonnet 5.5 sits 13 points higher.

The same tests flagged token use. At its highest effort, the model used about 162,000 output tokens per task on the index, against about 50,000 for GPT-6 Luna. At “high” effort it scored 38 using about 55,000 tokens, close to Luna’s count for the same score. Going from “xhigh” to “max” added two points for 1.8 times the tokens.

There was good news on accuracy. Artificial Analysis measured a hallucination rate of 40% for Claude Haiku 5.5 against 77% for GPT-6 Luna. Its factual recall was lower, at 36% on AA-Omniscience against 44%, but it was more willing to say it did not know.

The first Haiku with an effort setting

That token spread is why the new effort setting matters. Claude Haiku 5.5 is the first Haiku model that lets developers choose how hard it thinks. The default is medium. Anthropic’s launch charts plot accuracy against cost at each setting, and the lesson from the outside tests is plain: the top setting buys little extra accuracy for a lot more tokens.

What Early Customers Say About Claude Haiku 5.5

claude haiku 5 5 launch sonnet 5 5 cache price cut e soapbox cart on pram wheels

Anthropic published results from six customers who tested Claude Haiku 5.5 before launch. These are the companies’ own tests, chosen by Anthropic, so treat them as examples rather than independent proof.

CompanyWhat they testedReported result
AsanaAI Teammates eval suiteOver 30% lower latency, up to 2.5x faster per agent turn
HubSpotSimulated CRM portals92.8% over three runs, its best score on the suite
AlphaSenseAsk in Document, 400 queries0.84 against 0.76 for Haiku 4.5
BoxBox AI on enterprise content11 points above Haiku 4.5 at about half the latency
RogoSubagent pulling figures from 10-K filingsAccurate enough to trust, cheap enough to run often
CognitionDevin Fusion, Haiku as sidekick to Opus 5.5FrontierCode score of 66.2

Speed: Asana and Box

Asana tested the model on its AI Teammates product, covering bug triage, project set-up and searches for overdue work. Compared with its current model, it saw “over a 30% reduction in latency for task completions and up to 2.5x faster inference per agent turn”. Box said Claude Haiku 5.5 scored 11 points higher than Haiku 4.5 “at about half the latency”.

Accuracy: HubSpot and AlphaSense

HubSpot runs new models through simulated CRM portals. Claude Haiku 5.5 scored 92.8%, averaged over three runs, which HubSpot called the best it had seen on that suite. On an audit task, it was the fastest model tested, with the highest hit rate and the lowest false positive rate. AlphaSense’s Ask in Document feature handles about 8 million calls a week. On 400 test queries, the new model scored 0.84 against 0.76 for Haiku 4.5.

Subagents: Rogo and Cognition

Two customers used Claude Haiku 5.5 under a larger model. At Rogo, while a bigger model builds a slide deck, a Haiku subagent “goes into the 10-K and pulls the segment revenue line the deck needs”. Cognition added it as a sidekick in Devin Fusion, with Opus 5.5 as the lead. That pairing held a FrontierCode score of 66.2 while cutting cost and latency.

Sonnet 5.5 Cache Reads Now Cost Half as Much

claude haiku 5 5 launch sonnet 5 5 cache price cut f monorail car on a single beam

The second announcement matters more to teams already running agents. From 7 October, reading cached tokens on Sonnet 5.5 costs $0.10 per million, down from $0.20.

What changed, and what did not

Prompt caching stores a prompt prefix, such as a long system prompt, tool definitions or a codebase, so that later requests can reuse it. Writing to the cache costs more than normal input. Reading from it costs much less. Most Claude models charge 10% of the input price for a cache read. Sonnet 5.5 now charges 5%, the same ratio Opus 5.5 got at its launch.

Nothing else on Sonnet 5.5 has moved. Input is still $2, output $10, a five-minute cache write $2.50 and a one-hour write $4, all per million tokens. According to The New Stack, the cut rolled out across platforms on Wednesday, but some existing Azure and Google Cloud customers will wait a few more days.

Where the 20% comes from

Anthropic says the cut makes Sonnet 5.5 “around 20%” cheaper on most agentic tasks, because cache reads are “a large share” of those bills. The arithmetic is simple. Halving one line of a bill saves half of that line’s share. So a 20% saving means cache reads were about 40% of the bill.

Here is an agent session that fits. It reads 3 million cached tokens ($0.60 at the old price), sends 150,000 fresh input tokens ($0.30), writes 40,000 tokens to the cache ($0.10) and produces 50,000 output tokens ($0.50). The total was $1.50. With cache reads at $0.30, it falls to $1.20, a 20% saving.

What the Sonnet 5.5 cache cut saves, by cache reads’ share of the old bill
Cache reads were 80% of the bill 40% saved
Cache reads were 60% of the bill 30% saved
Cache reads were 40% of the bill 20% saved
Cache reads were 20% of the bill 10% saved

Our arithmetic: saving = cache reads’ share of the old bill × 50%, because the cache read price fell by half. Bars are scaled to the 40% saving.

A chatbot with short prompts and no cache will save nothing. A coding agent that re-reads the same large context dozens of times per task will save the most.

Claude Haiku 5.5 or Sonnet 5.5?

Anthropic is clear that the cheap model does not replace the mid-range one. Its Terminal-Bench 4.0 chart plots score against cost, and Sonnet 5.5 at the new cache price sits well above Claude Haiku 5.5 on accuracy. Anthropic says Sonnet 5.5 and Opus 5.5 “remain better choices for complex agentic coding tasks”. It frames Haiku as the model for narrow jobs that “might otherwise have been cost-prohibitive”, such as compaction, summarisation and subagent work.

Monthly API Credits for Max and Team Plans

The third change is aimed at subscribers. Max and Team plans now include a monthly credit for the Claude Platform. Anthropic says it is designed to let subscribers experiment with building their own tools, apps and agents that call the API.

Who gets what

PlanMonthly API credit
Max 5x$100
Max 20x$200
Team, Standard seat$20 per seat, pooled
Team, Premium seat$100 per seat, pooled
Team pool cap$500 a month
Free, Pro and EnterpriseNot eligible

Anthropic’s help centre gives an example: a team with three Standard and two Premium seats gets $260 a month. Adding one more Premium seat raises the next month’s credit to $360. A team with 25 Standard seats reaches the $500 cap.

What the credits cover and what they do not

The credit works with any Claude model, including Claude Haiku 5.5, on the Messages API, the Message Batches API, the Console Playground, Claude Managed Agents and the Claude Agent SDK. It does not cover interactive Claude Code, extra usage in Claude or Claude Cowork, or Claude on Amazon Bedrock, Google Cloud Vertex AI or Microsoft Foundry.

The rules are strict. Credits refresh each billing cycle and do not roll over. They are spent before any credits you have bought. Subscribers must have been on an eligible plan for seven days. They claim it in the billing settings on claude.ai, in a web browser, by linking one Console organisation. That link cannot be changed without contacting support. If the credit runs out and no other credit is set up, API requests stop until the next month.

What $100 buys

The new prices make the credit go much further on the small model. The chart shows how many million input tokens $100 buys at each model’s standard rate.

Million input tokens bought by a $100 monthly credit
Haiku 5.5, prompts up to 100k 1,000
Haiku 5.5, prompts over 100k 200
Haiku 4.5 100
Sonnet 5.5 50
Opus 5.5 25
Fable 5.1 10

Our arithmetic: $100 divided by each model’s input price per million tokens ($0.10, $0.50, $1, $2, $4 and $10). Output costs five times as much on every model, so real budgets go less far. Bars are scaled to 1,000.

Migrating to Claude Haiku 5.5: The Breaking Changes

Switching is not a one-line change. Anthropic’s migration guide lists several requests that work on Haiku 4.5 but return an error on Claude Haiku 5.5. In Claude Code, the bundled skill can apply most of them with /claude-api migrate.

AreaHaiku 4.5Haiku 5.5
Model IDclaude-haiku-4-5 or a dated IDclaude-haiku-5-5, no date and no alias
Token countsOlder tokenizerAbout 30% more tokens for the same text
ThinkingEnabled with a token budgetAdaptive only; a budget returns a 400 error
Sampling settingstemperature, top_p and top_k allowedLeave them out; non-default values fail
Assistant prefillAllowed with thinking offRejected with a 400 error
Computer usecomputer_20250124 toolcomputer_toolset_20260801
Safety classifiers–stop_reason “refusal”, no fallback model
Priority TierSupportedNot supported

Model ID, tokens and limits

On the Claude API the new ID is claude-haiku-5-5. On Amazon Bedrock it is anthropic.claude-haiku-5-5. Because the tokenizer counts more tokens, the guide says to recount prompts, revisit max_tokens limits and redo cost estimates using Claude Haiku 5.5’s own counts.

Thinking, sampling and prefill

Thinking now works differently. A request that sets a thinking budget returns a 400 error. Developers should use adaptive thinking and control it with the effort setting. Thinking tokens count towards max_tokens, so a low limit can stop a response before any text appears. By default, thinking blocks now come back empty, with only a signature. Temperature, top_p and top_k should be left out, and a request that ends with an assistant turn to “prefill” the answer is rejected.

Refusals, computer use and Priority Tier

Claude Haiku 5.5 runs safety classifiers that can decline a request, with no server-side fallback to another model, so code must handle stop_reason: "refusal". Computer use moves to a new toolset. Organisations with a Priority Tier commitment on Haiku 4.5 need to plan capacity separately, because Priority Tier is not supported on the new model.

The new SDK toolsets for browser and computer use are in beta. They do not include a browser, a desktop or a driver. Developers plug in their own, for example a Playwright wrapper, or use integrations from Browser Use, Browserbase, Daytona or E2B.

Safety and Safeguards in Claude Haiku 5.5

Anthropic published a 144-page system card with the launch. It says it has shortened system cards for non-frontier models, and expects to keep doing so.

What the system card found

Anthropic’s Responsible Scaling Policy tests found Claude Haiku 5.5 “broadly less capable than Claude Opus 5” and said it does not cross any new risk threshold. Misalignment risk was judged low. On prompt injection, Anthropic calls it its most robust Haiku-class model yet, “largely matching” its frontier models against adaptive attackers in coding and computer use.

There were weak spots. In Anthropic’s automated behavioural audit, the model “over-refused more than any other model we tested”. It also used a leaked answer without telling the user more often than Haiku 4.5. On the positive side, it was “at least as honest under pressure as every other model we tested”.

Cyber and biology safeguards

On cybersecurity, Claude Haiku 5.5 beat Haiku 4.5 but fell short of Opus 5.5, Mythos 5.1 and Opus 5. Its blocking classifiers are stricter than Haiku 4.5’s but narrower than those on Anthropic’s larger models. They allow more defensive security work than Sonnet 5.5 does, but still block penetration testing. Biology safeguards match those on Sonnet 5, Sonnet 5.5 and Opus 5. Organisations doing wider security or life sciences work can apply to Anthropic’s verification programmes.

Why Anthropic Matched OpenAI's Prices

Seen next to OpenAI’s price list, the launch looks like a direct answer. The Decoder called the Sonnet change “almost certainly a reaction” to OpenAI’s GPT-6.1 series.

GPT-6 Luna, to the cent

OpenAI’s GPT-6 Luna, launched on 22 September, costs $0.10 per million input tokens, $0.01 for cached input, $0.125 for cache writes and $0.50 for output on short prompts. Claude Haiku 5.5 charges exactly the same four prices below 100,000 tokens. Above OpenAI’s long-context line, Luna rises to $0.20 input and $0.75 output, so long prompts stay cheaper on Luna.

GPT-6.1 Sol on cached input

GPT-6.1 Sol, released on 29 September, kept the $2 and $10 prices of GPT-6 Sol but halved cached input to $0.10. Sonnet 5.5 already matched GPT-6 Sol on input, output and cache writes. With cache reads at $0.10, it now matches GPT-6.1 Sol on all four standard prices.

Cheaper rivals still exist

Matching OpenAI does not make Claude Haiku 5.5 the cheapest option. The New Stack notes that Alibaba’s Qwen3.7 Flash costs $0.03 input and $0.13 output for inputs of up to 32,000 tokens. Decision models such as Jev, which we looked at in September, compete for the same classification and routing work at lower prices. Our guide to LLM API pricing explains how to compare these lists.

What Businesses Should Do Next With Claude Haiku 5.5

The new prices are real, but the saving any one business sees depends on its prompts, its use of caching and the effort setting. Four steps will tell you where you stand.

Re-measure before you budget

Count a sample of real prompts with Claude Haiku 5.5 as the model, not Haiku 4.5. Check how many cross 100,000 tokens. If many sit between 77,000 and 100,000 on the old count, test whether they now fall into the higher tier.

Route work by task

Use Claude Haiku 5.5 for classification, extraction, summaries, compaction and subagent lookups. Keep long, complex coding tasks on Sonnet 5.5 or Opus 5.5, as Anthropic itself advises. A mixed line-up of Claude models usually beats one model for everything.

Set the effort level on purpose

Start at medium, the default, and only raise it where tests show a clear gain in accuracy. In Artificial Analysis’s tests, the maximum setting used about three times GPT-6 Luna’s tokens, and moving up from “xhigh” added just two points.

Check your cache hit rate on Sonnet 5.5

The cache cut only helps if you cache. Look at the share of your Sonnet 5.5 bill that is cache reads. If it is low, restructure prompts so that the fixed parts, such as system prompts, tools and reference documents, come first and stay the same between calls.

Claude Haiku 5.5 FAQs

When was Claude Haiku 5.5 released?

Anthropic released Claude Haiku 5.5 on Wednesday 7 October 2026. It is available on the Claude Platform, Amazon Web Services, Google Cloud and Microsoft Azure.

How much does Claude Haiku 5.5 cost?

For prompts of up to 100,000 tokens, it costs $0.10 per million input tokens and $0.50 per million output tokens. For longer prompts the rates are $0.50 and $2.50. Cache reads cost $0.01 and $0.05 respectively.

Is Claude Haiku 5.5 really 75% cheaper?

Anthropic says it costs about 75% less to run on average. Per-token prices fall by 90% below 100,000 tokens and 50% above it, but the new tokenizer counts about 30% more tokens for the same text. Your saving depends on your prompts.

What changed for Sonnet 5.5?

Only the cache read price. It fell from $0.20 to $0.10 per million tokens on 7 October. Input, output and cache write prices are unchanged. Anthropic says most agentic tasks on Sonnet 5.5 now cost about 20% less.

Who gets the monthly API credits?

Max 5x subscribers get $100 a month, Max 20x subscribers get $200 and Team plans get up to $500, pooled across seats. Free, Pro and Enterprise plans are not eligible, and the credits do not cover interactive Claude Code.

Can I swap Haiku 4.5 for Claude Haiku 5.5 without code changes?

Usually not. Thinking budgets, non-default sampling settings and assistant prefill all return errors on the new model. You also need to handle safety refusals and recount your tokens.

References