Anthropic unveiled Claude Haiku 5.5 on 7 October 2026, the new version of its small model, and sold it with three adjectives: the cheapest, fastest and most capable of that tier so far. The figure that sums the announcement up is the price: according to the company, running a workload on Haiku 5.5 costs on average 75% less than on Haiku 4.5. The model is already available across all platforms — including Amazon Web Services, Google Cloud and Microsoft Azure — under the identifier claude-haiku-5-5.
What was announced
Pricing is tiered by prompt size. For requests of up to one hundred thousand tokens, a million input tokens costs ten US cents and a million output tokens fifty cents; cache writes come to twelve and a half cents and cache reads to one cent. Above that threshold everything is multiplied by five: fifty cents for input, two dollars and fifty cents for output, sixty-two and a half cents for cache writes and five cents for cache reads. The comparison is with Haiku 4.5, which charged one dollar per million input tokens, five dollars per million output tokens, one dollar twenty-five for cache writes and ten cents for cache reads.
That is where the 75% average comes from. In the small print of the announcement, the actual cut is 90% for requests below one hundred thousand tokens and 50% above it; on Haiku 4.5, nine out of ten requests fell into the first group. The calculation also accounts for Haiku 5.5 shipping a new tokeniser, similar to those of Sonnet 5.5 and Opus 5.5, which burns slightly more tokens for the same piece of work.
It is also the first model in the small family with an adjustable effort setting, the same control the larger models already had to choose between optimising for cost or for intelligence.
These are the evaluation figures published by Anthropic, with the models reordered for readability:
| Benchmark | Haiku 4.5 | Haiku 5.5 | Sonnet 5.5 | GPT-6 Luna |
|---|---|---|---|---|
| GDPval-AA v2.1 (professional work) | 735 | 1620 | 1840 | 1437 |
| AA-Briefcase v1.1 | 614 | 1578 | 1824 | 1336 |
| OSWorld 2.1 (offline subset) | 15.7% | 72.4% | 83.9% | 48.9% |
| Humanity's Last Exam (no tools) | 10.2% | 45.9% | 56.9% | — |
| Terminal-Bench 4.0 | 0.0% | 39.2% | 70.6% | 16.4% |
| FrontierCode 1.1 (Main) | — | 46.4% | 52.1% | 42.4% |
| Chartography (no tools) | 6.4% | 46.4% | 61.6% | 29.1% |
The Sonnet 5.5 figure on FrontierCode corresponds to its highest effort configuration. With tools, Haiku 5.5 rises to 57.4% on Humanity's Last Exam against 18.7% for Haiku 4.5.
Two further changes come with the launch. Cache reads on Sonnet 5.5 drop by half, from twenty to ten US cents per million tokens, which the company says makes typical agentic work with that model around 20% cheaper. And this week a monthly API credit starts rolling out to Max and Team subscribers: one hundred dollars a month for Max 5x, two hundred for Max 20x and up to five hundred for teams, pooled across their users and usable with any model. The Python and TypeScript SDKs add beta support for computer use and browser use.
Where it comes from
The move fits the direction the company has been taking: first the cut to cache pricing already covered on this blog, then the tariff reductions that came with this generation's large models, and now the cheap end of the range. Anthropic itself frames the launch as a review of the value of its whole model family rather than an isolated product.
The jump over Haiku 4.5 is biggest precisely where the older model was useless: on Terminal-Bench 4.0 it scored a flat zero. On computer use it goes from 15.7% to 72.4% of the OSWorld 2.1 offline subset, according to the company's own testing.
What it means
The practical change is not that there is a new model, it is that tasks which did not pay off now do. Compacting context, summarising, classifying, firing off database queries, spawning subagents that fetch one specific figure from a long document: all of that used to be billed at mid-tier prices and now costs a tenth of that in the short tier. The hundred-thousand-token threshold is the boundary worth watching when designing prompts, because crossing it multiplies the per-token bill by five.
Anthropic draws the line itself: for complex agentic coding — the kind Terminal-Bench measures — Sonnet 5.5 and Opus 5.5 remain the better choices. Haiku 5.5 is positioned as a supporting part, not a replacement. The pattern its customers describe is exactly that: a big model handles the main task and a small one runs the errands.
For anyone already using Sonnet 5.5 in agentic flows, the cache cut requires no changes at all: the saving simply shows up. For anyone coming from Haiku 4.5, Anthropic publishes a migration guide; cost estimates are worth redoing with the new tokeniser, which consumes a little more.
On safety, the company claims improvements in almost all of its alignment evaluations compared with Haiku 4.5, with fewer misaligned behaviours and less willingness to cooperate with misuse. Cybersecurity safeguards are tighter than Haiku 4.5's and somewhat looser than those of other recent models: they allow a wider range of defensive tasks than Sonnet 5.5's, but still block penetration testing. Biology safeguards match those of Sonnet 5, Sonnet 5.5 and Opus 5.
What the announcement does not say
All performance figures come from Anthropic and its own methodology, detailed in the model's system card. There is no independent verification.
The customer testimonials are equally interested claims. Asana reports a latency reduction of over 30% and up to 2.5 times faster inference per agent turn against the model it uses today, without naming it. HubSpot cites the best result its CRM task suite has recorded, 92.8% averaged over three runs. AlphaSense, which says one of its features handles around eight million calls a week, reports an improvement from 0.76 to 0.84 over four hundred queries. Box puts the advantage over Haiku 4.5 at eleven points with roughly half the latency. Cognition places Haiku 5.5 as a sidekick in Devin Fusion with a FrontierCode score of 66.2, led by Opus 5.5.
Nor is it specified how long the monthly credit for Max and Team will last, or whether it is permanent. And the claim of being the house's fastest model carries its own asterisk in the text: it is, at standard speed, but the Opus models in fast mode are quicker.
"Haiku 5.5 is best suited to more narrowly scoped tasks", the company concedes when comparing its small model with the large ones on agentic coding.
