Anthropic is living through the most consequential weeks in its history: according to QZ.com, the company plans to start its investor roadshow in the week of November 9 and list before Thanksgiving at a valuation between $1.8 and $2 trillion. Claude Haiku 5.5 landed right in the middle of that picture, released on Wednesday, October 7. For Anthropic , the timing is no accident: in an intensifying race with OpenAI, the company wants its smallest and fastest model to capture developers' everyday workloads.
The launch completes the 5.5 family trilogy: Opus 5.5 arrived on September 22, Sonnet 5.5 on September 28. As Decrypt summarized it, Haiku 5.5 is positioned as the cheapest and fastest model the company has ever shipped. The presenter makes no secret of his surprise that the refresh finally came: Haiku 4.5 had gone almost a year without an update and was lagging behind its rivals.
Pricing: a steep cut with a 100k-token threshold
The pricing is aggressive: $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. According to Anthropic's announcement, that list sits 90% below Haiku 4.5's $1 and $5 rates; above 100,000 tokens the discount narrows to 50% ($0.50 and $2.50). The company's blended figure puts running costs about 75% lower , since roughly 90% of the old model's requests stayed under the threshold. SimonWillison flags a subtle catch: the new tokenizer splits the same text into roughly 1.25x more tokens, so the real-world discount is softer than the list suggests.
The price lands exactly on top of OpenAI's small model GPT-6 Luna, released September 22 at the same $0.10/$0.50 rates. Under 100,000 tokens the two models cost the same; above the line the picture flips, because Luna holds its $0.10 and $0.50 rates up to 272,000 tokens and only then steps up to $0.20 and $0.75. The benchmark rows reported by Decrypt justify the money: Haiku 5.5 leads Luna on all six published tests. SimonWillison's conclusion is crisp: Haiku wins on short context, Luna on long context.
Benchmarks: a big leap for a small model
The table in the official announcement shows the scale of the jump: on OSWorld 2.1, which measures real computer use, Haiku 5.5 scores 72.4% while Haiku 4.5 sits at 15.7% and Luna at 48.9%. The agentic coding test Terminal-Bench 4.0 is starker still: 39.2% versus 0.0% , with Luna at 16.4%. Knowledge-work scores read 1620 versus 735 on GDPval-AA (Luna 1437), 46.4% versus 42.4% on FrontierCode, and 46.4% versus 6.4% on the visual-reasoning test Chartography (Luna 29.1%). The full compilation by Technology Org shows the one strong rival left out of the comparison, Sonnet 5.5 (70.6% on Terminal-Bench, for instance), still playing a league above.
Numbers from early-access customers back up the table: according to SiliconAngle, Asana saw task-completion latency in its agent product drop by more than 30% , with inference running up to 2.5x faster per agent turn. HubSpot recorded 92.8 on its CRM test suite, the best score it has seen on that battery. The Box report is more granular: 60 versus 49 on complex-work evaluation, about 27% fewer tokens, and roughly half the wall-clock time, with the gap widening to 19 points on report-drafting work that demands numerical derivation.
Adjustable effort and where to get it
Haiku 5.5 is the first Haiku model with an adjustable effort setting : developers pick low, medium, high, or maximum reasoning for the job at hand; reasoning cannot be switched off entirely, and the default is medium. Anthropic's published curves show accuracy climbing with cost as effort rises. SimonWillison's playful experiment makes the trade concrete: a pelican sketch at low effort took 7 seconds and cost 0.09 cents, while the maximum-effort version took 5 minutes 9 seconds and cost 3.38 cents.
Availability is broad: the model is live in the Claude chatbot, inside Claude Code, through the API, and on Amazon, Google, and Microsoft clouds, with the US government cloud on the list too. SiliconAngle notes two further discounts: Sonnet 5.5's cache-read price is halved from $0.20 to $0.10, taking roughly 20% off most agentic work on that model. Monthly API credits sweeten subscriptions as well: $100 for Max 5x, $200 for Max 20x, and up to $500 pooled for Team plans. Decrypt points out the credits match the subscription fee and do not roll over.
Hands-on tests: from games to design
The presenter's own tests go beyond the numbers: the plain Mac clone Haiku 4.5 produced 11 months ago sits next to the Minecraft clone Haiku 5.5 generated, complete with cave systems and water physics. The most extreme trial yields a Call of Duty-style zombie game skeleton; an isometric room drawing earns praise for texture and lighting. The standout community example comes from a user named Theo, who built a dungeon crawler with shifting atmosphere in a single hour. Motion design impresses most: an animation that takes frontier models 40 minutes comes out of Haiku 5.5 in about 15 minutes .
The presenter's verdict sits squarely on price-performance: the model costs as little as Luna yet clearly outscores it, and it finishes the job faster in head-to-head runs against Grok. Near-Sonnet behavior on simple tasks makes it a natural subagent inside Claude Code: the large model draws the plan while the small model takes the repetitive subtasks. The daily-driver claim makes sense in that architecture — not a model without limits, but a cheap teammate that keeps working.
| Feature | Value |
|---|---|
| Price | $0.10 input and $0.50 output |
| OSWorld 2.1 | 72.4%, ahead of Luna |
| Effort setting | First adjustable Haiku |
Key moments
AI commentary
"The small-model contest is no longer decided on price alone; accuracy now sits at the same table. Haiku 5.5 looks like the first cheap model that needs no excuses, though nobody should commit before seeing the long-context bill."
AI assessment
The strongest objection stands in speed's shadow: Decrypt's team notes the model answered a simple logic question almost instantly — and wrongly. Blind trust gets expensive, especially in customer-facing work. The presenter's enthusiasm deserves a discount too: the tests run on hand-picked tasks at hand-picked settings, and not every workload will reproduce the same picture.
The limits are clearly drawn: above 100,000 tokens the price edge passes to Luna (SimonWillison), Anthropic itself points complex agentic coding toward Sonnet and Opus (SiliconAngle), and penetration testing plus attacker-leaning techniques stay blocked. So Haiku 5.5 is not the model for every job; it is the model for narrow, repetitive, speed-sensitive jobs.
The company's motive deserves a look as well: the $1.8 to $2 trillion listing target reported by QZ.com sits next to 2025 revenue of $4.6 billion and a net loss near $42 billion (mostly accounting adjustments). Slashing prices this hard reads as a move to grow the developer base before the IPO. Anthropic's customer quotes are also company-selected examples; without independent verification, it is early to take the table at face value.
The practical takeaway for readers is crisp: teams with high-volume summarization, classification, database queries, or live support can start with a small trial and measure the cost directly. For agent architectures the recipe is set: the big model plans, Haiku repeats. The monthly API credits on Max and Team plans make the trial ticket cheaper still.
Sources
8 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — WorldofAI
- @anthropic.com Anthropic — Introducing Claude Haiku 5.5
- @simonwillison.net Simon Willison — Introducing Claude Haiku 5.5
- @decrypt.co Decrypt — Anthropic Launches Haiku 5.5
- @technology.org Technology Org — Claude Haiku 5.5 price and benchmarks
- @qz.com QZ — Anthropic pre-Thanksgiving IPO
- @siliconangle.com SiliconAngle — Haiku 5.5 and Sonnet cache cut
- @blog.box.com Box Blog — Haiku 5.5 accuracy and latency
claude haiku · anthropic · artificial intelligence · llm · api pricing