Back to feed

Claude Haiku 5.5: Anthropic's Cheapest and Fastest Model Reshapes the Small-Model Race

Anthropic released Claude Haiku 5.5 on October 7: with input pricing at $0.10 per million tokens, an adjustable effort dial, and benchmark scores ahead of GPT-6 Luna, the economics of small models are shifting.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — 6gfACqBAKfw
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

Anthropic is living through the most consequential weeks in its history: according to QZ.com, the company plans to start its investor roadshow in the week of November 9 and list before Thanksgiving at a valuation between $1.8 and $2 trillion. Claude Haiku 5.5 landed right in the middle of that picture, released on Wednesday, October 7. For Anthropic , the timing is no accident: in an intensifying race with OpenAI, the company wants its smallest and fastest model to capture developers' everyday workloads.

The launch completes the 5.5 family trilogy: Opus 5.5 arrived on September 22, Sonnet 5.5 on September 28. As Decrypt summarized it, Haiku 5.5 is positioned as the cheapest and fastest model the company has ever shipped. The presenter makes no secret of his surprise that the refresh finally came: Haiku 4.5 had gone almost a year without an update and was lagging behind its rivals.

Pricing: a steep cut with a 100k-token threshold

The pricing is aggressive: $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. According to Anthropic's announcement, that list sits 90% below Haiku 4.5's $1 and $5 rates; above 100,000 tokens the discount narrows to 50% ($0.50 and $2.50). The company's blended figure puts running costs about 75% lower , since roughly 90% of the old model's requests stayed under the threshold. SimonWillison flags a subtle catch: the new tokenizer splits the same text into roughly 1.25x more tokens, so the real-world discount is softer than the list suggests.

The price lands exactly on top of OpenAI's small model GPT-6 Luna, released September 22 at the same $0.10/$0.50 rates. Under 100,000 tokens the two models cost the same; above the line the picture flips, because Luna holds its $0.10 and $0.50 rates up to 272,000 tokens and only then steps up to $0.20 and $0.75. The benchmark rows reported by Decrypt justify the money: Haiku 5.5 leads Luna on all six published tests. SimonWillison's conclusion is crisp: Haiku wins on short context, Luna on long context.

Benchmarks: a big leap for a small model

The table in the official announcement shows the scale of the jump: on OSWorld 2.1, which measures real computer use, Haiku 5.5 scores 72.4% while Haiku 4.5 sits at 15.7% and Luna at 48.9%. The agentic coding test Terminal-Bench 4.0 is starker still: 39.2% versus 0.0% , with Luna at 16.4%. Knowledge-work scores read 1620 versus 735 on GDPval-AA (Luna 1437), 46.4% versus 42.4% on FrontierCode, and 46.4% versus 6.4% on the visual-reasoning test Chartography (Luna 29.1%). The full compilation by Technology Org shows the one strong rival left out of the comparison, Sonnet 5.5 (70.6% on Terminal-Bench, for instance), still playing a league above.

Numbers from early-access customers back up the table: according to SiliconAngle, Asana saw task-completion latency in its agent product drop by more than 30% , with inference running up to 2.5x faster per agent turn. HubSpot recorded 92.8 on its CRM test suite, the best score it has seen on that battery. The Box report is more granular: 60 versus 49 on complex-work evaluation, about 27% fewer tokens, and roughly half the wall-clock time, with the gap widening to 19 points on report-drafting work that demands numerical derivation.

Adjustable effort and where to get it

Haiku 5.5 is the first Haiku model with an adjustable effort setting : developers pick low, medium, high, or maximum reasoning for the job at hand; reasoning cannot be switched off entirely, and the default is medium. Anthropic's published curves show accuracy climbing with cost as effort rises. SimonWillison's playful experiment makes the trade concrete: a pelican sketch at low effort took 7 seconds and cost 0.09 cents, while the maximum-effort version took 5 minutes 9 seconds and cost 3.38 cents.

Availability is broad: the model is live in the Claude chatbot, inside Claude Code, through the API, and on Amazon, Google, and Microsoft clouds, with the US government cloud on the list too. SiliconAngle notes two further discounts: Sonnet 5.5's cache-read price is halved from $0.20 to $0.10, taking roughly 20% off most agentic work on that model. Monthly API credits sweeten subscriptions as well: $100 for Max 5x, $200 for Max 20x, and up to $500 pooled for Team plans. Decrypt points out the credits match the subscription fee and do not roll over.

Hands-on tests: from games to design

The presenter's own tests go beyond the numbers: the plain Mac clone Haiku 4.5 produced 11 months ago sits next to the Minecraft clone Haiku 5.5 generated, complete with cave systems and water physics. The most extreme trial yields a Call of Duty-style zombie game skeleton; an isometric room drawing earns praise for texture and lighting. The standout community example comes from a user named Theo, who built a dungeon crawler with shifting atmosphere in a single hour. Motion design impresses most: an animation that takes frontier models 40 minutes comes out of Haiku 5.5 in about 15 minutes .

The presenter's verdict sits squarely on price-performance: the model costs as little as Luna yet clearly outscores it, and it finishes the job faster in head-to-head runs against Grok. Near-Sonnet behavior on simple tasks makes it a natural subagent inside Claude Code: the large model draws the plan while the small model takes the repetitive subtasks. The daily-driver claim makes sense in that architecture — not a model without limits, but a cheap teammate that keeps working.

Visualization: nodesdaily AI
FeatureValue
Price$0.10 input and $0.50 output
OSWorld 2.172.4%, ahead of Luna
Effort settingFirst adjustable Haiku

Key moments

  1. Opening: Anthropic on the IPO road
  2. Haiku 4.5 versus 5.5
  3. Price and the 100k threshold
  4. Luna comparison
  5. Benchmark results
  6. Effort dial from low to maximum
  7. Theo's one-hour game

AI commentary

"The small-model contest is no longer decided on price alone; accuracy now sits at the same table. Haiku 5.5 looks like the first cheap model that needs no excuses, though nobody should commit before seeing the long-context bill."

AI assessment

The strongest objection stands in speed's shadow: Decrypt's team notes the model answered a simple logic question almost instantly — and wrongly. Blind trust gets expensive, especially in customer-facing work. The presenter's enthusiasm deserves a discount too: the tests run on hand-picked tasks at hand-picked settings, and not every workload will reproduce the same picture.

The limits are clearly drawn: above 100,000 tokens the price edge passes to Luna (SimonWillison), Anthropic itself points complex agentic coding toward Sonnet and Opus (SiliconAngle), and penetration testing plus attacker-leaning techniques stay blocked. So Haiku 5.5 is not the model for every job; it is the model for narrow, repetitive, speed-sensitive jobs.

The company's motive deserves a look as well: the $1.8 to $2 trillion listing target reported by QZ.com sits next to 2025 revenue of $4.6 billion and a net loss near $42 billion (mostly accounting adjustments). Slashing prices this hard reads as a move to grow the developer base before the IPO. Anthropic's customer quotes are also company-selected examples; without independent verification, it is early to take the table at face value.

The practical takeaway for readers is crisp: teams with high-volume summarization, classification, database queries, or live support can start with a small trial and measure the cost directly. For agent architectures the recipe is set: the big model plans, Haiku repeats. The monthly API credits on Max and Team plans make the trial ticket cheaper still.

Sources

8 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

claude haiku · anthropic · artificial intelligence · llm · api pricing

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review source passages, versions and origins.

READ WITH SOURCES

Understand this story.

Checking your account…