Back to feed

Haiku 5.5: Anthropic finally fixes its small-model problem

Theo's review finds Haiku 5.5 doubling its predecessor while matching Luna's short-context price; computer use and agentic coding tilt toward Anthropic.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — 38_6C0dkKmU
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

Anthropic owns the top of the coding-model ladder. From Fable and Opus through the Sonnet 3.5 era, its tool calls turned AI-assisted development into the daily workflow software engineers now take for granted. But the lineup always had a weak link: the small models . Haiku 4.5, launched in October last year, was pricey at the time and looks laughable today. Theo went so far as to teach every tool, agent, and rules file in his setup to avoid Haiku 4.5 at all costs. A few weeks ago he wrote it plainly: Anthropic has no small model worth using.

The critique is historical. Anthropic pours nearly everything into the flagship and distills it poorly into the cheaper tiers, so Opus quality sagged through the mid-generation releases and the dull Opus 5. That curse broke with Opus 5.5 and Sonnet 5.5, and expectations moved to today's Haiku 5.5. The small model's job is clear: file search, triple-checking diffs, everything an expensive model should not burn tokens on. Theo had lost so much trust that he routed his Claude setup to a rival Sol model to save money. The question now: did Haiku 5.5 end the embarrassment?

Price and the 100k line

Haiku 5.5 answers clearly: the price is right, the performance looks solid, and there are a few tricks that flatter so small a model . Up to 100,000 tokens, input runs $0.10 per million and output $0.50; above the line both rise fivefold. Roughly 90% of legacy Haiku requests stayed under the threshold. List prices fall 90% for short prompts and 50% for long ones, averaging about 75% cheaper once the new tokenizer's appetite is factored in. This price card comes from the Anthropic announcement, which frames the release for high-volume, cost-sensitive work.

The threshold is doing two jobs at once. On one side it makes quick reads and classification absurdly cheap, taking on text-editing APIs; on the other it steers Claude Code usage toward short responses so long-context bills do not bankrupt the lab. The calendar reads: Opus 5.5 on September 22, Sonnet 5.5 on September 28, now Haiku 5.5 completing the 5.5 family. Staffer Lydia publicly suggests pinning the auto-compact window to 100k, a per- model setting that covers subagents too. This launch sequence is confirmed in the SiliconAngle write-up.

The comparison that matters is Luna. On short context Haiku 5.5 matches OpenAI's Luna to the decimal ($0.10/$0.50); past 100k the tables turn, since Luna's step-up only hits at 272k ($0.20/$0.75). Add the new tokenizer, which turns the same prose into roughly 25% more tokens, and headline price and real invoice quietly diverge. Fit inside 100k and Haiku leads on price-performance; spill over and Luna is the better deal. SimonWillison stresses exactly this split: parity below the line, Luna far cheaper above it.

Arithmetic at scale sharpens the point. On a sample turn reading 80,000 cached tokens plus 2,000 fresh input and 1,000 output tokens, Haiku and Luna cost about $0.0015 against Sonnet's $0.022; per thousand turns, $1.50 versus $22. Where verification is cheap, retry logic wins: Terminal-Bench pass rates of 39.2% for Haiku, 16.4% for Luna, and 70.6% for Sonnet imply roughly $0.004, $0.009, and $0.031 per verified success. Same sticker price, different hit rate. This math comes from the CTOL analysis, which recommends routing: the cheap model attempts, a verifier judges, failures escalate.

Benchmarks and the intelligence-per-dollar curve

The benchmark jump is violent. Knowledge-work GDPval more than doubles from 735 to 1620; Humanity's Last Exam climbs from 10.2% to 45.9% without tools and 18.7% to 57.4% with them; computer-use OSWorld leaps from 15.7% to 72.4%; agentic-coding Terminal-Bench goes from zero to 39.2%. Luna sits at 16.4% on the coding test, less than half of Haiku. Visual reasoning jumps from 6.4% to 46.4%. These figures are Anthropic 's published table; an independent index also ranks the newcomer first among small models at 43 points, while flagging its token hunger.

OpenAI led computer use for a long stretch; the gap is closing, though complex agentic coding still belongs to Sonnet and Opus 5.5. Haiku 5.5 is the first small model with an adjustable effort dial, from low to max, letting the workload set the intelligence-cost tradeoff. Context grows from 200,000 to one million tokens. Live support, voice agents, in-app assistants, and browser work all want that speed. This picture is summarized in the Technology roundup as well.

The funniest segment is the intelligence-per-dollar chart. The Pareto frontier runs through Luna, Haiku, Sonnet, and Opus until 6.1 Sol bends it, offering a giant intelligence leap for a tenth of a cent more per task. Luna is always cheaper and, on these numbers, always weaker. The quiet gem is open-weight Mimo, surprisingly strong for its price and one of few stragglers above the line. For teams staying on one provider, Haiku is now a sensible stop: no second subscription, models that understand each other's prompts.

The single-vendor case is practical. A team standardized on Anthropic inference runs everything under one roof, one bill, one support desk, without doubling subscriptions. The claim that Opus prompting Haiku beats Opus prompting Sol is unmeasured but felt in practice. Cost-target teams should judge by cost per verified success, not sticker price, and keep the Luna alternative on the table.

Demos and design taste

The subagent architecture is shifting. Theo plans to relax the route-everything-to-Opus rule in his rules file, letting Opus decide which narrow subset Haiku may handle cheaply. He has not stress-tested the idea yet but will rewrite the rule. The same direction appears in the AICoder record: Claude Code v2.1.293 makes Haiku 5.5 the default small model with a one-million-token window, aimed at summaries, compaction, queries, and classification.

Then come the demos. A community video by Arshan, generated for about sixty cents in API spend in under twenty minutes, looks insanely good for the money. Theo credits part of the secret to taste baked into Anthropic models : carefully curated design data, aggressively trimmed, giving the model strong reference points. Dara's whichAI front-end study carries the same polish.

The whichAI comparison buries Grok: better type choices, animation tricks never seen in these demos, a blueprint panel with hover behaviors that reshape content below. Kill the design scale and output collapses into classic language- model ugliness, which reveals where the gap comes from: the design plugin. Theo had uninstalled it and now plans to bring it back. With the plugin off, Grok looks even worse, and the segment dissolves into laughter.

The dark spot is a live rate-limit lesson. A script fetching 181 pull requests in parallel blows through the 5,000-requests-per-hour GitHub quota, saves the error body as if it were data, then chokes parsing JSON. Nothing re-triggers the wait, so the thread sits dead until Theo resumes it, locking him out of integrations for nearly an hour. Small models corner themselves and make extra work for the user: a live exhibit of the host's distrust.

Field evidence and verdict

Field reports are strong. Asana's evaluation notes task completion over 30% faster with up to 2.5x quicker inference per agent turn; engineer Aaron Vinh calls it noticeably snappier. HubSpot's CRM suite gives its best score yet at 92.8, fastest completion with the highest hit rate and lowest false positives. AlphaSense, running eight million calls a week, records a significant jump from 0.76 to 0.84 across 400 queries. These customer results are published in the Anthropic announcement.

Box completes the picture with an eleven-point lead at about half the latency. Safety is tightened: alignment testing shows far fewer deviations than Haiku 4.5, defensive work gets more room than Sonnet 5.5 allows, penetration testing stays blocked, and wider access runs through the Cyber Verification Program. This safety frame is confirmed in the SiliconAngle write-up, which also lists distribution via the Claude Platform, AWS, Google Cloud, and Azure.

Launch day brings two gifts. Sonnet 5.5 cache reads halve from $0.20 to $0.10 per million, cutting most agentic bills by about a fifth. Monthly API credits land for subscribers: $100 on Max 5x, $200 on Max 20x, up to $500 pooled for Team; no rollover, no interactive Claude Code. The context is told in the TechCrunch record too: Sonnet 5.5 from September 28 runs about 30% faster, beats Opus at multi-agent work, and carries Fable-grade cyber safeguards. The company had flagged the new small model weeks ahead.

The verdict: Haiku 5.5 is the model for narrow, verifiable, repetitive work, from summaries and compaction to classification, queries, and subagent duty. The smart model plans, the cheap one attempts, the verifier judges. Cost-per-success math agrees: cheap tries plus selective escalation. Heavy lifting stays with the flagships. The small- model curse looks broken for the first time, and Theo is off to rewrite his rules file.

Visualization: nodesdaily AI
MetricResult
OSWorld 72.4%Was 15.7%
Terminal 39.2%Was 0.0%
Short input price$0.10 per million

Key moments

  1. Opening case against tiny Anthropic options
  2. Price card and the 100k threshold
  3. Benchmark surge versus Luna
  4. Intelligence-per-dollar frontier chart
  5. Subagent layout and rules file
  6. Community demos and design taste
  7. Quota lesson and stalled thread
  8. Verdict and routing advice

AI commentary

"The small-model race just changed: price is no longer the excuse, measurement is. Still, routing matters; throwing everything at the cheap tier grows disappointment, not savings."

AI assessment

Strongest counterargument is token appetite: at max effort the newcomer burns about 162,000 output tokens per task where Luna needs 50,000 for a similar score. Equal sticker price, unequal invoice. Past 100k Luna's own step-up is gentler, so long context favors Luna. Parity pricing is not parity capability, and retries cannot solve tasks the small tier fundamentally cannot.

Gaps remain: complex agentic coding still belongs to Sonnet and Opus, and the vendor says so openly. Hallucination runs lower than Luna (40% versus 77%) with a healthier habit of admitting ignorance, yet factual knowledge trails (36% against 44% on the omniscience probe). Cyber permissions are wider on defense than Sonnet's, closed on offense; reasonable balance, but a hard wall for red teams.

Possible interest deserves a note: the video carries a Coderabbit sponsorship, and small- model praise flatters a sponsored coding tool. Community demos are a selected showcase; failed attempts never air. Even so, the benchmark table is official, customer names are public, and the figures line up with independent sources.

Practical takeaway: give Haiku work that fits under 100k with cheap verification, pin the auto-compact window to 100k, watch metered billing. Keep the flagship for long context and heavy lifting. One vendor simplifies the invoice, but the Luna alternative should stay on the table.

Sources

8 links; 5 of them also cited by 4 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

artificial intelligence · claude · anthropic · llm · agents

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review source passages, versions and origins.

READ WITH SOURCES

Understand this story.

Checking your account…