Back to feed

Haiku 5.5 Redraws Cost Efficiency for Small Models

Anthropic launched Haiku 5.5 at 10 cents input per million tokens up to 100,000 tokens, scoring 72.4 on OSWorld 2.1 and beating rival Luna.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — PVe-ibLRLaw
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

Anthropic introduced Haiku 5.5 on October 7, 2026, and the price list is the first thing that catches the eye. Jobs of up to 100,000 tokens cost 10 cents per million input tokens and 50 cents per million output tokens. Past that threshold the tariff jumps fivefold, rising to 50 cents and $2.50. The charts in the video therefore tell two stories: on token efficiency Haiku looks smarter than its rival GPT-6 Luna, while on cost efficiency Luna pulls ahead. A concrete example makes it vivid: the same job costs 21 cents through Haiku and 7 cents through Luna. The premium buys higher intelligence and slightly faster generation. This price table appears exactly as published in the official announcement by Anthropic, confirmed by the charts the company shared.

Everyone asks why such a cheap and fast small model is even needed, and the answer hides in high-volume work. Repetitive duties like summaries, compactions, classification requests and database queries should run without tiring the flagship models. Haiku 5.5 steps in here as a subagent that supports its bigger siblings Opus 5.5 and Sonnet 5.5. Speed-sensitive duties such as live customer support and browser use sit on the target list too. The host Caleb shows a living example: Atticus, a personal agent running in his home. This positioning is also stressed in the launch analysis published by SiliconAngle, whose examples of the target audience match the video.

Atticus: A Personal Agent for Daily Flow

Atticus follows spoken instructions and performs two duties in the demo. While Caleb reads a research paper, he simply speaks up, and Atticus gathers related papers in the background without breaking his flow. Then it looks at the desktop, reads the work open in the VS Code window, finds the PyTorch documentation and pins it to the left of the screen. Caleb keeps writing while the reference text waits beside him. Each interaction averages $0.001, so running a personal agent costs next to nothing. That math moved Caleb from local hardware to the cloud: he used to run a Qwen build over his home network and now calls the Haiku line through the API. His reason is the Michigan electricity bill at $0.23 per kWh. The video carries a Micro Center sponsorship, with the company handing out 128 GB flash drives at its Austin store opening. The local-to-cloud math lines up with measurements shared by SimonWillison, who finds small jobs finishing for fractions of a cent.

The most debated number sits in the computer use score. Haiku 5.5 took 72.4 points on the offline subset of OSWorld 2.1, ahead of Luna at 48.9 and far above the older Haiku 4.5 at 15.7. Even the bigger sibling Sonnet 5.5 stands at 83.9, so the small model has closed the gap remarkably. The other rows read strongly too: 1620 points on the GDPval-AA v2.1 knowledge-work test beats Luna at 1437, and 39.2 points on the Terminal-Bench 4.0 agentic coding test more than doubles Luna at 16.4. On the FrontierCode main test, 46.4 points edges past Luna at 42.4. Read together with the price and speed tables in the ArtificialAnalysis directory, these rows clarify the balance of power among small models.

What the OSWorld Test Measures

The OSWorld test checks whether a model can operate a real computer with mouse and keyboard. Virtual environments are staged, and the model is asked to finish long-horizon duties such as formatting a slideshow or booking a trip. The environment feeds fresh screenshots at every step, the model reads each image and plans the next move. The loop repeats until the job is done; the sample slideshow takes 62 steps, and the list holds more than a hundred tasks. The test design is documented in detail in the OSWorld 2.x papers released by the XLang community, where the task list can be verified.

The home trials show both the bright and the rough side of that loop. In the success story Caleb speaks, Haiku finds the Excel application and switches it to dark mode, clicking through menus as screenshots travel back and forth until the job lands. In the rough story the host asks for filtering on the ArtificialAnalysis website: models above the $1 line should be switched off. The model first unchecks the Opus 5.5 and GPT-6 rows, then quits halfway. After a follow-up instruction it moves a little further yet never reaches the goal, and the whole attempt takes 33 seconds. Despite 173 tokens per second, the host admits there is road ahead in both speed and reasoning, adding that his 4K monitor may have slowed things down. These breaking moments prove that spoken instruction plus computer control is not yet a finished craft.

Speed Alone Falls Short

The closing argument is crisp: cheaper, smarter and faster models unlock new downstream uses. Anthropic reports average running costs down 75 points, with a 90-point discount on jobs that fit inside the 100,000-token limit. A new effort setting lets users pick inference intensity from low to max; independent trials drew a pelican sketch for 0.09 cents at the lowest level and 3.38 cents at the highest. Yet the new tokenizer splits the same text into 1.25 times more tokens, hiding a price bump. Above the threshold Luna looks like the better deal, since the rival rises only to 20 and 75 cents past its 272,000-token mark. These details are backed by trial notes shared by SimonWillison, whose math matches the video.

Visualization: nodesdaily AI
MetricValue
Input up to 100,000 tokens10 cents per million tokens
OSWorld 2.1 score72.4 points, ahead of Luna
Cost per interaction$0.001 through Atticus

Key moments

  1. Two-level price table
  2. Why a small model matters
  3. Sponsor message
  4. Computer use score
  5. Personal agent home trial
  6. Closing verdict

AI commentary

"Small models keep getting smarter while their bills shrink, opening the road for personal agents. The above-limit price jump and unfinished tasks form the cautious side of the story."

AI assessment

The strongest counterargument comes from long-context pricing. Past the 100,000-token line Haiku turns five times dearer while Luna holds the 10-cent level up to 272,000 tokens, staying at the 20 and 75-cent line even beyond. Add the 1.25-fold tokenizer inflation and the table flips for long documents. Moreover the 72.4-point score belongs to the offline subset; nothing yet proves the same success across end-to-end workflows that take humans 1.6 hours. The gap between the tidy virtual rig and messy real websites calls for a cautious reading of the headline number.

The gaps list runs longer. Nothing concrete is shared on safety, data retention or enterprise auditing, while a personal agent that watches home screens leaves the privacy question open. The balance between local privacy and speed on one side and cloud cost on the other shifts per user, with no single formula. The sponsored hardware narrative muddies the picture further, placing costly accelerator rigs next to penny-priced cloud calls in the same video. Readers should therefore measure their own workload before deciding anything.

The practical takeaway compresses into three steps. First split the workload at the 100,000-token line: Haiku below it, the rival or a bigger sibling above it. Then experiment with the effort setting , drafting fast at the low level and reasoning carefully at the high one. Finally consider the subagent pattern, where the flagship model plans and the small model grinds through repetitive steps. Figures like $0.001 per interaction sound delightful, yet the monthly total deserves watching, because high volume fattens tiny fees quickly.

Sources

6 links; 2 of them also cited by 3 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

haiku 5.5 · anthropic · osworld · ai agent · cost efficiency

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review source passages, versions and origins.

READ WITH SOURCES

Understand this story.

Checking your account…