Anthropic introduced Haiku 5.5 on October 7, 2026, and the price list is the first thing that catches the eye. Jobs of up to 100,000 tokens cost 10 cents per million input tokens and 50 cents per million output tokens. Past that threshold the tariff jumps fivefold, rising to 50 cents and $2.50. The charts in the video therefore tell two stories: on token efficiency Haiku looks smarter than its rival GPT-6 Luna, while on cost efficiency Luna pulls ahead. A concrete example makes it vivid: the same job costs 21 cents through Haiku and 7 cents through Luna. The premium buys higher intelligence and slightly faster generation. This price table appears exactly as published in the official announcement by Anthropic, confirmed by the charts the company shared.
Everyone asks why such a cheap and fast small model is even needed, and the answer hides in high-volume work. Repetitive duties like summaries, compactions, classification requests and database queries should run without tiring the flagship models. Haiku 5.5 steps in here as a subagent that supports its bigger siblings Opus 5.5 and Sonnet 5.5. Speed-sensitive duties such as live customer support and browser use sit on the target list too. The host Caleb shows a living example: Atticus, a personal agent running in his home. This positioning is also stressed in the launch analysis published by SiliconAngle, whose examples of the target audience match the video.
Atticus: A Personal Agent for Daily Flow
Atticus follows spoken instructions and performs two duties in the demo. While Caleb reads a research paper, he simply speaks up, and Atticus gathers related papers in the background without breaking his flow. Then it looks at the desktop, reads the work open in the VS Code window, finds the PyTorch documentation and pins it to the left of the screen. Caleb keeps writing while the reference text waits beside him. Each interaction averages $0.001, so running a personal agent costs next to nothing. That math moved Caleb from local hardware to the cloud: he used to run a Qwen build over his home network and now calls the Haiku line through the API. His reason is the Michigan electricity bill at $0.23 per kWh. The video carries a Micro Center sponsorship, with the company handing out 128 GB flash drives at its Austin store opening. The local-to-cloud math lines up with measurements shared by SimonWillison, who finds small jobs finishing for fractions of a cent.
The most debated number sits in the computer use score. Haiku 5.5 took 72.4 points on the offline subset of OSWorld 2.1, ahead of Luna at 48.9 and far above the older Haiku 4.5 at 15.7. Even the bigger sibling Sonnet 5.5 stands at 83.9, so the small model has closed the gap remarkably. The other rows read strongly too: 1620 points on the GDPval-AA v2.1 knowledge-work test beats Luna at 1437, and 39.2 points on the Terminal-Bench 4.0 agentic coding test more than doubles Luna at 16.4. On the FrontierCode main test, 46.4 points edges past Luna at 42.4. Read together with the price and speed tables in the ArtificialAnalysis directory, these rows clarify the balance of power among small models.
What the OSWorld Test Measures
The OSWorld test checks whether a model can operate a real computer with mouse and keyboard. Virtual environments are staged, and the model is asked to finish long-horizon duties such as formatting a slideshow or booking a trip. The environment feeds fresh screenshots at every step, the model reads each image and plans the next move. The loop repeats until the job is done; the sample slideshow takes 62 steps, and the list holds more than a hundred tasks. The test design is documented in detail in the OSWorld 2.x papers released by the XLang community, where the task list can be verified.
The home trials show both the bright and the rough side of that loop. In the success story Caleb speaks, Haiku finds the Excel application and switches it to dark mode, clicking through menus as screenshots travel back and forth until the job lands. In the rough story the host asks for filtering on the ArtificialAnalysis website: models above the $1 line should be switched off. The model first unchecks the Opus 5.5 and GPT-6 rows, then quits halfway. After a follow-up instruction it moves a little further yet never reaches the goal, and the whole attempt takes 33 seconds. Despite 173 tokens per second, the host admits there is road ahead in both speed and reasoning, adding that his 4K monitor may have slowed things down. These breaking moments prove that spoken instruction plus computer control is not yet a finished craft.
Speed Alone Falls Short
The closing argument is crisp: cheaper, smarter and faster models unlock new downstream uses. Anthropic reports average running costs down 75 points, with a 90-point discount on jobs that fit inside the 100,000-token limit. A new effort setting lets users pick inference intensity from low to max; independent trials drew a pelican sketch for 0.09 cents at the lowest level and 3.38 cents at the highest. Yet the new tokenizer splits the same text into 1.25 times more tokens, hiding a price bump. Above the threshold Luna looks like the better deal, since the rival rises only to 20 and 75 cents past its 272,000-token mark. These details are backed by trial notes shared by SimonWillison, whose math matches the video.
| Metric | Value |
|---|---|
| Input up to 100,000 tokens | 10 cents per million tokens |
| OSWorld 2.1 score | 72.4 points, ahead of Luna |
| Cost per interaction | $0.001 through Atticus |
Key moments
AI commentary
"Small models keep getting smarter while their bills shrink, opening the road for personal agents. The above-limit price jump and unfinished tasks form the cautious side of the story."
AI assessment
The strongest counterargument comes from long-context pricing. Past the 100,000-token line Haiku turns five times dearer while Luna holds the 10-cent level up to 272,000 tokens, staying at the 20 and 75-cent line even beyond. Add the 1.25-fold tokenizer inflation and the table flips for long documents. Moreover the 72.4-point score belongs to the offline subset; nothing yet proves the same success across end-to-end workflows that take humans 1.6 hours. The gap between the tidy virtual rig and messy real websites calls for a cautious reading of the headline number.
The gaps list runs longer. Nothing concrete is shared on safety, data retention or enterprise auditing, while a personal agent that watches home screens leaves the privacy question open. The balance between local privacy and speed on one side and cloud cost on the other shifts per user, with no single formula. The sponsored hardware narrative muddies the picture further, placing costly accelerator rigs next to penny-priced cloud calls in the same video. Readers should therefore measure their own workload before deciding anything.
The practical takeaway compresses into three steps. First split the workload at the 100,000-token line: Haiku below it, the rival or a bigger sibling above it. Then experiment with the effort setting , drafting fast at the low level and reasoning carefully at the high one. Finally consider the subagent pattern, where the flagship model plans and the small model grinds through repetitive steps. Figures like $0.001 per interaction sound delightful, yet the monthly total deserves watching, because high volume fattens tiny fees quickly.
Sources
6 links; 2 of them also cited by 3 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube YouTube — Caleb Writes Code
- @anthropic Anthropic — Introducing Claude Haiku 5.5
Also cited by: Same Sticker, Different Bill: Haiku 5.5 vs Luna in 12 Runs · Haiku 5.5: Anthropic finally fixes its small-model problem · Claude Haiku 5.5: Anthropic's Cheapest and Fastest Model Reshapes the Small-Model Race
- @siliconangle SiliconANGLE — Anthropic releases Claude Haiku 5.5
- @simonwillison SimonWillison — Claude Haiku 5.5
Also cited by: Haiku 5.5: Anthropic finally fixes its small-model problem · Claude Haiku 5.5: Anthropic's Cheapest and Fastest Model Reshapes the Small-Model Race
- @xlang XLang — OSWorld 2.0 Benchmark
- @artificialanalysis ArtificialAnalysis — Model Comparison
haiku 5.5 · anthropic · osworld · ai agent · cost efficiency