It was a crowded week for frontier models. Claude Opus 5.5 arrived while OpenAI talked about GPT-6 Sol and Luna, and Mimo 2.6 was mentioned but set aside for being too slow to test fully. The core question in the video was whether 5.5 is a small increment like 5.1 or a real jump. Internal chatter had floated 5.2, and Fable had shipped as 5.1 without much impact. On the author's private coding leaderboard Opus 5 had trailed Astro medium and even Solve high by fractions — not far, but not first.
Method: One Bash Loop, Four Projects
The method was intentionally plain. A Bash script called the Claude Code CLI for 24 coding prompts. Twenty tasks were scored up to 20 points each for edge cases and imperfect paths across different projects; two were PHP, one Dart Flutter, one Go. Separately, a React and TypeScript frontend and a PHP Laravel backend were judged by GPT 5.6 up to 20 points for code quality. The first run finished all checks in about two minutes, the second even faster. Single-run variance is high, so the author may move to a three-run average in the future, but for now cost kept it to one reported run.
Speed and cost were the immediate surprise. With Opus 5, both medium and high averaged over a dollar per prompt even at the 20-dollar Anthropic plan; with Opus 5.5 the entire 24 prompts stayed under a dollar. The author usually spent three to four five-hour windows to test both Opus 5 variants. This time a single window was enough. On social media the model was framed as Fable-level but faster and cheaper than Opus 5; the first impression and the detailed numbers both pointed that way.
Scoreboard: 20 Out of 20 on Front and Back End
The score detail revealed a curious twist. In the spreadsheet both Opus 5.5 variants sat side by side. Medium scored a touch higher than high on code quality — plausible with a single try. Yet both were well above the 100-point ceiling others hit; three perfect 100s out of six runs were seen, averaged and then divided by five for the 20-point scale. Over a three-run average medium still edged high by a small margin, but high recovered on other axes.
The overall leaderboard update produced two new leaders above GPT-6 Astra. High took a clean 20 out of 20 across the four projects. Medium tripped once and logged 18 out of 19 elsewhere, still slightly ahead on code quality but behind high overall. Compared with Opus 5 the jump is stark, much closer to the 60 out of 60 ceiling. It is the second model to clear 20 out of 20 after Astra medium, but price and time are the bigger story. One viral post had argued we need not better models than Fable or Astra but normal models that are cheaper — Opus 5.5 matches that brief. Medium averaged 56 cents per API call versus over a dollar for either Opus 5 variant, and even Opus 5.5 high was cheaper than Opus 5 medium.
Why So Much Faster and Cheaper?
The official API price is only a little lower; the main saving comes from needing less compute, so the model is tuned for speed and cost overall. This was echoed by heavy users who said pushing hard barely dented their limits — Theo was cited. On timing, medium landed near two minutes per request, high near three, about twice as fast as Opus 5 and faster than Astra medium; the whole table has no other entry under two minutes except Luna low, which is off the leaderboard. Usage on the 20-dollar plan also eased: one spike over 100% was absorbed by a new cached reset, the test week stayed at 91% and then used 58% of a second window to finish, versus three to four windows for Opus 5. Anthropic also lifted five-hour window limits by 20% officially.
Community signal lined up. The author said he now trusts community feedback more than generic suites because fixing real repositories mirrors daily work. Opus 5.5 scored near the top there, cheaper than Fable, and social feeds filled with Halleluja-style praise about finally ignoring the usage cap on Fable. What is Fable's role at its premium price, or Sonnet 5 at a supposedly lower price? Sonnet 5 sits lower on the table and is neither much cheaper nor much better. Anthropic's note is clear: Opus 5.5 is the first in the Claude 5.5 family, so Sonnet and perhaps Haiku or even Fable 5.5 may follow soon. The author plans to keep adding new models to his benchmark and will cover GPT-6 Sol and Luna next, with deeper dives on his AI Coding Dealer site.
| Dimension | Finding |
|---|---|
| Quality | High 20/20, medium close behind |
| Speed | 2-3 min/req, half of Opus 5 |
| Cost | Avg $0.56, single 5h window |
| Model | Avg API Cost | Time / Request |
|---|---|---|
| Opus 5.5 Medium | $0.56 | ~2 min |
| Opus 5.5 High | ~$0.90 | ~3 min |
| Opus 5 Medium/High | >$1.00 | ~5-6 min |
Key moments
- Opening — why 5.5 not 5.1, 5.2 chatter and Fable 5.1
- Method — 24 tasks, Bash via Cloud Code CLI, 20+2+1+1 split
- Surprise — first run ~2 min, second faster, under $1 total
- Quality — React and Laravel judged by GPT 5.6 near 100/100
- Leaderboard — high 20/20, twin peaks above Astra
- Math — 56c vs >$1, Theo stress test and 20% lift
- Community — repo-fix signal, first 5.5 family and what's next
AI commentary
"What strikes me is that we were not looking for a more expensive best model but for a daily driver that is fast and cheap — Opus 5.5 looks like it fills that gap."
AI assessment
Steelman: Opus 5.5 delivers near-Fable quality with everyday speed and cost. Twenty-four tasks finished in one five-hour window, under a dollar, at two to three minutes per request. The argument is not for a better model but for the same quality made accessible — and that argument lands.
Limits are clear. A single-run report inflates variance; medium beating high is likely chance. The judge is GPT 5.6 alone, with no human review. The family is still incomplete — Sonnet 5 scores lower, Fable's value at a premium remains uncertain. Single anecdotes like a 680,000-line migration in a day do not generalize, and the project types tested are narrow.
The takeaway is operational efficiency more than raw capability. A small API price cut matters less than needing less compute overall. That translates into five-hour limits that feel larger in team use. Community trust is shifting toward practical repo-fix tests rather than generic suites; Opus 5.5 scoring near the top there is a stronger signal than paper scores.
Practically, start with Opus 5.5 medium for daily coding and agent work, try high for critical jobs. Track single-window consumption; with the limit lift and cached reset, expect three to four times fewer windows than with Opus 5. For a Fable decision, wait for Sonnet and Haiku 5.5; justifying Fable's premium is hard right now.
Sources
6 links; 3 of them also cited by 10 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — Opus 5.5 24-Prompt Coding Test
- @anthropic.com https://www.anthropic.com/claude-opus-5-5
Also cited by: The superintelligence race: Musk and Huang on energy, orbital compute and safe agents · Claude Opus 5.5: From Idea to Finished Work in a Single Session · Before Buying the $12,000 Mac Studio, Rent a Test for $2 an Hour · The AI Week That Packed World Models, Open Weights and Data Centers Into Orbit · OpenAI's "o" Assistant, Claude Sonnet 5.5 and MiniMax M3.1: Three Model Bets Before DevDay · The AI Price War: Opus 5.5, GPT-6 Sol and the Quiet Bottleneck of the Agent Era · Four Launches in One Day: the Opus 5.5 Comeback and the GPT-6 Sol and Luna Price Break · Decide First, Generate Later: The Jev Plus Opus 5.5 Playbook · The Market Is Underestimating This Massive AI Demand Shift · Claude Opus 5.5 in the Wild: One-Hour Award Site, 3D Space Voyage and a Single Dashboard
- @platform.claude.com https://platform.claude.com/docs/en/models/opus-5-5/overview
Also cited by: Claude Opus 5.5: From Idea to Finished Work in a Single Session · Four Launches in One Day: the Opus 5.5 Comeback and the GPT-6 Sol and Luna Price Break · Claude Opus 5.5 in the Wild: One-Hour Award Site, 3D Space Voyage and a Single Dashboard
- @techcrunch.com https://techcrunch.com/2026/09/22/anthropic-releases-opus-5-5-with-lower-prices-and-fable-level-performance/
Also cited by: Claude Opus 5.5: From Idea to Finished Work in a Single Session · OpenAI's "o" Assistant, Claude Sonnet 5.5 and MiniMax M3.1: Three Model Bets Before DevDay
- @finopsllm.com https://finopsllm.com/research/claude-opus-5-5-cost-model
- @aipricing.guru https://www.aipricing.guru/news/claude-opus-5-5-api-pricing-september-2026/
opus 5.5 · claude code · coding benchmark · price performance · anthropic