Two leading labs announced models on the same Tuesday, pushing the week's count of major releases to four. SpaceX AI moved its own reveal earlier to avoid being overshadowed, so even the calendar became part of the contest. The first question was not which model was better in isolation, but which one fits whose workflow.
The Opus backstory matters for context. Version 4.6 had earned genuine affection, 4.7 and 4.8 read as minor steps, and Opus 5 never won people over. Then the Mythos and Fable families took the stage, leaving Opus in the shadows for a long stretch.
Opus 5.5: the official picture
The official notice carries a September 22, 2026 date: Opus 5.5 matches Fable 5.1 on most work while running 40 percent cheaper than Opus 5 on typical loads. Input costs 4 dollars per million tokens, output 20; cache reads fell 60 percent to 20 cents. Output generation runs over 30 percent faster, and five-hour limits rose across Pro and Max plans.
On safety, the model is the first release after the pacing-the-frontier note and went through outside review at Frontier Design and METR, posting the strongest alignment result Anthropic has recorded. Biology and cybersecurity carry Fable-style safeguards, with a life-sciences verification track open to vetted groups. Sam Bowman's team argues the model is safer than its predecessors and that shipping it lowers misalignment risk.
The scoreboard backs the claim with numbers: Terminal-Bench 4.0 puts Opus 5.5 at 66.4 percent, ahead of Fable 5.1 at 55.8, Opus 5 at 52.3, and GPT-6 Astra at 57.9. FrontierCode reads 54.4, CursorBench 57.8, GDPval-AA 1846 (Fable 5.1 at 1735, Opus 5 at 1708, Astra at 1542), tool-using HLE 67.7, OSWorld 81.8, Chartography 89.0, AutomationBench 40.0. Agentic science work scores 58.7, doubling the 29.0 of Opus 5; Astra leads at 64.6 but at roughly five times the per-task cost.
Independent tracking tells a similar story: the Artificial Analysis index scored Opus 5.5 at 58, with Fable 5.1 and Astra stuck at 53. The RoboNuggets roundup measured Opus 5.5 seven points above Opus 5. The Sol side reads differently, with GPT-6 Sol landing near 5.6 because OpenAI aimed at cost rather than capability. The Opus-over-Fable and Astra-on-top details each rest on single-test ground, which is worth remembering.
Sol and Luna: the cost front
OpenAI framed Sol and Luna as Astra advances in faster, cheaper form, with a 50 percent cut against GPT-5.6 promotional pricing. API identifiers gpt-6-sol and gpt-6-luna went live across ChatGPT Work and Codex for every paid tier from Plus through Edu. Luna also reached Free and Go users on the desktop app for trial. Sam Altman stressed the best-model-per-price-band goal and per-task pricing logic.
The price drop changes usage shape: higher limits plus lower cost push teams to run models more often and on longer jobs. On Codex, allowances lasting longer than on Claude Code creates a practical edge for OpenAI in daily work. As coding agents take on extended tasks, sustained-use cost becomes decisive; Fable and Claude suit long unattended runs, GPT models suit short loops.
Community verdicts split: Sol earned the S-tier phone analogy, delivering much of Astra near a fifth of the price. The Every team picked Sol for daily reading and writing, Opus 5.5 for ambitious code and visual work. Writers declared Claude was back, Jeffrey Emanuel reported bugs that had resisted Fable and Astra for weeks getting caught, and Chubby argued that hunting a single winner asks the wrong question.
Early user evidence
The Box measurement under Aaron Levy is concrete: 63 percent fewer tokens , 42 percent less verbosity, 30 percent more speed against Opus 5. Finance tasks gained 39 percent accuracy, cloud-cost analysis 65; consumer products added 17 percent, clinical diagnosis 15. Optiver matched quality in half the turns at 40 to 50 percent lower cost, Kiro made 40 percent fewer calls on half the tokens. A 200,000-line audit closed in under three hours where Opus 5 had needed over twenty. The HAProxy port took 9.5 hours against 12 at 51 percent lower cost, and a 680,000-line migration finished within a day.
The creative showcase filled social feeds: Alex Albert generated clay animation from one instruction, Chase Lin built an interactive coral reef, a Golden Gate scene came out of model code alone. A Market Street view rendered in 1906 texture inside Blender, Jake Eaton shared paintings made of 7,500 lines of pixels. Tarek had his personal site redrawn with its update deck, and the Daily Brief author found the Opus 5.5 review of his own site clear and actionable.
The writing front opened its own victory page: the Every team called it the most readable output they had seen, and the removal of long dashes became symbolic. Key facts move to the front, long sessions stay followable, feedback loops feel enjoyable. Makayla Wrigley fused the beloved 4.6 persona with Fable-grade taste, the Macallan team apologized for the older model and pointed at the new one. In the Eric Tech comparison, Opus 5 output read long and tangled while 5.5 came back short and trackable.
A three-minute duel and effort economics
RoboNuggets sent the same landing-page brief with identical design references to both models. Sol returned a clean, professional page with a tidy animated flow. The Opus side won the aesthetics round with dot animations, a try-it-yourself interactive app, info cards, and more readable charts. The author's verdict was crisp: Opus for interface taste, Sol where plain professionalism suffices.
Eric Tech laid out effort economics: Opus 5.5 at medium effort matches Fable at high effort. Agentic coding leadership at low effort sits with Opus, business workflows stay with Astra. Real-world knowledge scores rise with value; the author tied the Sol duel to a viewer poll and ranked the communication cleanup among the biggest upgrades.
Animation class and engineer notebook
Andy Lo produced hand-drawn teaching animation from one brief: Hawking radiation as story, a catapult as interactive physics with adjustable settings. The prompt skeleton holds a style block, color palette, puppet-theater frame, girl narrator, paper transitions, and a seven-scene timestamped plan. Robotic narration got fixed through Fish Audio, with Claude Code rendering separate audio per scene.
Duncan Rogoff reduced the engineer playbook to three rules: hand over the whole task, define done, state when to stop and ask. Deleting think-carefully phrases saves tokens and context, and the effort slider covers speed needs. Mid-run follow-ups typed while work continues cost less than waiting and re-reading full context. CLAUDE.md gains continue-and-stop rules, destructive-action guards, subagent splits, and a tasks.md checklist.
Generic taste warnings fail on the design side; banned patterns like cream backgrounds must be listed one by one. Concrete picks such as italic headline accents, numbered section labels, and pill buttons stop the slide into default AI aesthetics. The rule is simple: write what you exclude specifically, not vaguely.
The limits section belongs on record too: Fable-style biology and cyber curbs lock out some legal and life-science jobs in practice. Teams awaiting verification complain they can no longer fall back to Opus where Fable refuses. Over-refusal hits few people, but where it hits, the model turns unusable.
Terminal-Bench 4.0 scores
- Opus 5.566.4%
- GPT-6 Astra57.9%
- Fable 5.155.8%
- Opus 552.3%
| Topic | Summary |
|---|---|
| Opus 5.5 | Fable-class power, 40 percent lower cost |
| Sol and Luna | Half price, aimed at everyday work |
| Choice | Sol for daily tasks, Opus for ambition |
| Model | Terminal-Bench | GDPval-AA |
|---|---|---|
| Opus 5.5 | 66.4% | 1846 |
| Fable 5.1 | 55.8% | 1735 |
| Opus 5 | 52.3% | 1708 |
| GPT-6 Astra | 57.9% | 1542 |
Key moments
AI commentary
"Two leading labs shipped on the same Tuesday, so I read five separate reviews against the official announcements and independent scores before writing. I keep my own editorial voice here: no hype, just what the numbers and early users actually show."
AI assessment
The strongest version of the cost story runs like this: at a fixed performance level, model cost has fallen roughly 47 percent per quarter since 2023, outpacing DNA sequencing fourfold, computing sixfold, and batteries eighteenfold. Aaron Levy's Jevons point follows naturally, because longer limits plus lower prices genuinely expand the set of jobs worth handing to agents.
The limits deserve equal weight: much of the excitement rests on single-test readings, and the new guardrails already block some professional work, with legal teams reporting more refusals and weaker time management. Astra still leads agentic science, so the Opus advantage is real but not universal.
Attribution matters for trust: prices and headline scores come from Anthropic and OpenAI announcements plus secondary write-ups, independent grounding comes from the Artificial Analysis index, and the enthusiasm wave comes from early-access builders. The promised Sol duel has not happened yet, so independent confirmation on the Sol side is still missing.
My practical read splits two ways: Sol and Luna win routine reading, writing, and light coding on price, while Opus 5.5 takes long agent coding and ambitious visual builds. Teams embedded in Codex stay put when scores are close, so the choice hinges more on workflow loyalty than on raw supremacy.
Sources
10 links; 3 of them also cited by 11 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — AI Daily Brief: Opus 5.5, Sol and Luna
- @youtube.com YouTube — RoboNuggets: Sol and Opus 5.5 in 3 Minutes
- @youtube.com YouTube — Andy Lo: Teaching Animation with Opus 5.5
- @youtube.com YouTube — Eric Tech: Opus 5.5 Coding Test
- @youtube.com YouTube — Duncan Rogoff: Opus 5.5 Engineer Guide
- @anthropic.com https://www.anthropic.com/claude-opus-5-5
Also cited by: The superintelligence race: Musk and Huang on energy, orbital compute and safe agents · Claude Opus 5.5: From Idea to Finished Work in a Single Session · Before Buying the $12,000 Mac Studio, Rent a Test for $2 an Hour · The AI Week That Packed World Models, Open Weights and Data Centers Into Orbit · OpenAI's "o" Assistant, Claude Sonnet 5.5 and MiniMax M3.1: Three Model Bets Before DevDay · The AI Price War: Opus 5.5, GPT-6 Sol and the Quiet Bottleneck of the Agent Era · Decide First, Generate Later: The Jev Plus Opus 5.5 Playbook · The Market Is Underestimating This Massive AI Demand Shift · I Tested Opus 5.5 on 24 Coding Tasks: Fable-Level Quality in Half the Time and Cost · Claude Opus 5.5 in the Wild: One-Hour Award Site, 3D Space Voyage and a Single Dashboard
- @platform.claude.com https://platform.claude.com/docs/en/models/opus-5-5/overview
Also cited by: Claude Opus 5.5: From Idea to Finished Work in a Single Session · I Tested Opus 5.5 on 24 Coding Tasks: Fable-Level Quality in Half the Time and Cost · Claude Opus 5.5 in the Wild: One-Hour Award Site, 3D Space Voyage and a Single Dashboard
- @labellerr.com https://www.labellerr.com/blog/claude-opus-5-5-vs-opus-5/
- @openai.com https://openai.com/index/introducing-gpt-6-sol-and-luna/
Also cited by: Space Bunny Alpha: Inside OpenRouter's Free Anonymous AI Experiment · The AI Week That Packed World Models, Open Weights and Data Centers Into Orbit · The AI Price War: Opus 5.5, GPT-6 Sol and the Quiet Bottleneck of the Agent Era · The Market Is Underestimating This Massive AI Demand Shift
- @community.openai.com https://community.openai.com/t/announcing-gpt-6-sol-and-gpt-6-luna-in-the-api-codex-and-chatgpt/1399925
opus 5.5 · gpt-6 sol · gpt-6 luna · ai models · coding agents · model pricing