Back to feed

Opus 5.5 Claimed: Anthropic Pushes Coding Peak, Price and Prose at Once

Prompt Engineering presents Opus 5.5 as beating Fable and Astra on Terminal-Bench 4.0 at lower spend, pairing top agentic scores with sharply more natural prose, a modest API price trim and a new banked quota reset.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — 04qy4OWteio
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

Prompt Engineering frames Opus 5.5 as the opening entry in a new Claude 5.5 line and leads with a bold claim: near-parity or an edge over Fable 5.1 on almost every suite, while costing less than Opus 5. What the host cares about more than the numbers is voice. Earlier large models often fell into a uniform, mechanical register that readers learned to recognise as machine writing; the promise here is prose that scans like a person wrote it, which matters more than a leaderboard bump in long agentic sessions.

On pricing the video describes a modest but telling adjustment: per-million input and output charges and the cache read/write fees tick down a little. The host argues that intelligence per dollar now shapes model choice as much as raw intelligence. Even a small cut, in that reading, signals that Anthropic feels competitive heat and wants the value narrative to be clear, since heavy users feel a per-token change accumulate quickly.

The benchmark the host trusts most is Terminal-Bench 4.0, presented as still unsaturated and therefore still discriminating. Curated by Stanford and Harbor, it poses 89 hard tasks inside isolated terminals that mirror real workflows. In the video’s charts the new build moves ahead of Astra and Fable 5.1, and when the same test is sliced by effort setting the lead persists at lower spend. The host notes that no lab test fully mirrors real coding, yet a non-saturated test still gives a useful expectation, and the broader agentic panels point the same way.

Across the wider suite the story repeats: lead the board while shrinking cost. Earlier cycles pushed raw capability; this cycle tries to make that capability cheaper to actually deploy. Viewed effort by effort, the cost curve sits below rivals, a pattern the host highlights on every coding-agent panel. The same podium claim appears on the knowledge-work wall where Elo and cost per job are plotted together, with a hope that answers also become more concise.

The signature for hands-on coding is a check loop that goes beyond emitting files. In the host’s favourite prompt the system writes code and then opens a browser to inspect the rendered result, catching its own mistakes. The host frames that visual self-check as the hinge for unattended, long-horizon work: an agent that can look at its own output can iterate without a human in the loop, turning a clever snippet into something shippable.

One chart gives the host pause: on Frontier Code, medium effort outranks high and extra-high and ends up close to max. Similar non-monotonic wiggles appear for Fable and some GPT variants on this specific evaluation. The host’s hypothesis is a distribution mismatch — the test may sit outside the data mix the models were tuned on — so a single inversion does not overturn the broader hierarchy but does counsel caution in generalising from one board.

Knowledge-work claims sit alongside the prose story. The host expects shorter, tighter answers and points to the Eo and cost wall as evidence of efficiency. Yet the headline for many will be how Opus writes. Anthropic says it reworked the voice after repeated feedback that Opus 5 felt mannered and hard to scan in long sessions, with testers saying the new messages are easier to parse at a glance. The before/after illustration is a metric explanation: where the prior tone wandered into tangled hedging, the newer wording plainly states that a free-tier shift accounts for only about one point of the August drop, a comparison the host calls markedly more natural.

Shipping notes are brief and practical: update Claude Code and the desktop app, or call the API under the 5.5 label. The video flags the usual annexes on retention, distillation and safety and points interested viewers to the docs rather than reciting them. The tone is measured transparency, presented as footnotes rather than headlines.

Demos carry the weight of the argument. A black-hole simulation follows a very detailed brief and yields a control-rich output where density and brightness can be tuned; the host finds the fidelity impressive even while noting that a busy CPU limits smoothness. A live tracker for the International Station makes real API calls, plots the current path and adds a small sun on the side that the host singles out as a delightful detail. A self-researched launch site stitches together research, animated figures and a comparison between Opus 5 and the newcomer without external assets, respecting the default medium effort setting and suggesting high for thorny tasks. A final 3JS voxel piece is presented as one of the most literal and richly detailed renders the host has seen, ticking every line of the brief using only model-generated visuals.

The close turns to capacity. The five-hour allowance has grown, the weekly budget has not yet, but users gain a bankable reset that the host reads as inspired by OpenAI and welcomes as a sign that labs must now compete on price and perks as well as capability. Taken together — clearer prose, stronger coding, a slightly lower bill and fewer tokens per job — the update is framed as a pragmatic step up rather than a fireworks moment, and competition, in the host’s view, is good for customers.

Visualization: nodesdaily AI

AI commentary

"What I find most honest in the video is not the leaderboard but the language shift: moving away from Opus's heavy, hard-to-scan voice matters more in long sessions than a few extra points. A small price drop plus a browser-grounded check that lets an agent verify its own render makes the idea of leaving agents to work overnight a touch more believable."

AI assessment

Steelman the counter-case and the crown stays provisional. The charts come from one narrator, with no official card, no independent leaderboard reproduction and no disclosed dollars-per-job methodology. The price trim is modest in absolute terms; a rival cut could erase it quickly. Voice improvement resists a single number — without blind human ratings the before/after anecdote remains an anecdote, however welcome.

Method limits are also clear: the three demos are curated, single-shot wins; failures are unseen. The Frontier Code inversion — medium beating high — suggests the evaluation is not monotonic across families and may reward a particular distribution, so winning there does not entail winning everywhere. Notes on brightness and CPU-bound smoothness in the black-hole scene remind that fidelity is not a pure model property but a system one.

On interest and verifiability, the tester is also the judge; independent reruns are needed. The internal codename wafer-eap and the Tuesday window circulate from a single X account relaying a leaker, with no second outlet, price sheet or system card — until Anthropic ships, 5.5 is a checkpoint label, not a product. Cost axes in the video mix effort levels without disclosing token counting, which makes the cheap-at-every-effort claim hard to audit.

Practically, teams that live in Claude Code and rely on long-horizon agents will feel immediate value from browser-grounded self-checks and tighter prose, especially at medium where spend is lowest. Yet treat the release as a candidate, not a default: benchmark your own repository, measure answer length and tool-use accuracy on your data, and test retention and safety notes against your compliance needs before swapping a production path. If the weekly quota is your bottleneck, wait for a clear weekly raise.

Sources

6 links; 2 of them also cited by 4 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

opus 5.5 · anthropic · claude code · terminal-bench · ai coding

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…

Opus 5.5 Claimed: Anthropic Pushes Coding Peak, Price and…