The pace suddenly accelerates. The video opens by noting the arrival of the Opus 5 family and then introduces a package presented as Opus 5.5 , followed by a report that GPT-6 arrived as Sol and Luna . Rather than settling the race on a leaderboard, the creator wants a scene that code itself produces, and sets a small game for viewers: two versions of the same cloth simulation, written by two different models side by side, with the audience invited to guess which is which.
What is said to change on tokens and cache
The headline promise is doing the same work with fewer tokens . The video says the new Opus variant generates about 40% lower bills than its immediate Opus predecessor, so subscription quotas drain more slowly and users keep more effective headroom . On top of that, the discount for reusing a session's prefix — the cache read — is described as 60% deeper , adding an extra cut on top of an already discounted line item. With roughly 20% off input-output on average, total cost in long conversations or in sessions resumed after a pause is said to fall noticeably; these rates are the video's reporting and do not replace the official rate card on their own.
Think of cache as a session's own memory. Before generating, the model turns the prompt into an internal state; a follow-up request sharing the same prefix can read that state instead of recomputing it, and the read is priced far below the write. Official docs describe 5-minute and 1-hour lifetimes : a 5-minute write costs 1.25x base input, a 1-hour write 2x, while the read is about 10% of base (even lower in some families). Opus 5 is officially listed on 24 July 2026 at $5 input / $25 output with a 1-million-token window , while the 5.5 label has no official GA entry — it reads as a single-source leak, so the reference for pricing and window remains the Opus 5 card.
On benchmarks the video stays cautious. It relays that the new Opus variant is presented as the best model on several tracks and beyond Fable 5.1 in some workflows, yet at the same time GPT-6 looks stronger on workflow-centric tasks . That is why the cloth simulation is staged not as a side demo but as the place where the claim meets a visible scene: besides planning, sustaining long-horizon agent work and debugging, the quality of the code is judged by the physics it can render.
The stage: interactive cloth
The scene is a fully code-generated interactive cloth simulation . What matters is not surface gloss but whether physical details are set up correctly. Both generations run at maximum effort ; the video notes one model at maximum for its effort setting and the other at its default maximum. Wind is added on screen, the simulation is run, and the cloth is pulled to watch how it tears . In one version the rip propagates with a slow, continuous glide as you keep pulling; in the other the tear shows a slightly cleaner surface along the edge yet the motion feels more clipped. That difference in continuity becomes decisive in the judgment.
Through the creator's eyes the split sharpens: one cloth carries more realistic continuity and wind dispersal , preserving a sense of inertia after the pull ends and delivering a visually more pleasing flow. The other offers a marginally smoother surface at the tear edge but feels less natural to manipulate, with the physics advancing in more staccato steps. The narrator weighs visual quality and physics quality separately: sharpness alone does not save the scene; the coherence of billowing in wind and tension at the moment of tearing is what makes the simulation convincing.
Speed, preference and a cautious close
The closing is sharp on both time and expectation. The preferred version is said to have been built in about 20 minutes , the other in about an hour — so the favored physics also arrived faster. The creator frames the Opus family's lag in this duel as a disappointment given its long-held place at the top, but adds that a single simulation is not a verdict: the model can do many other things and needs broader testing, especially on agentic coding and workflow strengths . The remaining note is sober: the Opus 5.5 tag remains an unconfirmed leak , and the GPT-6 Sol/Luna story is also an unconfirmed relay without an official model card , so the video's price and performance claims should be read alongside the official rate card and independent evaluations, not stretched from a single physics scene.
| Point | Note |
|---|---|
| 40% bill claim | Same work with fewer tokens; quotas drain slower, clear on long sessions. |
| Continuity on cloth | Wind and tear flow together; surface sharpness alone is not enough. |
| 20 vs 60 min | Favored physics also faster; one scene must not generalize, check official card. |
| Dimension | Opus 5.5 claim | GPT-6 in video |
|---|---|---|
| Billing | ~40% lower | — |
| Cache read | 60% deeper discount | — |
| Input-output | ~20% off | — |
| Sim time | ~60 minutes | ~20 minutes |
| Physics continuity | sharp surface, choppy flow | smoother, continuous |
Key moments
AI commentary
"What I take from the video is not a price sheet but a paired test: the promise of doing the same work with fewer tokens against whether the cloth actually drapes, tears and catches wind convincingly. The 40% bill cut and 60% cheaper cache read sound attractive, but the way the fabric flows as it rips is what tells you if efficiency converted into real output."
AI assessment
The strongest counter-case warns against stretching a single physics scene into a verdict: graceful billowing and continuous tearing are appealing, yet workflow automation, long-context planning and debugging are not compared in this video. If independent scores keep the Opus 5 line steady on planning and long-horizon agent work, a momentary lag on a cloth stage does not reprice the whole model; the 40% bill and 60% cache story may then matter more for subscription economics.
Limits are clear: a narrow sample — one generation per model, a single maximum effort point, no force parameters beyond wind and tear, and timing reported as about 20 minutes versus about an hour . Moreover, the 5.5 and Sol/Luna labels are unconfirmed without official model cards , so price comparison must be triangulated against Opus 5's official $5/$25 and 1M window and performance against independent workflow and agent evaluations.
The practical takeaway looks past the sticker price to the production line. Raising cache hit rate, calibrating effort to the task and reserving 1-hour writes for genuinely long waits is what turns the video's discounts into real savings. For model choice, do not decide on a single physics demo; run the two models side by side on a repeatable task from your own workload and compare time, tokens and output quality — otherwise a cache that looks cheap can turn into expensive writes in real use.
Sources
7 links; 2 of them also cited by 4 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — New Opus 5.5 features compared with GPT-6
- @anthropic.com https://www.anthropic.com/news/claude-opus-5
Also cited by: Claude Opus 5.5 in the Wild: One-Hour Award Site, 3D Space Voyage and a Single Dashboard · Opus 5.5 Claimed: Anthropic Pushes Coding Peak, Price and Prose at Once · Day 228: A 12-Minute Minecraft, a 3-Cent Model and $233K in ARR
- @claude.com https://platform.claude.com/docs/en/models/opus-5/overview
- @anthropic.com https://www.anthropic.com/news/claude-opus-4-6
Also cited by: Which AI Tools Are the Best in 2026? Dan Martell's Leverage Map
- @claude.com https://platform.claude.com/docs/en/build-with-claude/prompt-caching
- @anthropic.com https://www.anthropic.com/research/claude-opus-5
- @anthropic.com https://www-cdn.anthropic.com/files/4zrzovbb/website/14082576b71dc6b532c2d49094cf4661e957cd69.pdf
opus 5.5 · gpt-6 · artificial intelligence · pricing · cache · cloth simulation · physics