Back to feed

The AI Price War: Opus 5.5, GPT-6 Sol and the Quiet Bottleneck of the Agent Era

On September 22, Anthropic and OpenAI shipped new models roughly an hour apart. Opus 5.5 runs about 40 percent cheaper per task, while GPT-6 Sol and Luna arrive at half the price of their predecessors. In the same week Gemini Notebook added live conversation, xAI released Grok 4.7, and Meta said its Muse agent is coming to AI glasses. Prices are falling fast; the real bottleneck now is trust and oversight inside agent workflows.

Imported to Nodesdaily: (UTC+03:00)
Visualization: nodesdaily AI
Watch on YouTube — Q6uuvZmb0t8
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

On September 22, two of the largest AI labs shipped on the same day. Anthropic announced Claude Opus 5.5 and OpenAI published GPT-6 Sol and GPT-6 Luna roughly an hour later. The timing was not accidental: both companies are chasing the same buyers, the professionals who run coding and research work every day. Four models with lower price points inside a single week is the clearest sign yet that this market now competes on unit cost rather than raw intelligence.

Start with the numbers. Anthropic reports that Opus 5.5 costs roughly 40 percent less per task than Opus 5. There are two components. Input and output tokens dropped from 5 and 25 dollars per million to 4 and 20 dollars, and the cache read price fell from 0.50 to 0.20 dollars per million tokens. That second line matters most, because cache reads make up the majority of cost in long-running agent and coding workloads.

OpenAI's move was a straightforward price cut. Compared with GPT-5.6 Sol, GPT-6 Sol drops input from 4 to 2 dollars and output from 20 to 10 dollars per million tokens. GPT-6 Luna falls from 0.20 to 0.10 dollars on input and from 1.20 to 0.50 on output. That is a 50 percent reduction across both models. The company also states that GPT-6 Astra remains its strongest model overall, positioning Sol and Luna for price-sensitive volume work.

The context matters more than either announcement. Anthropic cut the price of its own flagship class on the same day its competitor halved its own pricing. Independent writer Simon Willison described the sequence as a familiar price war, and Grok 4.7 launching at 2 and 6 dollars per million shows the pressure is industry-wide. When supply-side competition persists, buyers gain negotiating power quickly, and the effective cost of intelligence starts to fall for everyone.

Now the part beyond price. The real promise of these models is not that they are smarter but that they do the same work with less computation. Anthropic attributes the gain to lower serving requirements and reports more than 30 percent faster output than Opus 5. OpenAI's same-day announcement pairs GPT-6 with better prompt caching, aimed at persistent agents that would otherwise repeat the same calls. Efficiency here is not a slogan; it is an infrastructure decision.

There is a trade-off. A smaller, faster model also carries less context. In testing described by the source, Opus 5.5 navigated the web unusually quickly but was weaker on some research tasks. The same source calls it the best model it has used for matching writing tone, while noting that another model remained slightly better on design work. In other words there is no single winner, and the task determines the answer.

Hands-on tests in the same source point the same way. Given raw footage to clean, both models used roughly one percent of a weekly quota, but the run times diverged sharply: about seven and a half minutes for one model against fifteen and a half minutes for the other. Quality and latency trade off from model to model, and a single benchmark table rarely captures that difference.

The second major thread this week is Gemini Notebook. Google announced real-time conversation with notebooks in its mobile app and interactive overviews under reports. Those overviews are not plain text; they combine generated assets such as infographics, quizzes, flashcards, mind maps and slide decks into a single flow. Usage limits also moved to a compute-based system on September 2, with the option to schedule generation for later when a quota runs out.

The source also describes using notebooks as a source inside Google Docs. That capability did not appear in Google's official announcements at the time this article was prepared, so it is best read as the source's own observation rather than a confirmed update. Several reading-assistant features are similarly spread across plans and time windows. The tools are changing quickly, and feature lists age even faster.

The third thread comes from xAI. Grok 4.7 shipped at the same price and speed as its predecessor, built on a larger base model and a longer reinforcement-learning run. The company says the model was trained to understand its own agent harness natively and offers a 500,000-token context window. Voice memos and voice calls for agents arrived the same week, following the company's earlier announcement that these agents run on their own cloud computer with enterprise access.

The fourth thread moves into hardware. At its Connect event, Meta said its Muse agent is coming to AI glasses, where you activate it by name and can ask about what you are looking at. It also announced a camera-free, audio-only glasses model, a small standalone agent device, and a target of more than one hundred glasses styles by year end. One detail deserves attention: the pocket-sized device is still in development and its materials are not final.

Behind all of this sits one economic fact: agents cost money. A model that reasons better does not raise the bill, but a model that takes more steps does. In long tasks, repeated calls, memory and tool use multiply cost horizontally rather than vertically. That is why these price cuts matter less as promotion than as a change in margin structure. Cheap tokens are what make sustained agent use viable at all.

The concern runs the other way too. Smaller and faster models also lower the cost of a wrong decision. If an agent sends the wrong email, deletes the wrong file or confirms an inappropriate action, a cheaper call does not shrink the damage. The meaningful unit of measurement is no longer the token but the cost of verifying a decision. Any pipeline that skips that step is automating a risk, not a process.

So what should a team actually do? I would avoid betting everything on a single model and instead stack by task type. Short, high-volume classification, routing and scoring work is well served by cheap fast models. Long and risky work needs a strong model paired with an audit trail and a way to roll back. The other lesson of this week is methodological: price and speed comparisons without a quality measure are misleading. Running the same task across three models and comparing the output blind leaves at least a transparent record.

The week does not produce a single winner; it produces a market where three companies cut prices in the same direction. Opus 5.5 lowers cost per task and leads on long context, GPT-6 Sol and Luna target price-sensitive volume, and Grok 4.7 is being woven into its own agent harness. Falling token prices are a precondition for the agent economy. They are not, on their own, a justification for running automation nobody checks.

Output price per million tokens

  • Grok 4.7$6
  • GPT-6 Sol$10
  • GPT-5.6 Sol$20
  • Opus 5.5$20
  • Opus 5$25
Output token list price; on identical workloads, cost scales with output volume.
MetricOpus 5.5GPT-6 SolGrok 4.7
Input ($/M tokens)422
Output ($/M tokens)20106
Context windowLong context focus1.05M tokens500k tokens
PositioningLong tasksGeneral purposeIts own agent harness
Change vs previous version40 percent cheaper50 percent cheaperSame price and speed

Key moments

  1. Two model launches on one day
  2. Three-model motion graphic test
  3. Video editing run times compared
  4. Small standalone agent device

AI commentary

"The real story this week is not the model count but the unit cost falling by half. Getting the same capability with fewer tokens changes infrastructure budgets directly. Price cuts alone are not enough, though: in agent workflows the decisive question is who pays when a decision is wrong."

AI assessment

The strongest claim here is the price war, and it holds up: Anthropic's own price table, OpenAI's 50 percent cut and an independent commentator's read all point the same way. But the 40 percent figure is the company's measurement at default settings on its own workloads. Long-context research, multi-file code and heavy tool use can produce a different ratio, so it should be read as a vendor scenario rather than a general law.

What is missing is independent confirmation that the models do the same work at the same quality. The comparisons described are single-person measurements of one person's workflow. Different context lengths, languages and long-running projects can flip the result. A model that wins on short tasks can fall behind on long ones, so model choice should rest on measurements inside your own workload rather than on a general ranking.

A fair counterpoint belongs here too. Price cuts can be driven by genuine competition, but how long that competition lasts is unclear. Companies may be absorbing unit losses through accelerated depreciation or temporary pricing. Re-checking pricing pages and contracts at decision time matters more than any number in this article. For teams making long commitments, quotas and latency guarantees are a safer metric than sticker price.

My practical takeaway is simple: classify the work first. Give short, repeated decisions such as routing, classification and scoring to cheap fast models. Run long and hard-to-reverse work with stronger models, log every step and keep human approval in the loop. Read the price drop as a verification budget rather than a windfall, because a cheaper call does not make a wrong answer less expensive; it only makes more wrong answers affordable.

Sources

9 links; 6 of them also cited by 20 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

artificial intelligence · model pricing · agents · token costs · claude opus · gpt-6 · nodedaily

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…