Paying the same money for half the service sounds like a bad deal, yet that is exactly what happened in the AI world last month. OpenAI cut the credit equivalent of its 200-dollar Pro plan in half, and Anthropic had made a similar move months earlier. Meanwhile a free model running on an ordinary laptop now rivals the best model of half a year ago. The speaker reads these two opposite moves as one picture: the top gets stingier while the bottom approaches free, and the whole logic of the bill is being rewritten.
The subscription squeeze: Astra and halved credits
September moved fast. OpenAI shipped its strongest model, GPT-6 Astra, on September 3, paused new sales of the 200-dollar Pro plan a week later as demand exploded, and on September 29 an employee named Thibault announced the plan was back but with half the API value per dollar. The catch is subtle: Astra keeps its unit price while GPT-6 Sol and Luna launched at half price, so users of the cheaper models keep their quota and Astra users lose half of theirs. Unveiled the same day, a 500-dollar tier with an eight-times-faster Ultrafast mode signals that raw speed will now be sold separately. This picture is confirmed by the September 29 report on the-decoder.com, which notes that the company is drifting from subscriptions toward pay-as-you-go billing.
The SemiAnalysis math: 14,000 dollars of usage for 200 dollars
Back in June, the research firm SemiAnalysis measured how much API-equivalent usage hides inside subscriptions, and the numbers were staggering. The 20-dollar Claude Pro equated to roughly 400 dollars of usage and ChatGPT Plus to about 700 dollars. The madness lived in the 200-dollar tier: Claude Max held around 8,000 dollars of credit and ChatGPT Pro about 14,000. The speaker argues this subscription subsidy could never last, that the losses had to surface somewhere, and he looks right. The detailed breakdown of this economy also appears in the June analysis on superhuman.ai, which reports that labs were losing thousands of dollars per subscriber.
Anthropic went through the same squeeze two months earlier. In July its flagship Fable left the 20-dollar plan for metered billing, while users on the 100- and 200-dollar Max plans could suddenly spend only half their limits on it. The current table on the company support pages sharpens the split: Fable 5 and 5.1 now count as a standard part of Max plans, while Pro users reach for their credit cards. Some of the speaker's details on version 5.1 do not match the latest table exactly, but the direction is identical: the best model is no longer the unlimited guest of a flat fee. This shift is confirmed by the official help article on support.claude.com, which states plainly that the promotional period ended in mid-July.
Four years from free trial to quota math
Today is the product of a four-year journey. ChatGPT opened free to everyone in late November 2022, the 20-dollar Plus price became the industry standard that February, and Anthropic copied it with Claude Pro the same year. The 200-dollar Pro launched in December 2024 promising unlimited use, and a month later Sam Altman admitted the company was losing money on it. Late January 2025 brought DeepSeek R1 at one twenty-seventh of o1-level pricing, and a week later Nvidia shed 589 billion dollars in a single day. What followed is familiar: Max tiers, weekly caps, agents running around the clock that each burn tens of thousands of dollars of usage, and a buy-credits button inside Codex limits. The speaker draws a blunt lesson: the experiment of selling unlimited intelligence for a flat fee collapsed in the age of agents.
The laptop claim: frontier quality moves home
The cheerful side of this gloomy picture is the laptop claim. The Alibaba-origin 27-billion-parameter Qwen model arrived in August and runs on ordinary laptops. The speaker benchmarked it against Opus 4.6, the frontier model of half a year earlier: the local model showed better taste on a landing page, a Tetris clone, and a fox-on-a-bicycle drawing, tying on a physics simulation. One correction is needed, since the video calls it Qwen 3.8 while the comparison page shows it is Qwen3.5 27B, the name used throughout this article. The source of that correction is the head-to-head comparison on artificialanalysis.ai, where Qwen3.5 27B sits at 23 index points, just below the Opus 4.6 score of 26.
The price of speed and a second round of tests
The local model has two big flaws, though: slowness and depth. Generating 10 to 15 tokens a second on a small machine, it stretches a landing page that takes 5 to 10 minutes on a subscription into a full hour. A second round of tests sharpens the picture: turning a car sketch into a 3D model, Blender work, and game building all leave the local model far behind the flagships, yet it nearly matches GPT-6 Sol on a weather app. On the intelligence index Qwen runs with GPT-6 Luna, DeepSeek V4.1, and Gemini 3.8 Flash, but it fails the class on long-horizon agent tasks and hallucination rates . The price gap is striking: Sol halved its predecessors with 2 dollars of input and 10 dollars of output per million tokens. That price break is confirmed by the launch report on thenextweb.com, which notes Luna chasing high-volume work at 10 cents of input.
Prices keep falling for three plain reasons. First, competition: Chinese models grow smarter and cheaper, and every release since the DeepSeek moment has forced the frontier labs into discounts. Second, efficiency: models can now tune how long they think, and Opus 5.5 on medium effort matches its predecessor with 65 percent fewer tokens, cutting costs by more than three quarters. Third, operating costs: in July OpenAI had its own models rewrite serving code on its own chips, shaving about 20 percent off, and the savings reached prices. The speaker favorite price-per-job chart sums it up: work that cost 9 dollars on Fable 5 now costs 30 cents on Sol.
Looking two years back, the picture turns even more dramatic. ARC-AGI, the visual-puzzle set that stumped AI for years, cost roughly 4,500 dollars to crack in December 2024, while a free model now matches that score for pennies. The set measures intelligence as the speed of learning unknown tasks, which implies ordinary intelligence has nearly become free. The source of that framing is the official definition on arcprize.org, whose stated goal is tracking the real human-machine gap scientifically. Naturally the price of growing capability stays put: at the very top, 10 dollars in and 50 dollars out per million tokens holds unchanged for both Astra and Fable.
The top stays expensive: hidden tests and new jobs
Narrow public-score gaps between open and frontier models should not mislead. The US government testing center CAISI evaluated DeepSeek V4 Pro in April 2026 and found its 5-percent public gap on SWE-bench ballooning to 16 points on private tests. The center, housed at the American standards body, reports the open-weight model trailing the frontier by about 8 months. These findings are documented in the official evaluation on nist.gov, which describes a broad frame spanning 35 models and 16 benchmarks. Open models washing out on the Terminal-bench science test while frontier models succeed tells the same story. Increasingly, what justifies flagship prices is entirely new workloads.
Video-editing trials look mediocre but improve fast, while on the gaming side Astra produces playable Minecraft-like demos in a single shot, workloads nobody even attempts on open models. The hallucination gap runs at two to one: frontier models are twice as good at not inventing answers they lack. Safety shows the same pattern, with self-hosted open models far more exposed to hijacking. The speaker closes with an honest balance: heavy flagship users will get less than before, yet most users never needed flagships at all. Buying a Mac will not replace a subscription in the near term, though as a long-term bet it makes sense.
Key moments
AI commentary
"The speaker makes a convincing two-front case: flagship models are getting stingier while ordinary work keeps getting cheaper. The real story is not prices but who foots the bill. Winners will be users who learn which job deserves which model."
AI assessment
The strongest counter-argument is that this picture presents a temporary balance as a permanent law. Shrinking subscriptions do not guarantee that cost declines reach users; providers may tighten quotas faster than they cut prices to defend margins. Plans once sold as unlimited turning metered within months suggests today cheap small-model prices could be a similar honeymoon.
The missing piece is the enterprise bill. While individual 200-dollar plans dominate the discussion, the real volume flows through corporate API spending, where it is unclear how much of each discount comes from competition versus efficiency. The IPO-pressure thesis also passes in a single sentence, with no independent data on how investor pressure shapes pricing.
The practical takeaway for readers is to sort work by merit instead of paying flagship rates for everything. Drafts, summaries, and simple coding belong on cheap models; long-horizon agents and production code belong on flagships. A local-model investment buys next year freedom rather than today bill, and anyone planning to quit subscriptions within six months will be disappointed.
Sources
8 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — The New Era of AI Has Just Begun
- @the-decoder.com The Decoder — OpenAI Pro plan half-credit report
- @support.claude.com Claude Help Center — Fable plan limits
- @superhuman.ai Superhuman — SemiAnalysis subscription economics
- @artificialanalysis.ai Artificial Analysis — Qwen3.5 27B vs Opus 4.6
- @thenextweb.com The Next Web — GPT-6 Sol Luna price cut
- @nist.gov NIST — CAISI DeepSeek V4 Pro evaluation
- @arcprize.org ARC Prize — What is ARC-AGI
artificial intelligence · subscriptions · api pricing · open source · local models · openai