Back to feed

AI Tier List Reset: GPT-6 Astra Takes the Crown as Subscription Math Rewrites the Ranks

Avenox updates his AI tier list with GPT-6 Astra on top, subscription math deciding most tiers, and open local models plus harnesses getting their own ladders.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — djlec83pVoU
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

Avenox frames this as an update video, not a cold ranking. The format mixes ordering with background notes and personal judgments, which matters because dozens of new models arrived at once and a bare S-to-D board would hide the reasoning.

GPT-5.6 Luna keeps the S tier on price-performance. A 20-dollar membership behaves like near-unlimited use, worth hundreds of dollars in metered spend, so budget-sensitive users get the best deal here. An incoming GPT-6 Luna only strengthens the case for staying in this lane.

Terra drops to C while Sol rises to S, with a correction attached. Terra costs roughly ten times Luna without earning that gap, while Sol's displayed price was wrong and sits near 30 dollars on output. Sol's expected upgrade path alongside Astra lifts it back to the top tier.

Claude Fable 5.1, meaning the latest Fable line, stays S despite cost and quota complaints. The whole series reads consistent from top to bottom, and model quality outweighs the pricing pain in this account.

DeepSeek 4.1 Flash lands at A after a price hike. Peak-hour pricing touches the 1.2 dollar band and the old Pro-era cheapness is gone. Against a Luna subscription worth 300 to 400 dollars of API use, DeepSeek can cost its core audience up to thirty times more, though the lab itself still earns respect.

GLM 5.3 Flash takes S on price plus subscription logic. Metered API credits rarely decide real costs, and a subscription route makes it far cheaper in practice. MiMo falls to C as dated even for cheap bulk cleanup, while Gemini 3.8 Flash holds B for speed and video viewing against weak tool calling and frequent hallucinations.

MiniMax M3 settles at B as adequate without excitement. A free tier on OpenRouter plus cheap apps and a growing ecosystem help students and light daily work. The doubt is forward-looking, with Claude and GPT lines expected to grow larger, plus a disclosure that no company sponsors the channel except an ongoing talk with Kimi. Meta Muse Spark 1.3 Contributor follows as the cheap-API pick at ten to twenty cents with data sharing, sensible mainly under 20 dollars of monthly use, exactly as predicted two months earlier.

Sonnet 5 moves from B to A on a cut toward the 10 dollar level. At six or seven dollars it would be S, since the model itself is fine and only the price was wrong. At Terra money Sonnet wins every time. Opus 5 stays a permanent S as the favorite daily series for range and manners, while Kimi K3 joins S for open weights, enterprise self-hosting without data leakage, a liked 200-dollar plan, and three trillion parameters. The missing piece is a smaller Kimi line for subordinate agents in the one-to-two dollar band.

Grok 4.6 takes A on price, ecosystem, and Cursor-side perks. Long unsupervised runs still do not inspire the same trust as Astra, Fable, or Kimi. Hy4 Preview is an uneventful C, Nemotron a straight D, and GLM 5.3 repeats at S. DeepSeek's problem is now crowded competition from Luna, Muse Spark at a sixth of its price, and GLM Flash until it ships its own subscription.

GPT-6 Astra claims the top seat and, in this account, the best-model-so-far title. Three-dimensional scene reading, fine detail spotting, rule following, strong reasoning, and no mid-plan quota cut carry the argument. Working styles differ, so some users may disagree, and chatting still feels more enjoyable on Fable. The launch is even read as damage to Anthropic's IPO, with a new Opus or bigger Fable expected soon.

Local open models get their own ladder. Qwen Max sits at B while the small 35B and 27B local Qwen build jumps to S for offline use, bank and insurer fine-tuning, and work that must not leak data. Gemma 4 31B moves up on size-adjusted quality, well ahead of Gemini itself. Gemini 3.1 Pro is C, and Mistral is A only because Europe has no rival buyer for banks; the weights alone would be D.

Harness rankings start with Antigravity at B, lifted from D by free access and older Opus availability despite constant renaming. Claude Code rises from B to A on remote control, agent-to-agent messaging, and mid-task effort changes that no longer blow the cache, with a good price-to-allowance ratio until limits fall. Codex is a straight S on a closed 200-dollar plan with heavy allowances and resets, usable inside Hermes Agent and other agents, plus a browser and computer use that configured Cloudflare and Workspace mail and built a site with drag and drop. Cursor moves to A on strong infrastructure and rising quotas after the SpaceX deal, with Grok quality and a team that understands harnesses.

Hermes Agent is S as a personal assistant rather than a pure coder, run on a server and driven from Telegram, presented as capable of automating much startup service work. Pi agent shares the top area for simplicity and adaptability, with OMP as its client-side twin from a Turkish builder solving similar problems. OpenCode is B because the harness feels like an app pretending to be a terminal, Kilo Code and Cline are aging C tiers, Devin is D, Windsurf is left unranked from lack of use, and GitHub Copilot is D with shrinking allowances.

Media models close the board. ElevenLabs is S with no real rival, Veo 3 is B because Seedance opened a visible gap, though Veo may still lead on Turkish prompts and costs less; one e-commerce anecdote has a user switching to Seedance and staying there. Nano Banana 2 drops to B as GPT Image 2.5 pulls ahead on quality, price, and updates, used here inside Codex under Astra control. GPT Image 2.5 is S, Suno is S as the unmatched music maker, Seedance 2.5 is S in a champions league of its own, Gemini Embedding 2 is S for joint video, audio, and text indexing that simplifies RAG, and Qwen embeddings sit at B. The closing math is personal: 400 dollars a month on Claude Code plus Codex could be cut to a third or quarter at 95 to 97 percent quality, but the orchestration work plus a five to ten percent quality gap makes paying more the rational choice for heavy output and the wrong choice for students.

Visualization: nodesdaily AI

Monthly payment comparison

  • Luna Plus20 dollars
  • Sonnet level10 dollars
  • Kimi plan200 dollars
  • Codex plus Claude400 dollars
Monthly outlay logic from the video: one flat subscription can replace multiples of metered API spend for heavy users.
ModelTierPrice signal
GPT-6 AstraSIn-subscription, no cut
Claude Fable 5.1S10 dollars in / 50 out
Kimi K3S200 dollar plan
DeepSeek 4.1 FlashAPeak around 1.2 dollars
Claude Sonnet 5AAround 10 dollar level
Gemini 3.8 FlashBFast but hallucinates

AI commentary

"I see this update as a price-performance ledger more than a capability ranking, and on that ledger the subscription math now beats raw benchmark scores."

AI assessment

The strongest objection to this tier list is that it prices one person's subscription stack, not the models in a vacuum. Flat-rate AI plans are already bending under agent workloads, with providers quietly trimming allowances and GitHub Copilot moving to metered use. If the 20-dollar all-you-can-use math breaks, several S tiers priced on subscription generosity would need a recount.

What the video does not show is method: no fixed test set, no durations, no repeat runs. Prices quoted here move by the hour, from DeepSeek peak windows to a 75 percent cut on Fable cache reads and a per-task figure near nine cents on GLM Flash. I treat every number below as a September snapshot and re-check before any buying decision.

Provenance matters on two claims. The Kimi sponsorship talk is disclosed as a conversation, not a deal, and the 2.8 trillion parameter open-weight claim needs independent testing; outside reviews call K3 the most capable open-weight option but also the priciest and slowest among them. The Seedance verdict lands while studios are openly threatening suits over training footage, so I keep the video-quality praise separate from the legal risk.

For my own stack, the split is simple. If I billed under 20 dollars a month of metered use, I would live on Luna-style subscriptions plus Qwen or Gemma local models for private data. Since my work pays for sustained agent runs, Codex plus Claude Code earns its keep, with Kimi K3 as the open-weight escape hatch and Seedance plus ElevenLabs only for paid video jobs.

Sources

10 links; 5 of them also cited by 24 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

tier list · astra · price/perf

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…