Back to feed

Opus 5.5 Leak and China's 600-Billion-Parameter Wave: Qwen, Kimi and MiniMax Crowd the Same Week

Anthropic's Wafer-coded Opus 5.5 leak points to a Tuesday window, while China's StepFun 600B, MiniMax M3.1, Alibaba's Qwen family and Moonshot's Kimi K3.1 all surfaced in the same week. Cheaper frontier pricing and giant open-weight schedules turned the week into a straight-line race.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — i00isgmGgGg
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

The frame this week is not a single announcement but a compressed schedule. The video stitches five separate leaks into the same 72-hour window: a new Opus trial on Anthropic's side and, from Beijing and Shanghai, fresh signals from StepFun, MiniMax, Alibaba and Moonshot. Like overlapping stages at a festival, every lab seems to have pulled its clock forward. For me this clustering is not coincidence; it is the intersection of price pressure and open-weight timing.

The lead on the Anthropic side is Opus 5.5. The trial, referenced under the shorthand Wafer, has retired the earlier 5.2 tag and is being tested internally as claude-wafer-eap. The leak points to a Tuesday window and it remains open whether this arrives alone or as a bundle with updated Sonnet and Haiku. The core idea is not to activate the whole model at every step but to run a dense frontier model more efficiently. Think of a hospital that pages only the relevant specialist instead of opening every department for every case.

The pricing claim carries the headline. The leaked table talks about 4 dollars per million input tokens and 20 dollars per million output tokens for Opus 5.5, with cache reads at 20 cents and writes at 5 dollars. Against the current Opus 5 anchor of 5 and 25 dollars, that is roughly a 20% cut. If the number holds and the performance chatter around GPT-6 Astra level holds with it, the offer becomes sharp. For a legal team that wants to summarize a ten-page agreement and then reason over a twenty-page appendix in one go, the same job closes with a lower invoice. A small context window reduction is also noted, but the team appears to have traded a narrow cut for speed and cost.

The demo reel explains the excitement. The video shows single-prompt SVG scenes, procedural three-dimensional environments and a BMW M5 rendering stitched side by side. Wheels and a few joints are still imperfect, but spatial coherence and material separation are visibly ahead of the prior trial. A coffee machine example makes the same point: the model connects not only shapes but functional part relations. In a direct side-by-side with the earlier Opus 5.2 trial, 5.5 produces cleaner geometry from the same prompt set. It feels like moving from hand sketches to a parametric model in the workshop.

StepFun's Step 5 Preview is the most concrete launch of the week. A 600-billion-parameter sparse mixture-of-experts (MoE — only the relevant experts fire for each token, not the whole brain) with 27 billion active per token. Context window is 1 million tokens and image input is native. On Artificial Analysis Intelligence Index it sits at 44, in the band of GLM 5.3 and Kimi K3. Pricing is the sharp part: StepFun's own chart puts per-task cost about 65% lower than rivals, essentially promising the same research job at roughly one third. How it works in four steps: 1) the query hits a router, 2) the router picks the most relevant expert slice, 3) only that slice computes, 4) results are fused into one answer.

Why the architecture matters is straightforward: running all 600 billion parameters for every question demands extreme hardware, while the sparse design spends energy only where needed. StepFun stresses agentic work, software engineering and financial expertise, with emphasis on long-horizon execution — the ability to carry a fifty-step code-base repair without dropping the thread. Picture carrying only the relevant shelf to the desk instead of the whole library. Preview is already available via API and open weights are dated October 15. The practical limit is physical: running 600-billion-parameter weights locally is unrealistic for most teams even with quantization, so the cloud API is the primary path.

MiniMax tells a two-track story. For M3.1, a hidden commit found in test files and traces inside the open-sourced code depot are read as near-launch signals, with late this month or early next month mentioned. In the August 26 interim results call the chief executive said M3.1, M3 Pro and H3.1 are close to completion. M3.1 is framed less as pure scale and more as reliability for real work: stability, output quality, inference efficiency and agent generalization, aiming to power broader infrastructure at scale. M3 Pro expands the narrative: a scale toward roughly three trillion parameters and an architecture step that pushes pre-training scale, training efficiency and extreme long-horizon tasks. It is less a wider highway and more a new multi-level interchange on top.

On Alibaba's side the Qwen family looks locked to the Apsara Conference calendar, scheduled September 22 to 24. Qwen 2.5 and Qwen 3 both debuted on that stage, which fuels expectation. An agenda session described as Advancing Qwen into the Agentic Era signals a flagship that lifts coding and everyday office work, introduced together with a new omni model and live translation across many languages. The same showcase lists Happy Oyster 2 world-model preview and audio models for ASR, TTS and real-time. In other words Qwen 4 is not a solo; it arrives as a bundle. This is not a single-product launch but an attempt to line office plus code plus audio plus world model on one stage.

The surprise is already shipping: Qwen-Image 2.1. A 7-billion-parameter visual backbone (DiT — Diffusion Transformer that builds an image by progressively denoising) that handles both generation and editing, supports up to ten reference images, high-fidelity edits and native RGBA (transparency channel) for clean transparent backgrounds. Community tests comparing it with Google's Nano Banana 2 (Gemini 3.1 Flash Image) describe it as neck and neck despite the tiny size; one reported score puts it at 60.28 versus 59.82. The license detail matters: a shift from Apache to research-only, which limits selling the raw output as product. For a marketing team this means you may self-host the weights and experiment freely, yet packaging the generated image as a product needs an extra permission layer.

Moonshot offers the most elegant teaser for Kimi K3.1: the official Kimi account posted the digits of pi after a leading 3.1 prefix, and the crowd solved it in one sentence — the next model is Kimi K3.1. No official release yet, but marketing heat has clearly risen. Read together, the week is clear: a US push to offer the same intelligence cheaper, and a Chinese push to run giant expert mixtures cheaper and ship the weights. Both answer the same question from opposite ends: will frontier intelligence become accessible to everyone, or stay as rented API? This week is one of the first where the second path, via open weights, closed noticeably on the first.

Visualization: nodesdaily AI

AI commentary

"What I see this week is not a single model race but two philosophies colliding: Anthropic tries to make the same intelligence cheaper and more reachable, while Chinese labs try to run giant expert mixtures cheaply and ship them as open weights."

AI assessment

Steel-manning the counter-case first: most of the excitement rests on unverified price and performance tables. The Tuesday window and Astra-level claim for Opus 5.5 remain speculation until an independent run exists; StepFun's 65% saving chart is its own measurement. The healthier test is to run the same task with the same prompt on two providers and compare invoice and quality side by side. The video is a crisp compilation, but a compilation is not proof on its own.

Limits are also clear. Running StepFun 600B locally is impractical for most teams; MiniMax M3 Pro's three-trillion target expands training scale as much as inference cost. Qwen-Image 2.1 is impressive at 7 billion, yet the RGBA and ten-reference support comes with a research-only license that brakes commercial use. MiniMax M3.1's emphasis on stability reads well on paper, but agent generalization is proven not in one long task but across dozens of tools and permission layers, and those public tests are not yet out.

Through a provenance lens each claim wants a different check. Anthropic's price only becomes real on the official price page, performance only on an independent replay of the same SVG and procedural prompts; StepFun's saving claim wants the invoice itself; Alibaba's Qwen 4 wants a dated announcement and model card on the Apsara stage; Kimi K3.1 wants a dated model identity after the pi teaser. A single point on an independent index (for instance Intelligence Index at 44) is a single measurement, not a guarantee of long-task success, which needs its own benchmark.

The practical takeaway splits cleanly. Teams running long agentic jobs and finance-heavy analysis via API can usefully try Step 5 Preview today, but waiting for the October 15 open-weight drop before locking production makes sense. For those sensitive to closed-frontier cost, the Opus 5.5 leak is a wait-and-see signal; even if Tuesday happens, first-day pricing and access tiers can wobble. For image work, self-hosting Qwen-Image 2.1 to experiment is low risk, but anyone planning to sell the output should clear the license first.

Sources

8 links; 2 of them also cited by 2 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

opus 5.5 · anthropic · qwen · kimi · minimax

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…