Back to feed

Xiaomi MiMo v2.6: The $3 Million Live RL Run Behind the Best Open Model Claim

Xiaomi’s MiMo v2.6 family — Pro and Flash — chases the top open-model slot with 46 points on the Artificial Analysis Intelligence Index after a 6-day live reinforcement-learning run; weights and training code are on Hugging Face, while the 310-billion-parameter Flash sibling runs on 15 billion active parameters, fitting two DGX boxes.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — VSh8M3CUP88
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

Open models catching the frontier is not a press-release slogan here but a live demo: Xiaomi’s MiMo v2.6 arrives with two siblings — Pro and Flash — and, unusually, a streamed reinforcement-learning run with a $3–3.5 million price tag shown on screen. The narrator, who says he rarely covers Xiaomi, flags the pair as GPT-5.6-class on several agent benches while also publishing weights and training code together with competitive 3D scene generation — a combination that signals how far China’s open ecosystem has stretched when training itself becomes the show.

Why a live RL run matters

During reinforcement learning the team flipped the usual playbook and shared six days of live logs instead of a post-hoc chart; you can watch performance climb day by day, a transparency choice the official MiMo note frames as the result of half a year of research buildup. Described internally as one of the largest compute allocations for a domestic open model, the run is presented less as a flex and more as a lab notebook — like a marathon with split times on a public scoreboard, turning cost into methodology rather than marketing.

The family itself is two native multimodal models. The video frames Pro at around one trillion total parameters with about two billion active per token, and Flash at roughly 310 billion total with 15 billion active — numbers that echo the documented MiMo MoE design (mixture-of-experts — only the relevant experts fire per token, like calling a subset of an orchestra) that previously scaled to one trillion total with 42 billion active and one-million-token context in V2.5. Weights live on a dedicated Hugging Face collection, three variants are already on OpenRouter, and a distilled 9-billion-parameter coder (Quinn 9B) rounds out the lineup for on-device experiments.

Architecture and hardware still heavy

Sparsity explains the scale: MoE keeps per-token compute modest while total capacity stays large, and hybrid attention — sliding-window local plus sparse global attention, extended with multi-token prediction heads — holds long context without exploding KV-cache memory (about a 40% saving versus full attention in the documented design). Even so, cache grows with context because both weights and context traces share VRAM. A chart built with Claude’s help in the video suggests Flash can fit two DGX boxes in full quantization, while the trillion-scale Pro wants roughly eight H100s in full precision — quantization (packing numbers into fewer bits, like squeezing a wardrobe into a cabin bag) trims that, but trades a little fidelity for footprint.

On performance the narrator avoids single-score hype. In the Artificial Analysis Intelligence Index, MiMo-V2.6-Pro scores 46, overtaking Kimi K3 and Qwen3 Max as the strongest open model in that snapshot, though still behind the top closed systems. The less saturated Terminal-Bench 4.0 is highlighted as the better agent yardstick, where both Pro and Flash sit near GPT-5.6 Sol level and hold up in the narrator’s own checks. On the cost frontier, Pro is framed as Pareto-efficient: llm-stats pricing around $0.43 input and $0.87 output per million tokens undercuts several peers by multiples, implying more intelligence per dollar.

Access is kept compatible with the open claim: open weights on Hugging Face, API on Xiaomi’s Open Platform and OpenRouter using lowercase `mimo-v2.6-pro` / `flash` identifiers, with pricing carried over from V2.5 and an UltraSpeed mode promising up to 20× throughput. Context stays generous at one million tokens in, 128k out, with cache-hit input pricing falling to thousandths of a cent per million when prompts hit the cache — plus a Team Token Plan and Batch API now powered by V2.6.

In-harness verification, not one-shots

The narrator’s favorite stress test runs inside the DeepSeek harness — a standardized rig where agents browse, act and are scored on unseen tooling. Asked to research itself and build a launch page with verification, the model burns roughly 45 million tokens, but a 98% cache-hit rate keeps the bill flat, like reusing photocopies instead of reprinting the book. In multi-modal (text plus image) medium mode it compulsively searches, screenshots and re-checks, and the first-pass page impresses by mirroring the blog’s themes and even animations for both Pro and Flash — framed as evidence that verified loops, not single shots, reflect real engineering.

A sponsored interlude splits search into two worlds: general web search, which returns summaries optimized for relevance and often leans on blogs and SEO pages, versus Firecrawl’s Developer Index, which indexes live GitHub issues, merged pull requests, changelogs and raw docs to return primary sources. In an XM 7-to-8 migration example, both approaches find the syntax fix, but the developer index surfaces the merged PRs and catches a wildcard trap where blogs pin the wrong version — the point being that for code, the primary source is the answer, and agents should load it directly via MCP with 500 free credits to start.

Capability demos then probe whether verification survives contact with reality: a live API call tracks the International Space Station accurately on a day-night map, with a checkerboard overlay tracing the orbit where stars would be, and the agent even starts fetching ISS imagery to model the station in Blender before being stopped. Blender assets are created and animated via code, procedural 3D via 3D Gaussian Splatting is judged decent but not best-in-class — DeepSeek 4.1 Flash still leads in that harness — and spatial reasoning is flagged as merely okay. Games and visuals are now table stakes; the test is whether the agent sustains a verified loop rather than a flashy one-shot.

The closing lens widens from tool use to science: the team’s experiments report formal math solved with Lean, plus early signals in materials science and biochemistry, notably without domain-specific reinforcement learning. The narrator is careful to mark his non-expertise but frames the implication plainly — coding is now common ground for many models, the next frontier is the physical world and research, and early evidence that open-weight models are becoming useful there is what excites him, provided they remain genuinely open and reproducible rather than open in name only.

Visualization: nodesdaily AI

AI commentary

"My take is that MiMo v2.6’s real signal is the method, not the score: streaming six days of live training instead of dropping a trailer turns opacity into accountability, and framing a $3.5 million bill as transparency lets frontier-level performance be discussed alongside bargain pricing."

AI assessment

Steel-manned, MiMo v2.6 argues that open weights plus training code plus a live RL notebook can bring reproducibility close to the frontier without a frontier bill; the 46-point index and a strong Terminal-Bench 4.0 posture give that claim a measurable anchor rather than a slogan. Paired with verified loops in a harness, the implication is that the model can be a reliable sub-agent for well-scoped tasks even if it is not yet the best orchestrator for everything.

The limits sit in the footnotes: Pro’s need for roughly eight H100s in full precision and a KV cache that swells with context mean open does not mean locally runnable for most teams; the model is enterprise-open, not laptop-open. The Firecrawl comparison is a single XM 7→8 window and carries a sponsorship lens without an independent ablation, and the video itself grades 3D and spatial reasoning as merely decent, trailing DeepSeek 4.1 Flash in the same harness.

On provenance, pricing and benchmark entries are verifiable on Xiaomi’s pricing page, llm-stats and Artificial Analysis, but the live-RL narrative rests on a single org’s logs with no independent reproduction yet; the promised separate test of the distilled 9B coder also signals that evidence is still arriving in installments. Independent reruns of the index and Terminal-Bench will be the real litmus test.

Practically, Flash is attractive for high-frequency agent calls, long-context document pipelines and batch work where cost per intelligence matters; for heavy reasoning and frontier research, Pro may still trail the top closed models. Base your call on your own cache-hit rate — not the 98% showcase — and on a hardware bill that includes cache, then A/B against a closed baseline before switching.

Sources

7 links; 2 of them also cited by 1 other story. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

xiaomi · mimo v2.6 · open model · reinforcement learning · ai · agent · 3d

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…