Open models catching the frontier is not a press-release slogan here but a live demo: Xiaomi’s MiMo v2.6 arrives with two siblings — Pro and Flash — and, unusually, a streamed reinforcement-learning run with a $3–3.5 million price tag shown on screen. The narrator, who says he rarely covers Xiaomi, flags the pair as GPT-5.6-class on several agent benches while also publishing weights and training code together with competitive 3D scene generation — a combination that signals how far China’s open ecosystem has stretched when training itself becomes the show.
Why a live RL run matters
During reinforcement learning the team flipped the usual playbook and shared six days of live logs instead of a post-hoc chart; you can watch performance climb day by day, a transparency choice the official MiMo note frames as the result of half a year of research buildup. Described internally as one of the largest compute allocations for a domestic open model, the run is presented less as a flex and more as a lab notebook — like a marathon with split times on a public scoreboard, turning cost into methodology rather than marketing.
The family itself is two native multimodal models. The video frames Pro at around one trillion total parameters with about two billion active per token, and Flash at roughly 310 billion total with 15 billion active — numbers that echo the documented MiMo MoE design (mixture-of-experts — only the relevant experts fire per token, like calling a subset of an orchestra) that previously scaled to one trillion total with 42 billion active and one-million-token context in V2.5. Weights live on a dedicated Hugging Face collection, three variants are already on OpenRouter, and a distilled 9-billion-parameter coder (Quinn 9B) rounds out the lineup for on-device experiments.
Architecture and hardware still heavy
Sparsity explains the scale: MoE keeps per-token compute modest while total capacity stays large, and hybrid attention — sliding-window local plus sparse global attention, extended with multi-token prediction heads — holds long context without exploding KV-cache memory (about a 40% saving versus full attention in the documented design). Even so, cache grows with context because both weights and context traces share VRAM. A chart built with Claude’s help in the video suggests Flash can fit two DGX boxes in full quantization, while the trillion-scale Pro wants roughly eight H100s in full precision — quantization (packing numbers into fewer bits, like squeezing a wardrobe into a cabin bag) trims that, but trades a little fidelity for footprint.
On performance the narrator avoids single-score hype. In the Artificial Analysis Intelligence Index, MiMo-V2.6-Pro scores 46, overtaking Kimi K3 and Qwen3 Max as the strongest open model in that snapshot, though still behind the top closed systems. The less saturated Terminal-Bench 4.0 is highlighted as the better agent yardstick, where both Pro and Flash sit near GPT-5.6 Sol level and hold up in the narrator’s own checks. On the cost frontier, Pro is framed as Pareto-efficient: llm-stats pricing around $0.43 input and $0.87 output per million tokens undercuts several peers by multiples, implying more intelligence per dollar.
Access is kept compatible with the open claim: open weights on Hugging Face, API on Xiaomi’s Open Platform and OpenRouter using lowercase `mimo-v2.6-pro` / `flash` identifiers, with pricing carried over from V2.5 and an UltraSpeed mode promising up to 20× throughput. Context stays generous at one million tokens in, 128k out, with cache-hit input pricing falling to thousandths of a cent per million when prompts hit the cache — plus a Team Token Plan and Batch API now powered by V2.6.
In-harness verification, not one-shots
The narrator’s favorite stress test runs inside the DeepSeek harness — a standardized rig where agents browse, act and are scored on unseen tooling. Asked to research itself and build a launch page with verification, the model burns roughly 45 million tokens, but a 98% cache-hit rate keeps the bill flat, like reusing photocopies instead of reprinting the book. In multi-modal (text plus image) medium mode it compulsively searches, screenshots and re-checks, and the first-pass page impresses by mirroring the blog’s themes and even animations for both Pro and Flash — framed as evidence that verified loops, not single shots, reflect real engineering.
A sponsored interlude splits search into two worlds: general web search, which returns summaries optimized for relevance and often leans on blogs and SEO pages, versus Firecrawl’s Developer Index, which indexes live GitHub issues, merged pull requests, changelogs and raw docs to return primary sources. In an XM 7-to-8 migration example, both approaches find the syntax fix, but the developer index surfaces the merged PRs and catches a wildcard trap where blogs pin the wrong version — the point being that for code, the primary source is the answer, and agents should load it directly via MCP with 500 free credits to start.
Capability demos then probe whether verification survives contact with reality: a live API call tracks the International Space Station accurately on a day-night map, with a checkerboard overlay tracing the orbit where stars would be, and the agent even starts fetching ISS imagery to model the station in Blender before being stopped. Blender assets are created and animated via code, procedural 3D via 3D Gaussian Splatting is judged decent but not best-in-class — DeepSeek 4.1 Flash still leads in that harness — and spatial reasoning is flagged as merely okay. Games and visuals are now table stakes; the test is whether the agent sustains a verified loop rather than a flashy one-shot.
The closing lens widens from tool use to science: the team’s experiments report formal math solved with Lean, plus early signals in materials science and biochemistry, notably without domain-specific reinforcement learning. The narrator is careful to mark his non-expertise but frames the implication plainly — coding is now common ground for many models, the next frontier is the physical world and research, and early evidence that open-weight models are becoming useful there is what excites him, provided they remain genuinely open and reproducible rather than open in name only.
AI commentary
"My take is that MiMo v2.6’s real signal is the method, not the score: streaming six days of live training instead of dropping a trailer turns opacity into accountability, and framing a $3.5 million bill as transparency lets frontier-level performance be discussed alongside bargain pricing."
AI assessment
Steel-manned, MiMo v2.6 argues that open weights plus training code plus a live RL notebook can bring reproducibility close to the frontier without a frontier bill; the 46-point index and a strong Terminal-Bench 4.0 posture give that claim a measurable anchor rather than a slogan. Paired with verified loops in a harness, the implication is that the model can be a reliable sub-agent for well-scoped tasks even if it is not yet the best orchestrator for everything.
The limits sit in the footnotes: Pro’s need for roughly eight H100s in full precision and a KV cache that swells with context mean open does not mean locally runnable for most teams; the model is enterprise-open, not laptop-open. The Firecrawl comparison is a single XM 7→8 window and carries a sponsorship lens without an independent ablation, and the video itself grades 3D and spatial reasoning as merely decent, trailing DeepSeek 4.1 Flash in the same harness.
On provenance, pricing and benchmark entries are verifiable on Xiaomi’s pricing page, llm-stats and Artificial Analysis, but the live-RL narrative rests on a single org’s logs with no independent reproduction yet; the promised separate test of the distilled 9B coder also signals that evidence is still arriving in installments. Independent reruns of the index and Terminal-Bench will be the real litmus test.
Practically, Flash is attractive for high-frequency agent calls, long-context document pipelines and batch work where cost per intelligence matters; for heavy reasoning and frontier research, Pro may still trail the top closed models. Base your call on your own cache-hit rate — not the 98% showcase — and on a hardware bill that includes cache, then A/B against a closed baseline before switching.
Sources
7 links; 2 of them also cited by 1 other story. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — Xiaomi MiMo v2.6 Launch
- @mimo.mi.com https://mimo.mi.com/docs/en-US/news/latest/v2-6
Also cited by: Xiaomi MiMo-V2.6 Pro Tops Open-Weights: 1M Context and 20x UltraSpeed at Record Price-Performance
- @huggingface.co https://huggingface.co/collections/XiaomiMiMo/mimo-v26
- @artificialanalysis.ai https://artificialanalysis.ai/models/mimo-v2-6-pro
Also cited by: Xiaomi MiMo-V2.6 Pro Tops Open-Weights: 1M Context and 20x UltraSpeed at Record Price-Performance
- @mimo.mi.com https://mimo.mi.com/docs/price/pay-as-you-go
- @xiaomi-mimo-ai.com https://www.xiaomi-mimo-ai.com/research/mimo-architecture.html
- @llm-stats.com https://llm-stats.com/models/compare/mimo-v2-6-pro-vs-mistral-large-3-2509
xiaomi · mimo v2.6 · open model · reinforcement learning · ai · agent · 3d