This is the host's first long briefing after a month away from his home setup, and it compresses eight separate leaks into one sweep: Gemini 4 Pro, GPT-6 Soul, Grok 4.7, Union Alpha, Typesafe's Jev, Google's Dream RSI framework, Projects in Claude Code, and Xiaomi's Mimo 2.6 training run. None of them arrived as a formal launch; each surfaced through ghost tests on Arena, quota listings on Google Cloud, free harness trials, or public research logs. For me that is the real story of the week: how models are tested has become as newsworthy as the models themselves.
Gemini 4 Pro Ghost-Testing on Arena
The clearest trace for Google is a checkpoint circulating on Arena under the name Gemini 3.8 Flash. When a user ran an infinite JSON stress test asking for a thousand planets with dense detail, the real Flash summarized and quit while this checkpoint kept producing non-repetitive, detailed records until it hit the browser's character limit and crashed the page. That stamina lines up with a leaked 256k output token limit attributed to Gemini 4 Pro, which strengthens the guess that this is the flagship checkpoint. The video also notes how to try it — sign in to Arena, open Battle mode and pick Code — and stresses that even at Arena's lower reasoning setting the outputs look surprisingly strong, implying launch-time capacity could be even higher.
The checkpoint's creative outputs form the core of the briefing. A smooth Mario Kart-style game with world generation, a Formula 1 car rendered in 3D with varied lighting and camera angles, a voxel-art temple complete with falling cherry-blossom leaves, fish in the river and a full day cycle, and a pure SVG PlayStation 5 — all are presented as unusually polished for Arena. The most talked-about example is an animated SVG of a Nintendo Switch where the controllers snap into the body and the screen boots with the logo; the host says Gemini 4 Pro clearly bested Opus 5 on that prompt. A ten-minute SVG generation that shows depth, multiple viewpoints, the mute-button glow and even a skeleton layer, plus a side-by-side pelican-riding-a-bicycle prompt where GPT-6 Astra still leads on visual depth, together make the case that the base reasoning level of the ghost checkpoint is already competitive with top models — fueling the Google comeback narrative while still leaving independent verification open.
GPT-6 Soul: Cutoff, Tiers, and Early Feel
On the OpenAI side, the briefing frames GPT-6 Soul as a near-term efficiency tier ahead of Dev Day. Its knowledge cutoff is described as April 30, the same date attributed to GPT-6 Astra, hinting at a shared lineage, with lower tiers possibly called Luna and Terra. Soul is said to have entered gray testing inside the ChatGPT web app, where early testers report richer formatting and noticeably better memory and personalization. A PS5 controller SVG generated with high reasoning effort in about three minutes and a 3D house interior coded in a 3GS-like environment are offered as signs that even the cheaper tier approaches Astra-like quality. The host also relays claims that Soul beat the unreleased Claude Opus 5.2 on a couple of prompts while using fewer tokens — but cautions that these are unverified early impressions, not benchmarks, and that pricing versus Astra and real speed will only be clear at launch.
Grok 4.7 Surfaces in Cloud Quotas
For Grok, the signal is its appearance inside Google Cloud quotas and system limits. Combined with Elon's earlier 'coming soon' note, the listing is read as a release being days away, even if the original ten-day window has already slipped. Early numbers cited in the video put Grok 4.7 at 6.4 on Colony Bench Mega, a clear jump over Gemini 3.8 Flash at 2.7 and GPT-5.6 Soul at 2.6, yet still well behind GLM 5.3, DeepSeek 4.1 Flash, GPT-6 Astra Pro and Claude Fable 5. On the CUDA side the picture is tighter: 9.5 on the GLM 5.2 fused test, just above Grok 4.6's 9.4, with small gains on other CUDA workloads. The host's take is balanced — a real step over 4.6, but often marginal — and suggests waiting for independent reruns before extrapolating.
Union Alpha Unmasked: Pareto 269 and the Harness Truth
The week's biggest mystery, Union Alpha, is resolved right after the main edit. Initially presented as a stealth model available for free via OpenCode, OpenRouter and Klein, with a 256k context window, multimodal support and agentic coding focus, it was credited with near-Astra and near-Opus 5 performance at roughly one-eighteenth of the expected cost. The follow-up reveal identifies it as Pareto 269 from Unbiased — not a single model but a system that routes work across several open and frontier models, in the spirit of the earlier Sukuna setup. Listed pricing is $2.50 per million input tokens and $0.25 for cached input, $7.50 for output, claimed to be less than a quarter of Astra's list price. That routing explains why it looked so close to frontier models on some tasks: the harness forwards work to those models when needed. My read is that Union Alpha should be judged as an orchestration product rather than a weight; the critical question is how transparent the routing logic is, not just the headline price.
Typesafe's Jev, developed with Dio, sits on a different axis entirely. Unlike classic large language models that generate text token by token, Jev takes unstructured input and a predefined set of possible outputs and returns structured decisions with probabilities, computing everything in parallel. Its training approach is called RLCD, reinforcement learning for calibrated decisions. The company's own benchmarks claim 20 to 200 times quicker while costing 40 to 400 times less across the decision workflows they tested, with pricing at four cents per million input tokens and free outputs — claims that will need independent validation. The key limitation is explicit: Jev does not generate normal text, so it is not a replacement for GPT-6 Astra or any Claude model; instead it is positioned as a companion that handles thousands of cheap, fast decisions around a frontier model — routing, classification, judging outputs, or picking the next agent action.
Dream RSI, Claude Projects, and Xiaomi's Public RL Run
Google DeepMind's Dream RSI is framed as a loop that is close to, but not yet, closed. Model weights do not directly improve; the agent that explores the policy learns from its own history of discovery attempts and gets better at searching for solutions over time. The video says the loop was demonstrated across algorithms, GPU kernel engineering and model development, suggesting the policy improvement generalizes beyond a single task. The caveat matters: this is recursive improvement of the search policy, not recursive self-improvement of weights, so while the label RSI is exciting, the result is a controlled step toward it rather than full autonomous weight-level evolution.
Claude's new Projects inside Claude Code is presented as a practical upgrade for on-the-go work: give Claude several tasks in one conversation and it routes them into parallel threads that keep running after you close your laptop, with shared memory and files staying in sync across threads. It is currently in limited beta for Pro and Max subscribers, with a wider rollout anticipated — potentially together with Opus 5.2. Xiaomi's Mimo 2.6, after nearly six months of silence, is described as mid-way through a massive reinforcement learning run that scales RL compute, multitask agentic environments and overall compute, generating roughly two billion tokens per training step across thousands of parallel rollouts. The team is said to be streaming the run publicly and to open-source more training details in the coming weeks, though the model itself is still one to two months away.
The briefing closes with a short robot demo where GPT-6 Astra powers dance steps, a symbolic reminder that large models now stretch from SVG craftsmanship to physical choreography. Taken together with the host's plugs for his newsletter, benchmark and vibe-coding platform, the picture is clear: Google is quietly auditioning a comeback via Arena, OpenAI is layering a family to lower cost, xAI is warming a release through cloud quotas, and Xiaomi is earning trust by opening the training log. My lesson from the week is to read it not as a single winner but as competing transparency strategies — whose ghost test is most honest, whose quota leaked earliest, whose log teaches most — that is where the real contest is playing out.
Key moments
- Opening — eight headlines in one briefing
- Gemini 4 Pro ghost test and 256k claim
- Mario Kart and Formula 1 3D generations
- Voxel temple with day cycle detail
- PS5 pure SVG — ten-minute depth
- Pelican on a bike — side by side with Astra
- GPT-6 Soul cutoff and gray-test traces
- Grok 4.7 Cloud quota and early scores
- Union Alpha and the Pareto 269 reveal
- Jev, Dream RSI, Claude Projects and Mimo 2.6
AI commentary
"What struck me this week is not a single model but how we learn about them: ghost runs on Arena, quota leaks on Google Cloud, and public RL logs all turned the testing infrastructure itself into the headline rather than any launch date."
AI assessment
The strongest steelman against this briefing is to temper the leak quality. Ghost tests on Arena run at a lower reasoning level and you land on a checkpoint by chance; the polished SVG or 3D we see may be a cherry-picked best case, not the average. Numbers like a 256k output limit come from labels, not independent measurement. Reading a full Google comeback from a single Switch animation therefore overreaches; what matters is repeatability across runs and the failure rate, not the highlight reel.
The second gap is method and time limits. The thousand-planet JSON test shows stamina but without metrics for diversity, coherence and true non-repetition it could also be a memorization display. The PS5 skeleton SVG is impressive, yet it took ten minutes and crashed a browser at the character limit — costs that may not be acceptable in real use and that raise practical questions about latency and stability. Grok 4.7's Colony Bench and CUDA scores are likewise early and narrow; the picture could shift on other domains, long context, or safety tests. The prudent step is to treat these as signals, not verdicts, until independent reruns appear.
On motive and verifiability, each headline carries expectation management: Gemini traces rely on anonymous user shares, GPT-6 Soul on gray-testing observations, Grok on a Cloud listing, and Union Alpha on a harness marketed at first as a free model. Pareto 269's pricing sounds very attractive, but if part of the work is routed to frontier models, the price you pay and the cost of the model that actually did the work are not the same; without a transparent routing log, cost comparisons mislead. Jev's claim of a 20 to 200-fold speed gain and a 40 to 400-fold cost reduction is also measured on the company's own workflows and needs independent benchmarking before generalizing.
My practical take is to make three habits routine right now: probe the ghost checkpoint on Arena with the same prompt a few times, watch Cloud quotas for Grok, and read Xiaomi's public RL logs. But do not bet on a single video's brightest sample at decision time. If you are prototyping, poke the Gemini ghost with the same SVG prompt repeatedly, measure Soul's memory and formatting gains on your own data, wait for Grok 4.7's official release, budget Union Alpha as orchestration rather than a weight, and pilot Jev only as a decision layer knowing it cannot generate text.
Sources
11 links; 3 of them also cited by 19 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube YouTube — Gemini 4 Pro Leaks and Weekly AI Briefing
- @officechai.com https://officechai.com/ai/gemini-4-0-pro-arena/
Also cited by: The costume on Arena: Gemini 4 traces leaking under the Gemini 3.8 Flash label
- @biggo.com https://finance.biggo.com/news/66df2f89-c067-43d7-b888-94deb071f472
- @openai.com https://openai.com/index/gpt-6-astra/
Also cited by: Brain and Body: A Single-Screen Agent Setup with GPT-6 Astra on Hermes · Space Bunny Alpha: Inside OpenRouter's Free Anonymous AI Experiment · Robot-Use Agents: Why General-Purpose Models May Win Robot Control · Building a Productive Card Collection App in Minutes with Base44 and GPT-6 Astra · Price War Begins: GPT-6 Sol and Luna Halve Model Costs · Gemini 4 Leak? 10 Interactive 3D Tests Against GPT-6 Astra and Fable 5.1 · 10,000 Agents, 88 Hours, $1 Million: AI Mastermind #39 From Code to Cash to Autonomy · Cloning a Channel With One Prompt: The $33K Video Factory Built on GPT-6 Astra and Higgsfield · GPT-6 Astra Guide: How Horizontal Power Turns the Model Into Work Done · From Hand Sketch to Realistic Villa: A Showcase Video with GPT-6 Astra and Higgsfield MCP · The $500-a-Day Claim With GPT-6 Astra: Building Three Business Models End to End · When Agents Take the Job: Investing in the Stack After GPT-6 Astra (+5)
- @ccleaks.com https://ccleaks.com/news/openai-gpt-6-astra-launch-sep-2026
- @digg.com https://digg.com/ai/wityfrzg
- @gigazine.net https://gigazine.net/gsc_news/en/20260918-union-alpha/
- @theregister.com https://www.theregister.com/ai-and-ml/2026/09/16/typesafe-ai-debuts-model-for-machines-that-plays-doom/
- @aimodeling.com https://www.aimodeling.com/en/news/slug/dream-rsi-replay-simulator
Also cited by: Did Google Just Dream Its Way to Self-Improvement? Inside Dream RSI and the Intelligence Explosion Debate
- @forkast.news https://forkast.news/xiaomi-mimo-v2-6-breaks-cover-a-1t-class-chinese-lab-trains-in-public/
- @digitaltoday.co.kr https://www.digitaltoday.co.kr/en/view/103667/musk-admits-grok47-not-better-than-rivals-teases-agi-with-grok5
gemini 4 pro · gpt-6 soul · grok 4.7 · union alpha · pareto 269 · dream rsi · mimo 2.6