Over the weekend a model listed as Gemini 3.8 Flash on Arena drew attention because it behaved nothing like the public build Google shipped on September 2nd. The known 3.8 Flash is a speed tier with approximate geometry, while this anonymous entry went quiet for about 10 minutes on a raw SVG PlayStation 5 request and then dropped thousands of lines capturing curves, shadows and even the small optical drive accent in one go. Given that the official Flash answers the same request in 10 to 15 seconds with rough shapes, observers framed the choice as either the slowest Flash ever or a costume thrown over something larger.
How the costume test was spotted on Arena
Arena itself began in 2023 from a UC Berkeley team as a blind test where two models answer side by side without branding and users vote, giving labs clean signal free of halo effects, and such anonymous runs are routine with Google often circulating Gemini checkpoints that way. A mid-August anonymous model called Ox Alpha surged into the top three within 48 hours and remains unclaimed to this day. What stood out here was not anonymity but reuse of a shipped model's name, which reads less like stealth and more like camouflage.
Why visual fidelity reads as a jump
Precise spatial coordinate work in vector code is a notorious failure point where most models return roughly aligned shapes, broken paths and simplified geometry, so the leaked side-by-sides carried weight. Voxel pagodas, BMW M4 renders, and the classic pelican riding a bicycle under a twilight sky full of stars arrived without clipping or misaligned shapes, while a variant handled day-to-night grading, working headlights, anatomical labels on the bird and cadence control in one prompt. Developer Pankaj Kumar noted the model parses long multi-part instructions in one pass without four rounds of babysitting, and developer B's 14-minute sketch-themed showcase where scrolling darkens the lines as a concept embodied the same leap.
The long pause is described not as struggle but as heavy test-time compute. The older Flash approach ships basic shapes and approximate coordinate math instantly, while this one plans, writes, executes, debugs and polishes internally before streaming anything. That gap reflects a different design philosophy for precision placement and advanced voxel and lighting math, unlikely to ship by accident under a speed badge.
Web builds and games in one shot
The creative queue kept growing. A retrofuturism-tinged cyberpunk hero layered a dimensional grid with a real-time metrics board and a 3D object in the header without visual collapse. An Airbus H145 in about 10 minutes, a pixel pagoda that feels present, and a flight simulation clearly ahead of official 3.8 Flash were each delivered in one pass. Taken together, sharing a name while running laps around the official build argues they are not the same model.
Games pushed the scale further. A Minecraft-style build whose completion from page architecture to individual modules is placed on par with published Astra results and a 3D cart racer said to look more modern than KartRider in its prime both emerged with complete interactive logic in a short window. A simple SVG cat test flagged by developer Vidi, where a competing frontier model returned a tiger-like shape, underscored the spatial fidelity point rather than a joke.
Leaked scores and the pricing shock
The viral piece is the leaked chart, explicitly held loosely because nothing is confirmed. It claims roughly 88 percent on Deep SWE 1.1 for agentic coding about two points clear of Astra, the only 2064 Elo on GDP Eval-A-V2 for real-world knowledge work, 95.3 percent first place on Terminal Bench 2.1 for terminal coding, and 86.8 percent on OSWorld 2.0 for computer use ahead of both Astra and Fable 5.1. If accurate, the sweep across coding, agents, reasoning and operation would be broad.
The pricing rumor may bite harder. The chatter is $2.25 per million input tokens and $11.25 per million output, which against Astra launched September 3rd and Fable 5.1 September 1st at $10 and $50 would be exactly 22.5 percent of the frontier tier. Against Google's own Gemini 3.8 Flash promotional $0.75 and $3.75 through December 2026 it is roughly triple, fitting a pro tier expectation and quietly supporting the not-Flash argument. Until Google posts rates, this remains speculation.
Back end, a tough year and RSI traces
The back end story comes from ghost-routed calls. Trackers including Qwena report a 10 million token input ceiling and a 256,000 token output ceiling versus the official 1,048,576 and 65,536, about 9.5 times input and nearly four times output, which next to the typical 4K to 16K generation cap would allow entire repositories or long novels in one continuous pass. Added claims include native cross-session memory persisting across threads without a bolt-on vector store, direct internet access without an API, sandbox traps that auto-route untrusted code into isolated synthetic environments, and native robotic motor control translation pointing toward unified multimodal models meant to run on physical hardware. All remain unverified and as of mid-September Google has published no model card, parameter count or endpoint for Gemini 4 or 4 Pro.
Context makes the speculation legible because Google has had a rough run at the top. The last real pro update was Gemini 3.1 Pro on February 19th, about seven months ago, the Gemini 3.5 Pro previewed at I/O for June slipped repeatedly and in August reports from Semi Analysis and The Wall Street Journal said it was killed because the pro model was not improving faster than Flash and the architecture was scrapped. The confirmed line instead came on July 21st when Google's post announcing Gemini 3.6 Flash said its most ambitious pre-training effort yet had begun for Gemini 4, echoed two days later by Pichai on Alphabet's Q2 call that Google needs Gemini 4 to stay competitive at the frontier while pre-training evaluations looked strong and post-training proceeds.
That leads to the whisper shadowing the year, RSI for recursive self-improvement. DeepMind Chief Strategy Officer Jasjit Singh at the Agentech AI Summit in Berkeley said RSI has become core to the industry's capital expenditure thesis while admitting current AI revenue does not support spending at this scale, with the bet that the capability curve keeps climbing past exponential. The Information places realistic RSI around 2027 to 2028, while the rumor mill claims Gemini 4 finished pre-training early because DeepMind closed an RSI loop during training. Research published on September 14th called Dream RSI on agents that continually improve search strategies during exploration and carry experience forward does not prove the claim but confirms RSI sits at the center of the roadmap. The stakes are large with Alphabet around $4.23 trillion after crossing $4 trillion in January on Apple announcing Gemini for the new Siri, TPUs fabbed by TSMC and servers from Foxconn, Quanta, Inventec, Wistron and Wiwynn tying training scale to Taiwan order books, plus Gemini Sparks in Japan from July and steady enterprise onboarding onto Gemini Enterprise. Google has confirmed nothing about Arena, yet its pattern of public stress testing weeks before announcements fuels the consensus guess of a fall reveal alongside developer tooling most likely in October, so watching that Arena slot may signal the costume coming off before any press release.
Key moments
- Opening — the costume idea and 10-minute SVG
- What Arena is — UC Berkeley blind test
- PS5 and voxel — curves, shadows and alignment
- Pelican on a bicycle — spatial reasoning exam
- Compute policy — heavy test-time compute
- Showcase and Airbus — 14-minute sketch site, 10-minute helicopter
- Scores and pricing — leaked chart and 22.5 percent claim
- Back end and RSI — 10M window, Dream RSI Sep 14
AI commentary
"In my view this leak reveals a testing posture more than a model: a speed-badged model deliberately running slow reads like a signal about compute policy and release timing rather than raw aesthetics."
AI assessment
The strongest part of this leak is the pairing of a speed badge with deliberate slow generation that delivers surgical coordinate accuracy and single-pass parsing of long multi-part prompts as described by developers; if that heavy test-time compute is truly spent on planning and internal debugging, the philosophy gap between two models sharing one label becomes legible and points to a pro tier.
Its limits are that the entire chain remains unverified. The chart, the 10 million input and 256 thousand output ceilings, persistent memory and motor control all lack a Google model card or endpoint, screenshots and ghost route traces risk selective sharing, and questions remain about how 8 to 10 minute generations fit live product expectations, what pricing looks like after the December promo window, and how the sandbox for untrusted code would be governed at scale.
The takeaway is that the video frames the tough year correctly: the gap since February, the reported cancellation of 3.5 after failing to outpace Flash, and the shift to Gemini 4 pre-training from July are confirmed, while the claim that an RSI loop closed early and the Dream RSI paper showing search strategy evolution are adjacent signals that should not be conflated into causation and together create both opportunity and fragility if performance and cost improve at once.
Practically, checking that Arena slot week to week, rerunning the SVG cat and pelican prompts in your own tests, and treating leaked pricing as speculation until the October window and the December promo expiry is the healthiest stance; rather than locking a production plan before a model card lands, watch whether single-pass repository writes and 256 thousand token continuity hold stable and reproducible.
Sources
7 links; 4 of them also cited by 6 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — Gemini 4 RSI Leak: Costume Test on Arena
- @blog.google https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/
Also cited by: Gemini Managed Agents 09-2026 Leap: Anti-Gravity on Gemini 3.8 Flash Finally Works Like a Real Teammate · DeepSeek V4.1 Flash Meets Gemini 3.8 Flash: A Five-Prompt Coding Duel
- @arxiv.org https://arxiv.org/abs/2609.14858
Also cited by: Did Google Just Dream Its Way to Self-Improvement? Inside Dream RSI and the Intelligence Explosion Debate · Google's Dream RSI: Turning History into a Simulator for Recursive Self-Improvement
- @officechai.com https://officechai.com/ai/gemini-4-0-pro-arena/
Also cited by: Gemini 4 Pro Leaks, GPT-6 Soul in Testing: From Arena to Google Cloud, the Week's AI Shockwave
- @nokiapoweruser.com https://nokiapoweruser.com/gemini-4-pro-arena-ai-test-specs-leaks/
Also cited by: Gemini 4 Leak? 10 Interactive 3D Tests Against GPT-6 Astra and Fable 5.1
- @blog.google https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/
- @searchenginejournal.com https://www.searchenginejournal.com/pichai-says-google-needs-gemini-4-to-compete-at-the-frontier/583214/
gemini 4 · gemini 3.8 flash · arena · google deepmind · rsi · dream rsi · nodesdaily