Back to feed

DeepSeek V4.1 Flash's Insane Architecture: Shared Memory That Shrinks KV Cache 437x

Covered by Two Minute Papers, DeepSeek V4.1 Flash packs 552B MoE parameters and 1M-token context while shrinking KV cache 437x vs V1 and 4x vs the prior Flash — via cross-layer shared memory, CSA2 and FP4 compression.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — vIHw_2VjSUw
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

DeepSeek V4.1 Flash sits at the center of Two Minute Papers' 'Insane New Architecture' episode, which opens with sheer speed and two bold claims: on some benchmarks Flash can edge past Claude Opus-class models and Kimi K3, and it reliably beats DeepSeek V4 itself. The host frames cost to make the point tangible — running V4 Pro locally takes roughly $300k in hardware, while Flash does the same work for about a quarter. The curve points toward a pocketable agent within a year, offered as cautious optimism rather than hype.

What got faster, what got smaller: the anatomy of 437x

The most striking idea is that the model is huge and small at once. Huge because the backbone is 552 billion parameters with up to one million tokens of context — the video flags 'more than 500 billion.' Small because the KV cache that holds the context collapses: 437x smaller than V1 from three years ago and 4x smaller than V4 Flash from just months ago. The paper abstract in arXiv 2609.19969 pins the number: global footprint down to 890 bytes per token, roughly one quarter of DeepSeek-V4-Flash in HBM. With SWA Bounded Replay, the persistent footprint on SSD or host memory falls to about one eighth of V4-Flash. For long-horizon, input-heavy agent workloads, that directly cuts the bottleneck of compute, storage and bandwidth.

How is the shrink done? The video names it: Compressed Sparse Attention 2, CSA2. In a classic Transformer each layer keeps its own KV memory; DeepSeek V4.1 Flash shares it across layers. Not everyone has to remember everything — there is one shared memory. The difficulty is that different layers need different views of the same history. The fix is a 40-layer Causal Encoder-Decoder, CED: a 20-layer causal encoder projects a global memory from its final hidden states, and a 20-layer decoder reads from that global memory instead of deriving KV from its own states. Add FP4 KV caching and both the HBM-resident global cache and the SSD-resident persistent cache drop sharply.

Architecture numbers: 8B prefill, 16B decode, 45T tokens

The asymmetry matters for agent economics. The model activates 16B parameters per token during decode but only 8B during prefill, which halves prefill cost for input-heavy jobs. DeepSeek reports streamlining the V4 architecture, adding efficient extensions, and pretraining on a 45T-token multimodal corpus followed by broad post-training for text and multimodal agent scenarios. The 10 September 2026 announcement on deepseek.com positions it as the smallest member of the new family with native visual understanding, promising faster inference and higher throughput. Checkpoints are open on Hugging Face.

The video does not stop at architecture; it shows two concrete capability probes. First, native visual understanding — feed an iconic game menu image and have the model write a game that reproduces it. Second, physics reproduction — on the honey-coiling paper, GPT-6 Astra nails the result as 'absolutely stunning,' Opus 5.1 gets the physics right with slightly weaker rendering, and Flash is framed as a few papers away from matching that level. The host stresses showing the truth beyond headlines and notes, with a wry aside, that YouTube has stopped recommending his physics sim papers.

The catch and the bill: a model that thinks a lot and burns tokens

The catch at the end rebalances the story. V4.1 Flash likes to think and burns a lot of tokens. It is not expensive per token, but the sheer volume is high. With hardware you can run it for free; without, Lambda's GPU cloud is presented as the practical bridge — the host notes he uses Lambda to reproduce papers in minutes and to test ideas quickly. Every new DeepSeek paper, the argument goes, lowers inference cost for everyone, which directly helps doctors and scientists. Open weights plus an open paper keep the 'open science for the win' thread intact.

Stepping back, the insanity of V4.1 Flash is not raw muscle but memory economics. Cross-layer sharing, encoder-to-decoder projection and FP4 compression together shrink a one-million-token context to an 890-byte per-token trail while pushing performance above the baseline. Cutting a $300k local barrier to a quarter and halving prefill activation turns the 'in your pocket in a year' line from speculation into an engineering roadmap. By telling the shrink on two scales at once — 437x versus V1 and 4x versus V4 Flash — the video gives a compact frame for anyone tracking architecture.

ModelGlobal KV / tokenPersistent KVArchitecture
V4.1-Flash890 bytes1/8 of V4-FlashCED 20+20, CSA2+FP4
V4-Flash4x larger1x (ref)Standard MoE
V1 (3 yrs ago)437x larger—First gen

Key moments

  1. Intro — why V4.1 Flash excites
  2. Performance surprise — beating V4 and Opus-class
  3. Visual understanding — from menu image to game
  4. The shrink secret — CSA2 and shared memory
  5. Inside the architecture — CED 20+20 and FP4
  6. Physics demo — honey coiling, Astra vs Opus
  7. The catch — thinking hard, burning tokens

AI commentary

"What struck me most here is the inversion of 'bigger model' orthodoxy. DeepSeek grew the model while shrinking the memory — for me, that's the real threshold for agents to fit in our pockets."

AI assessment

Steel-manned, DeepSeek hits agent economics where it hurts. Long-context cost has long been stuck on 'prefill is expensive, KV is bloated, HBM/SSD is the bottleneck'; CED plus CSA2 plus FP4 loosens that knot at once, and does so with a measurable 890 bytes per token, not a slogan. The 8B prefill versus 16B decode split is one of the rare architectural wins that directly lowers the bill for input-heavy, real-world agents.

Limitations matter. The video is short and single-source; performance lines are hedged as 'on some tests' without the dataset, tolerance or statistical spread on screen. Cost anchors like $300k and a quarter swing quickly with hardware and memory configuration — viewers should not budget from the anecdote alone. The honey-coiling physics demo is striking, but one paper reproduction is not a generality proof; different physics and domains may not transfer with the same fidelity.

On stakes and verifiability, arXiv 2609.19969 and the Hugging Face checkpoints strengthen the base, and the 10 September 2026 DeepSeek announcement confirms the architecture and 45T-token pretraining. Still, how smoothly SWA Bounded Replay behaves in production, where FP4 KV starts to bite quality at very long context, and which workloads truly sustain 1M tokens await independent reproduction. The token-burn side also needs transparency: even if unit cost is low, total cost fluctuates as thinking traces lengthen.

Practically, teams running long-context agents, RAG or multimodal assistants should treat V4.1 Flash as a serious trial for memory and prefill budgets — locally if you have hardware, on rented GPU like Lambda if not. For short-context, latency-sensitive chat the gain is narrower, and the 'likes to think' character means you should account per response for tokens. Open weights plus an open paper make the trial low-friction — measure on your own dataset before scaling.

Sources

6 links; 2 of them also cited by 4 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

deepseek · artificial intelligence · kv cache · csa2 · moe · nodesdaily

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…