DeepSeek broke an unwritten rule with this release: the model called Flash is temporarily taking over work that used to go to the flagship Pro. Starting 04:00 UTC on September 14, 2026, Pro requests route to V4.1 Flash at Flash rates until the next Pro model lands. That is not a side option for cheap traffic; it is the main line wearing a Flash badge for a while.
The old hierarchy was simple. V4 Pro was the heavyweight at 1.6 trillion parameters with 49 billion active per token, the model you picked for maximum strength. V4.1 Flash arrives with a 552-billion-parameter backbone, which is smaller yet still enormous. The headline is therefore not the parameter count but how much of the model wakes up at each stage of the work.
The core idea is an asymmetric Causal Encoder-Decoder split. Roughly speaking, reading and writing are handled by different crews: about 8 billion active parameters organize the prompt, the codebase, and tool outputs while reading, then about 16 billion take over when the model writes answers and picks the next tool call. The mental image is one team sorting a huge case file and handing the distilled pieces to a stronger team that drafts the plan.
Memory is where the design pays rent. Long agent runs accumulate a growing record of processed context, and carrying that record gets expensive as windows stretch toward a million tokens. DeepSeek quotes around 890 bytes of global cache per token, about one quarter of the previous generation in active memory and one eighth in persistent storage. Smaller cache means less high-bandwidth memory per live run, less SSD per stored session, and less repeated work when the agent revisits earlier material.
Then come the tables. On Terminal-Bench 2.1 the video reports 90.6 for Flash against 89.1 for Claude Opus 5 and 88.8 for GPT-5.6 Sol; on DeepSWE it reports 74.2 against 74.0 and 73.0; on AutomationBench the gap widens to 54.8 against 50.3 and 45.8. Early independent chatter adds texture: one intelligence index scores Flash at 40 points above its own Pro at 36, with domestic rivals higher still, and a design-arena test puts it within a point and a half of a much pricier closed model at a fraction of the cost per finished design.
The model does not work alone, which is why the Harness update matters. DeepSeek frames the Harness as everything around the brain: tools, files, shell access, planning, subagents, sessions, storage, scheduling. The new version keeps expensive context alive when system instructions change mid-task, exposes cache-hit and token statistics more clearly, and refreshes optional subagent runtimes. Version 0.1.5 also leans into file uploads and sidebar previews, pushing the Harness from a runner toward a daily workbench.
Open weights plus MIT licensing round out the picture, with a catch the video states plainly. Developers gain inspection, supported inference stacks, and freedom from a single closed API, but the full backbone still has to live somewhere far beyond a standard laptop. So the decision splits three ways: API builders get a cheaper reader for long agents, infrastructure teams get control at the price of serious hardware, and anyone chasing a single champion for every task gets evidence of membership in the frontier conversation rather than proof of a crown. The thesis of this release is asymmetry: be cheap where reading allows it and strong where deciding demands it.
| Test | V4.1 Flash | Opus 5 | GPT-5.6 Sol |
|---|---|---|---|
| Terminal-Bench 2.1 | 90.6 | 89.1 | 88.8 |
| DeepSWE | 74.2 | 74.0 | 73.0 |
| AutomationBench | 54.8 | 50.3 | 45.8 |
AI commentary
"My read: the interesting move here is not a higher score, it is cheaper reading. If long agent runs are mostly context handling, spending less on reading and more on deciding is the right asymmetry — but I will only route production traffic after independent reruns confirm the vendor table."
AI assessment
The strongest objection writes itself: these are vendor-published tables released to justify retiring a flagship, and the same checkpoint reportedly spans 65.5 to 74.2 percent on DeepSWE depending on the harness around it. Independent signals are mixed so far: one index puts Flash at 40 points ahead of its own Pro at 36 but behind two domestic rivals, while other write-ups note weak spots on program-style and science-heavy sets. That does not make the release empty, but it makes the headline scores provisional rather than settled.
What the video does not test matters too: long-horizon stability across messy repos, behavior under shifting mid-task instructions over hours, real peak versus off-peak billing for sustained agents, and safety review data. The Harness update addresses part of this with cache-preserving instruction changes and better cache statistics, yet statistics visibility is not the same as a measured month of production incidents. I would also want memory and storage numbers from a neutral deployment, not only the one-quarter memory and one-eighth storage ratios from launch material.
The provenance shapes how I read every figure. DeepSeek published the wins while announcing that Pro requests move to Flash pricing until the next Pro arrives, with partners already lined up behind the new model. So my recheck list before any migration is short: rerun Terminal-Bench under a fixed harness version, compare the independent intelligence index rather than a single vendor table, confirm live provider prices for peak and off-peak windows, and read the Harness release notes instead of the launch prose. Numbers that survive that gauntlet earn traffic; the rest stay slides.
My practical take is threefold. If you build coding agents, Flash is worth a controlled trial because the cost structure of reading plus cache changes the unit economics of long runs. If you want to self-host, the MIT license gives control but the 552-billion-parameter backbone still demands serious accelerator hardware, so control is real and cheap it is not. And if you want one strongest model for everything, this release does not settle that; it earns Flash a seat at the frontier-agent table, nothing more and nothing less.
Sources
9 links; 3 of them also cited by 4 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube DeepSeek V4.1 Flash — episode video
- @deepseek.com https://www.deepseek.com/en/news/deepseek-v4-1-flash/
Also cited by: DeepSeek V4.1 Flash's Insane Architecture: Shared Memory That Shrinks KV Cache 437x · DeepSeek V4.1 Flash and the Split Brain: Challenging Frontier Models with 890-Byte Notes · AI Tier List Reset: GPT-6 Astra Takes the Crown as Subscription Math Rewrites the Ranks · DeepSeek V4.1 Flash: Cheap, Fast and Open — yet Rough on the Test Bench
- @huggingface.co https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
Also cited by: DeepSeek V4.1 Flash's Insane Architecture: Shared Memory That Shrinks KV Cache 437x
- @techtimes.com https://www.techtimes.com/articles/327163/20260910/deepseek-v41-flash-cuts-agent-memory-costs-fourfold-new-architecture.htm
- @orcarouter.ai https://www.orcarouter.ai/blog/deepseek-v4-1-new-base-model
Also cited by: DeepSeek V4.1 Flash: Cheap, Fast and Open — yet Rough on the Test Bench
- @technode.com https://technode.com/2026/09/10/deepseek-releases-harness-0-1-5-with-v4-1-flash-support-file-uploads-and-sidebar-previews/
- @kingy.ai https://kingy.ai/blog/deepseek-v4-1-flash-local-hardware-requirements/
- @decrypt.co https://decrypt.co/377917/deepseek-openai-gpt-6-astra-design-benchmark
- @gate.com https://www.gate.com/news/detail/deepseek-v41-flash-achieves-40-point-intelligence-score-costs-7x-less-than-24185950
deepseek · v4.1 flash · coding agents · benchmarks · open weights · harness