Back to feed

DeepSeek V4.1 Flash: Cheap, Fast and Open — yet Rough on the Test Bench

On September 10 DeepSeek released V4.1 Flash, a 552-billion-parameter open-weights model with a million-token context, an MIT licence and benchmark scores that challenge closed giants. Matthew Berman's hands-on tests reveal the rougher side of the story.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — U-rsvXds9ck
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

DeepSeek formally launched V4.1 Flash on September 10: a 552-billion-parameter model with a million-token context and MIT-licensed open weights. Over the API it answers to deepseek-flash, and the legacy v4-flash names now route to it. The headline claim is bold: scores sit in the same league as the previous generation's closed giants, Opus 5 and GPT-5.6 Sol. The video frames this inside a familiar six-month rhythm: the frontier sprints first, open models match that level half a year later, and six months after that the power lands on local machines.

The published scoreboard has clear highlights: 74.2 on DeepSWE v1.1 for software engineering, a hair above Opus 5 at 74.0; 88.1 on CyberGym for security; 54.8 on AutomationBench. The table cuts both ways, though: on Terminal-Bench 3.0 it stops at 30.0, well behind Opus 5 at 43.3. And in a third-party design arena it reaches 98 percent of Astra's score while spending about 2.3 cents per task; Astra bills $1.61 for the same work.

The architecture is the most interesting part. The model uses a new 40-layer encoder-decoder skeleton: 20 layers read the input, 20 produce the output. Reading wakes only 8 billion parameters per token, generation 16 billion; the rest of the 552 billion sit idle at that moment. That is the mixture-of-experts idea in a nutshell: decide which mini-specialist owns the question and hand it over. A compressed sparse attention mechanism squeezes long-text costs further.

The video tests the speed claim live: a thousand-word sample streams in roughly six seconds, on the order of 200 tokens per second. The host compares the feeling to the early Groq excitement, and notes the model outruns the previous flagship V4 Pro. For agent workloads that keep reading fresh input, that tempo may matter more than any raw score, because waiting time bills just like compute does.

The memory win is even more striking. The context cache takes about a quarter of the fast memory and an eighth of the SSD space its predecessor needed. The company's comparison against V1 is startling: cache per token has shrunk 437-fold. For agents that grind through long contexts step by step, that means less GPU memory tied up and a smaller server bill.

The move lands in the middle of a genuine crunch. HBM3E spot prices are said to run four to five times long-term contract levels, with Samsung locking up 70 percent of capacity through 2031. On September 10 Reuters reported Chinese chipmakers raising processor prices as memory inflation bites. Pricier phones and computers with trimmed memory are smoke from the same fire; DeepSeek is looking for the answer in algorithms.

The price list is aggressive: 15 cents per million input tokens off-peak, 30 cents at peak. Cached input is nearly free at 0.3 cents per million. Output runs between 60 cents and $1.20 per million. The dual tariff exists to pull corporate workloads into the hours when GPUs sit idle. And since the weights are open, nobody has to hand their data to DeepSeek.

The video's economic thesis is crisp: the price of the best answer and the price of a good-enough answer differ by orders of magnitude, with the closed frontier charging on the order of $50 per million tokens while this class talks in cents. For high-volume work like websites and PDF pipelines, which the host puts at 95 percent of usage, this tier is more than capable. The company even documents its techniques in an open report a startup could build on.

The test bench is less generous than the scoreboard. The Rubik's cube simulation writes itself in 12 seconds and looks fine at first glance, but pressing scramble makes colors change on their own rather than from movement. The solve button produces no real solution; it replays the moves backwards to the start. The same failure shows up on the company's site and through Codex, and the host says no model has collapsed on this test for a year and a half.

Two more trials follow. In the team's Paintbench exercise the model turns a reference portrait into an abstract, detail-starved piece, with none of Astra's layered brushwork. The bullet-through-water 3D demo ships inside a fine app shell with mediocre physics. The verdict is fair: a cheap, fast, efficient workhorse; whether it beats GLM 5.3 is a question for another video.

Visualization: nodesdaily AI

Design-task cost: two cents vs $1.61

  • V4.1 Flash$0.023
  • GPT-6 Astra$1.61
OpenDesign arena, cost per task.
BenchmarkV4.1 FlashBest rival
DeepSWE v1.174.274.0 (Opus 5)
CyberGym88.184.5
Terminal-Bench 3.030.043.3 (Opus 5)

AI commentary

"My read: this release proves open models keep their six-months-behind rhythm; the real news is not the benchmark table but the architecture that shrinks the memory bill. The Rubik's cube fiasco honestly marks this class's limits."

AI assessment

Steel first: these scores come from the company's own report. On launch day no independent measurement existed; Artificial Analysis had not tested the model and its Hugging Face page showed no serving provider. On tables like Terminal-Bench 3.0, Opus 5 leads by a wide margin. So the 'frontier-level' verdict stands on selected scoreboards and should be read as an unverified claim.

The video's own trials are single-prompt, single-attempt demos; they carry no statistical weight. More importantly, the company draws the limits itself: on the hardest science-flavoured agent tasks and on image analysis the gap with giant models persists. A dim score like 15 on Exploit Gym is the swept corner of the table. The collapse in the play tests may be an early signal of those limits.

Two caveats on verifiability. First, the '.1' suffix suggests a tune-up, yet underneath sits a brand-new base model while a 1.6-trillion-parameter flagship retires. Second, the prices shine brightest at off-peak rates. Before deciding I would test my own workload at peak pricing and against independent measurements.

My practice shapes up like this: for long-context agent jobs and memory-constrained setups this model is the first candidate to try; where final accuracy is bought with money, the closed frontier still reigns. I will trust my own measurements over the price tag, and the open weights make that measurement risk-free.

Sources

10 links; 5 of them also cited by 5 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

deepseek · v4.1 flash · open weights · moe · benchmarks · hbm · ai

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…