Back to feed

125 Billion Parameters at Home: How Strata Hit 93 Tokens a Second on a 12 GB Card

Strata, a new inference engine, runs the 125-billion-parameter Qwen 3.8 Flash-Next at 93 tokens a second on a 12 GB gaming card. The honest version of the 6x headline is 2 to 2.5x, and the real limit is not the graphics card but 64 GB of system memory.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — m0VHx73SAG0
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

Picture a spinning cursor watched for three minutes while the fans howl and a giant model manages a single word. That was the speaker a few weeks ago: a weekend spent getting a 125-billion-parameter model running on a home PC with llama.cpp, rewarded with 15 tokens a second on a good day and first-word waits past three minutes on long prompts. Then Strata showed up on GitHub on September 24th, with numbers like 93 tokens a second on a 12-gigabyte card making the rounds. It looks like classic thumbnail math at first glance, but once the inner architecture of the model comes into focus it turns into an engineering story worth taking seriously.

At full precision Qwen 3.8 Flash-Next is a 354 GB model, roughly 29 times the memory of the card it runs on, like threading a cruise ship through a kitchen window. Yet Strata installs with a double click on Windows or a single script on Linux, figures out which files the system can handle, downloads about 70 GB and opens a chat window. The same engine serves the same API endpoints OpenAI and Anthropic use, so tools like Claude Code can point straight at the local machine. The project gathered nearly 4900 stars in eight days, with forum threads full of users posting their own speed figures and comparing notes. Screenshots and speed tables are published in the repository of the Strata project on GitHub, whose README figures show 94 tokens a second on a 12 GB card with the Q2_0 build.

125 billion parameters reading only 4 GB per token

Contrary to what most people assume, a 125-billion-parameter model does not touch 125 billion parameters per word. The 2-bit Q2 build that Strata runs fastest is a 66 GB file, and about half of it consists of experts, small specialist modules. Each of the 48 layers of the model holds 512 experts, 24576 blocks in total, and per token every layer picks just 10 of its 512 plus one shared expert. Across all layers 480 experts work per token, about two percent of the total. Of 34 GB of expert data, a single token reads roughly 0.66 GB, and with the always-on parts the total touched comes to only 4 GB out of a 66 GB file. The 125B-parameter, 6B-active figures of this mixture-of-experts design check out against the Qwen page on HuggingFace.

Nearly half the file, a 29 GB portion, forms one 51-billion-parameter lookup structure keyed to the three most recent tokens in context. Because the model knows exactly which rows it will need before it needs them, Strata simply leaves the whole table on the SSD and reads only 16 rows per token, about 23 kilobytes. The striking part is that the Qwen team designed this model for exactly this kind of duty: the model card says the embedding structure is more amenable to offloading. Thirty-six of the 48 layers use an attention type with constant memory no matter how long the conversation grows, while the remaining 12 layers look back at no more than about 48 positions. These four innovations are described in the same terms in the Qwen repository on GitHub and in the architecture write-up from AlibabaCloud, and the Qwen team counts this model as a preview of the Qwen4 effort.

When a token is generated, each of the 48 layers has the graphics card do two jobs: the attention mechanism and a router picking the 10 experts this token needs. Once the list exists, the card checks which of those experts it already holds in memory; a 12 GB card keeps about 4500 of the 24576 total. Experts it holds run immediately on the GPU while the missing ones get computed by the CPU in system memory at the same time, then the card merges the two streams and advances to the next layer. On the test machine one full pass takes about 34 milliseconds, 15 spent by the GPU and 13 by the CPU, so the two processors genuinely work in parallel, waiting on each other at every layer. The 4500 cached experts are not random; they load from a profile, a record of the most used experts across many conversations. Serving about half of all expert reads from the first message, the cache climbs toward a 72 percent hit rate as Strata learns what the current conversation keeps asking for. This expert cache is the whole engine; where older engines treat card memory as the scarce resource and system RAM as a last resort, Strata runs both as equal players.

The heart of the engine: expert cache and draft layer

The second trick, behind a large share of the real-world gains, is a small prediction layer trained alongside the main model. Acting like a cheap fast draft, this layer guesses up to three tokens ahead, then the big model runs all 48 layers in a single pass to check the guesses. Every accepted guess is kept plus one token of its own, so each expensive pass yields between 2.4 and 3.2 tokens on average instead of one. On the same machine output rose from 47 to 57 tokens a second with the trick off to the 82 to 92 band with it on. Whether output quality suffers has a clear answer in the team test: even with deliberately wrong guesses forced through, the text came out identical to running without prediction at all. This multi-token prediction carries one small asterisk, since on IQ-named builds the GPU and CPU round the same expert math very slightly differently, so one prompt can end two ways; the team check against llama.cpp agreed on 97.5 to 99 percent of positions. General MTP support landing in the May 2026 period is documented in the report from LLMRequirements.

The six-times figure deserves an honest look. In mid-September, before Strata existed, its author ran the 3-bit build through a llama.cpp-based tool on the same PC later used for the headline numbers, an RTX 5070 with a six-core CPU and 64 GB of DDR5, and got 15 tokens a second, with a 20000-token prompt taking 3 minutes 11 seconds before the first word. By September 18th the same setup was tuned to 29 tokens a second. The 93 headline value of the Strata project was measured on the 2-bit build; dividing 93 by the slowest 3-bit run gives 6.2. Compare 3-bit to 3-bit and the ratio sits at 2.1. In what is probably the fairest test so far, a user result on an RTX 3090 put current llama.cpp at 31 tokens against 74 for Strata, a 2.4x gap, with the guess-and-check equivalent switched off on the llama.cpp side. Nobody lied; the slowest baseline was simply divided by the fastest result. The number that matters for real machines sits between 2x and 2.5x depending on hardware and build, and climbing from 29 to 62 tokens a second at home remains a meaningful jump.

What the speaker wishes someone had said upfront: the GPU is almost the wrong thing to focus on. Bigger memory keeps more experts on the card; the 32 GB RTX 5090 kept 17463 of the 24576 expert blocks in VRAM and reached 179 tokens a second on the 2-bit build while the 12 GB 5070 stayed at 79 on the same build. But the CPU side, doing half the work, is bounded by memory bandwidth : the processor pulls about 42 GB a second from memory for the two-bit build, roughly everything the DDR5 side can give. Since the DDR4 side offers half that bandwidth, the CPU half of every layer round takes roughly twice as long on a DDR4 machine; the brand-new graphics card waits on the CPU, which waits on its memory. The test machine even ran its memory at 5200 megatransfers with the speed profile switched off in BIOS; enabling XMP or EXPO takes about 30 seconds and buys measurable speed here. The threshold that matters most is 64 GB of RAM: the 2-bit build wants 34 GB for experts plus about 10 for the operating system and everything else.

The memory bill and the honest decision

Quality is where compression actually costs, and the quantization table tells the price. The team behind the quantized files benchmarked them against the full model: full precision scores 87.4 on the live coding test while the 2-bit Q2 build writing 93 tokens a second sits at 81.1, about six points behind. The 3-bit IQ3 lands at 86.3, one point off the full model, so the 3-bit file at 62 tokens a second is the one for coding work and the 2-bit file for chat. For perspective, at full precision Flash-Next beats the 27-billion-parameter Qwen 3.8 by less than a point on one coding test, 62.5 versus 61.7, so the lead the whole setup exists to reach is already sub-point while the fastest build gives back six on another test. The climb in the price index is compared against game console prices in the current DDR5 coverage from Tomshardware; 64 GB of DDR5-6000 was listed at 913 dollars in the autumn period against an all-time low of 159 dollars. The hosted version costs 0.47 dollars per million output tokens, so 913 dollars buys about 1.9 billion tokens; writing that many at 93 tokens a second takes over 240 days of nonstop running, and the hosted copy answers many requests at once. That tariff is confirmed as the current list price in the data from PricePerToken.

One buried development changes the picture: on October 1st llama.cpp merged multi-token prediction tuned for this exact model, taking decode from 28 to 44 tokens, about 1.55x, on a machine holding the whole model in card memory. The catch is that on machines where the model spills into system memory, which covers most viewers, the same feature ran slower; each extra guess pulls its own experts from memory at a cost above the gain. The Strata report shows its own gain at 1.6 to 1.8x because the expert cache absorbs that cost. The expert cache itself remains an open draft in llama.cpp since August 28th, whose author measured 18.4 rising to 24.2 tokens on two cards. Since Strata is partly born from llama.cpp code, the picture is less a takeover than a two-feature public race for the lead. The verdict is clear: a desk already holding 64 GB of DDR5 and a 12 GB-plus Nvidia card should install and run it; anyone buying the memory from scratch should try the hosted tariff first.

Visualization: nodesdaily AI

Key moments

  1. One word in three minutes: slowness pictured
  2. Strata on stage: 93 tokens on 12 GB
  3. Reading only 4 GB per token
  4. Splitting layers: GPU and CPU together
  5. Draft layer: 3 tokens per pass
  6. Honest math behind the 6x claim
  7. Real limit: the 64 GB memory bar
  8. 913 dollars of RAM or the API?
  9. Verdict: who installs tonight, who waits

AI commentary

"What excites me here is not the speed figure but a model designed around the hardware people already own. The speaker shows the baseline instead of hiding it, and once you know where the 6x comes from, the story stops being hype and starts being a race worth following."

AI assessment

The strongest counter-argument targets the headline: like for like, the gain is 2 to 2.5x, and the fastest build is the weakest at code. The 93-token speed belongs to a quantization sitting six points below the full model, while the 3-bit file that closes the quality gap runs at 62. The speed record and the quality record do not live in the same file, so buyers must choose which one they are paying for. The title math deceives nobody on close reading, but it can still send a hasty reader to the checkout with 6x expectations.

Gaps remain in the account given on screen: figures come from a single machine over a few weeks, with DDR4 users, laptop cards and the AMD side covered only in passing. Long-horizon quality, error rates under agent workloads and coherence at 128K context stay open questions. Strata is also an eight-day-old 0.1 release, so interfaces, profiles and debugging tools will move fast. The speaker's position deserves a note too: an avowed local-inference enthusiast sharing the excitement, but one who publishes the baseline instead of hiding it.

The practical rule for readers is simple: with 64 GB of DDR5 and a 12 GB-plus card already on the desk, the install is free speed, with the 3-bit build for code, the 2-bit build for chat, and the memory profile switched on in BIOS. If the memory must be bought new, a few million tokens on the hosted API should come before a 900-dollar purchase. Offload-first architecture looks like the durable shift here, while the engine gap closes through public pull requests; the race itself, not the number, is what to watch.

Sources

8 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

local ai · strata · qwen · memory · quantization · llama.cpp · hardware

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…