Back to feed

Giant Models on a $3,000 Mini PC: What 128GB Unified Memory Actually Runs

The Stack puts the 128GB Bosgame M5 (AMD Ryzen AI Max+ 395, Strix Halo) under the microscope: a pool five times larger than a 24GB graphics card lets huge models fit, but 256 GB/s of bandwidth dictates per-token speed. Mixture-of-Experts architectures stream far faster than dense models, long prompts hit a compute ceiling, and a price jump from $1,699 to $2,999 rewrites the value equation.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — CJz2edJ62HQ
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

Bosgame’s M5 mini desktop flips the familiar graphics-card logic: instead of a separate 24GB VRAM pool, the CPU and integrated GPU share a single 128GB pool . AMD’s Ryzen AI Max+ 395 (Strix Halo) calls this unified memory — the processor and graphics dip into the same physical memory — so one model can exceed the 24GB ceiling by more than five times . The Stack’s question is precise: does putting that much memory on a desk for $3,000 make it the best option for local AI, or is fitting the biggest model not the same as getting the fastest answer?

What Unified Really Means — and Why It Feels Slower

In a normal desktop the CPU lives in its own RAM while the graphics card holds a private stash of dedicated VRAM across the PCIe bus. Try to load a 70-billion-parameter dense model into a 24GB card and it spills over into sluggish system memory, crawling to a halt. The M5 removes that wall. Think of a model as a huge book that must lie perfectly flat before you can read a word — the shared desk is enormous, so giant models sit in one piece. The whisper on the spec sheet is bandwidth: TechPowerUp measures the shared pool at ~256 GB/s , several times slower than a discrete card’s ~1,008 GB/s on an RTX 4090. Producing each word means sweeping through the model’s numbers out on that desk, once per token, so pool speed, not just size, decides how quickly text appears. Under Linux the GPU does not automatically receive the full 128GB allocation; guides that cite ~110GB usable for graphics require you to raise the reservation with kernel-level tuning, otherwise large models bounce at the door.

Which models fit and which stay fast comes down to architecture. A dense model — every parameter wakes for every word — fits even at 70B but must drag all 70B numbers across the shared bus for each token, crawling at 4-5 tokens/s . A Mixture-of-Experts (MoE — a sparse design where only a few expert sub-networks fire per token) does the opposite: a 120-billion-parameter giant keeps a vast library resident yet routes each word to only 5-6B active parameters. Like opening only the right shelf in a huge library for each question. Channel measurements on this silicon show the effect: Qwen3 and GPT-OSS 120B swing from 15 to 56 tokens/s depending on the software backend, while a smaller 30B MoE pushes above 70 tokens/s . When the smaller MoE is the fastest of the three, architecture matters as much as capacity.

How Much Memory Do You Need? The 0.6 GB Rule and Conflicting Numbers

To estimate fit, Model Fit publishes ~0.6 GB per billion parameters at 4-bit quantization . Under that rule 80B needs ~50GB , 120B needs ~65GB , both comfortably below the 110GB ceiling. Yet no single authoritative speed exists. Mine Studio reports GPT-OSS 120B at ~30 tokens/s and a 35B model above 50 , while Model Fit lists an untested 11 tokens/s estimate for the larger model. The spread is compiler flags and runtime (llama.cpp, ROCm backend). The channel’s own rerun proves it: the same Qwen model crawled at 1.79 tokens/s on an unoptimized build and leapt to 15.56 tokens/s with the right flags on identical silicon. The rule remains: fewer active parameters per token means less data crossing the pool and faster streaming.

AMD’s ‘2.2× faster than an RTX 4090’ headline makes sense only in one lane. For Llama 3.1 70B that does not fit in 24GB, the 4090 must sip spilled weights through the PCIe bus like a marathon runner through a cocktail straw; the M5 keeps the whole model nestled in its slower unified pool and the arithmetic checks out. That CES 2025 comparison was built on that spill. The r/LocalLLaMA crowd spotted the trap immediately: Strix Halo did not dethrone the desktop king, AMD engineered the most favorable bench. When the model *does* fit in 24GB, the discrete card is substantially faster . You trade raw generation speed for the ability to house enormous models inside a compact enclosure.

Reading the Prompt Is Not Writing the Answer

Generating text is two jobs: prefill (ingesting the whole prompt at once) and decode (emitting tokens one by one) . Decode is memory-bandwidth bound and the M5 handles it reasonably; large MoEs stream at conversational speed. Prefill is pure parallel compute. Before the first token can appear, the model must digest every sentence, code snippet and clarification you pasted — all at once. That is where the integrated Radeon 8060S , ambitious as it is, runs out of steam against a desktop GPU. Reddit testers of Strix Halo call this phase compute-limited , a noticeable bottleneck on long prompts. Casual chat with brief prompts barely shows it. A local coding assistant does: on every turn it packages documentation, whole source files, project trees and compiler errors, dumping far more text than you typed. Asking it to rename a variable across a codebase means a heavy ingestion pause before it even starts replying — billed on every turn.

Setup is its own project. Windows treats the box like a polite desktop and reserves only a cautious slice for graphics; you must raise the allocation yourself before anything large loads, and the exact menu moves with machine and driver version. Then software matters more than silicon: llama.cpp on the same Qwen model jumps from 1.79 to 15.56 tokens/s with flags and backend alone. The ROCm stack is a fast-moving target: one developer spent 9 months across four ROCm releases on a 96GB box trying to get the VLM engine running and ended that stretch with no working setup, later reporting it works on newer releases without the workarounds that had defeated them — while others on identical hardware still need to build from source. Before buying for a specific tool, check that tool’s *current* ROCm/llama.cpp status rather than trusting any dated review, including this one.

The Price Reality: $1,699 to $2,999, a $3,847 Peak and Soldered Memory

Price alone is a lesson. Launch pricing was $1,699 for the 128GB build ; the Bosgame storefront today lists the same build at $2,999 — a $1,300 jump. Guru3D reported the intro price when only the 128GB variant was offered. The 96GB version is several hundred less, but availability is the real story: every older low-price listing is sold out, the only figure that matters is the config you can actually buy. Model Fit tracked a comparable box near $3,847 in July 2026 at the peak of the memory shortage; the same class had launched nearer $2,000 . A machine that is mostly memory moves with the price of memory. And that memory is soldered LPDDR5X — wired directly to the board because the fast interface demands it. You cannot pop the case and add sticks later. Paying $3,000 for a non-expandable box fundamentally changes the math, which is why the right move is to buy the exact memory size your largest planned model genuinely needs, not the cheapest tier you can talk yourself into.

Clear away the marketing noise and the conclusion sharpens. Bosgame’s ‘gaming equals RTX 4070’ claim and the big TOPS figure on the spec sheet belong in the vendor corner; the former is a vendor slide, the latter counts the NPU (neural processor) while every measurement here used the graphics side. More importantly, any vendor assembling a compact desktop around this AMD platform relies on identical processor and graphics IP ; seeing Ryzen AI Max+ 395 on the listing means the engine is identical — the M5 is one way to buy this chip, not the obvious choice. At its current price, shopping the field for build quality, support and warranty is well worth it. The single most important lesson: capacity (what you can load) and speed (how fast it streams) are two separate questions . The M5 offers a rare luxury — massive frontier MoEs held quietly and locally — combined with the sobering sight of it working through extensive prompts at a deliberately unhurried tempo. A generous spec on one tells you nothing about the other.

Visualization: nodesdaily AI

Bandwidth Comparison

  • M5 unified (256)256 GB/s
  • RTX 4090 (1008)1,008 GB/s
Unified pool bandwidth deficit — TechPowerUp and GDDR6X.
MetricM5 Strix HaloRTX 4090
Memory pool128GB unified24GB VRAM
Bandwidth~256 GB/s~1,008 GB/s
70B dense4-5 tok/sspills -> slow
120B MoE15-56 tok/sdoes not fit
List price$2,999 (128GB)~$1,599

Key moments

  1. The question: what does a $3,000 mini PC actually run?
  2. Removing the wall: unified memory, 128GB pool
  3. Desk analogy and the 256 GB/s reality
  4. MoE vs dense: 5-6B active vs 70B at 4-5 tok/s
  5. 120B MoE 15-56 tok/s, 30B MoE 70+ tok/s
  6. The 2.2× vs 4090 claim and PCIe straw
  7. Prefill vs decode: compute limit on long prompts
  8. Coding assistant ingests the repo every turn
  9. 1.79 to 15.56: flags decide speed
  10. $1,699 → $2,999 → $3,847: price and soldered LPDDR5X

AI commentary

"My take: this is not a speed demon, it is a capacity demon. If you want a huge MoE sitting quietly on your desk at conversational speed, it is compelling; if you expect an assistant that must ingest thousands of lines on every turn to still feel snappy, you are confusing capacity with speed and disappointment is baked in."

AI assessment

Steel-manned, the case is coherent: if you want a large MoE held quietly and locally, a 128GB unified pool that fits 120B in ~65GB and streams chat is a rare practical answer. The fulcrum is architecture — 5-6B active parameters per token eases bandwidth and unlocks the 15-56 tokens/s band seen in channel testing. Framed that way, the M5 is a deliberate trade of pure generation speed for capacity in exchange for privacy and always-on access.

Limits are stark: the pool is huge but slow (256 GB/s), dense models collapse to 4-5 tokens/s, and the integrated GPU is compute-limited on long prompts, producing a noticeable ingestion pause before every turn when you feed repository-scale context. Setup is sensitive — Windows allocation plus ROCm/llama.cpp flags explain the 1.79 to 15.56 tokens/s jump on identical silicon — and the 9-month VLM saga shows software is more volatile than hardware. There is also a methodology gap: no single authoritative speed, every source measures with different backends and flags, so generalizations are fragile.

For interests and verifiability, read the headlines carefully: ‘2.2× over 4090’ holds only when the workload spills past 24GB; when it fits, the discrete card wins. Gaming/TOPS claims count the NPU, outside the scope of the graphics-side measurements here. Pricing matters too: $1,699 to $2,999 today, $3,847 at the July 2026 memory peak, and soldered LPDDR5X ties cost to the memory cycle with no upgrade path. Independent repeat measurements exist in pieces (Mine Studio, Model Fit, in-channel reruns) but not as one blind controlled comparison, so every speed claim should be read as ‘which backend, which flags, which allocation.’

Practically: if your job is local chat with a large, sparse MoE and you value a small quiet box, the trio of sparse large model + correct backend + ~110GB allocation can work. For long-context, frequent re-ingestion work like coding assistants, the integrated compute ceiling bites — budget for the ingest pause on every ‘fix the whole repo’ command. When buying, size memory to the largest model you genuinely plan to run, not the cheapest tier you can talk yourself into; since every competitor around this AMD silicon uses the same Ryzen AI Max+ 395 engine, compare on build quality, support and warranty and keep capacity vs speed as two separate questions.

Sources

8 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

bosgame m5 · strix halo · unified memory · moe · local ai

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…