The opening question is direct: with new Mac minis and Mac Studios announced, should you buy a fresh machine or keep the one you have? The answer depends on four variables: what you want the model to do, how fast you want it, how much memory you have, and what your budget is. All analysis runs through a free comparison tool the creator built: pick a Mac and see which models fit, or pick a model and see which Mac it needs. If the video passes one hundred thousand views, the tool goes open source.
He starts with capability buckets instead of raw benchmark tables, asking what each model can actually do for you. Bucket one is 38 to 48 GB: Qwen 3.8 27B, the everyday model. Summarizing a fifty-page PDF, drafting or explaining a contract, writing scripts, working through a codebase one task at a time all sit here. He compares it with Opus 4.6, a frontier model from six months ago. The limit is clear: leave it alone for an hour on a big coding job or ask expert-level research questions and it falls behind the cloud.
Bucket two is around 128 GB: Qwen 3.8 Flash and peers. This unlocks multi-step office work with tools: research across sites, filling a spreadsheet from a pile of documents, stronger multi-file coding with tests running on its own. It is a tight fit on a 128 GB machine and struggles if other apps are open. He places it just below Claude Opus 4.8, near GPT 5.6 Terra. It is visibly faster and more capable than the previous bucket.
Bucket three needs at least 256 GB: GLM 5.3 Flash. This is the model that circulated anonymously and free on OpenRouter as Ox Alpha before Z.ai confirmed it as a GLM release. He calls it engineer-grade: project-level coding, test and fix loops, longer sessions, native multimodal input. Total size is large but the mixture-of-experts design reads little data per token, which brings speed and low per-token cost. DeepSeek V4 Flash is an alternative in this band but he rates it less capable.
Bucket four is 512 GB, the current single-machine ceiling: full GLM 5.2 and 5.3. With only about 54 billion parameters active, the MoE structure reaches flagship-adjacent capability: expert work, better coding, longer autonomous sessions. The reference point is GPT 5.6 Soul, with a caveat: longer sessions bring more hallucination and consistency loss than cloud flagships. Still a striking level for a locally run model.
Bucket five is beyond one Mac: chaining two, four or more machines to run giant models distributively. The example is Kimi K3, needing roughly 1.5 TB of memory at Q4, in other words about four top-tier Studio-class machines. The payoff lands between Opus 4.8 and Fable 5, the closest local point to the current frontier. He frames it as a ceiling example for those with budget and infrastructure, not for everyone. The message is explicit: everybody wants the biggest, but not everybody needs it.
These buckets are cross-checked against the Artificial Analysis Intelligence Index: Kimi K3 near 50 against Fable 5 at 53, the GLM family in the GPT 5.6 Soul band, Qwen 3.8 27B below Opus 4.8, GLM Flash above Opus 4.8. The more striking point is the time trend: the same size gets smarter quickly, and the Qwen 3.6 to 3.8 jump proves it. So a job needing 256 GB today may be solved in the 48 GB band in six to twelve months. Buying big fixes today but does not insure the future.
The physics of speed is the heart of the video. Every token reads the whole model from memory: Qwen 3.8 27B at Q4 means about 17 GB per token and up to forty tokens per second depending on machine; the Q8 file is 29 GB so per-token reads rise to 31 GB and speed falls. The GLM Flash file is near 200 GB at Q4 yet reads only about 30 GB per token thanks to MoE, so it stays fast despite its size. Quantization (Q4/Q8) shrinks files, eases fit, and trades a small quality loss for speed.
The ceiling is memory bandwidth: Mac mini M6 near 170 GB/s, MacBook Pro M5 Pro 307 GB/s, M5 Max 460 to 614 GB/s, M5 Ultra 1.2 TB/s. The worked example is Qwen 3.8 27B Q4 at 32k context: about 6.5 tokens per second on the M6 which feels very slow, 13 on the M5 Pro, 22 on the M5 Max, around 40 on the M5 Ultra. Context window is the hidden variable: cache is a fraction of a gigabyte at 8k but 17 GB at 256k, cutting speed roughly in half. Yet real agent work needs 128k and above; comparisons at tiny windows look fast but mislead.
Chip overhead, runtime and acceleration complete the picture. There is a fixed per-token tax around 4 ms on the M4 generation and 0.5 ms on M5; it barely matters for big models but dominates small fast ones. On identical hardware MLX runs up to forty percent faster than llama.cpp-based GGUF on MoE models, with a small gap on dense ones. Multi-token prediction lifts the Qwen example from 22 tokens per second into the 34 to 56 band, and draft-model speculative methods go further. In contrast, spare RAM beyond fit, CPU cores, or GPU core counts alone do not raise decode speed; bandwidth sets the ceiling and extra gigabytes sit idle.
Fit and speed maps from the tool follow. Qwen 3.8 27B Q4 at 32k context fits every listed Mac; Flash runs on a 64 GB machine via SSD paging with a large quality penalty; the comfortable threshold for GLM Flash is a 256 GB M5 Ultra with the slower M3 Ultra 256 GB as alternative. At 256k context the picture hardens: most machines below 128 GB must compromise, while 128k context fits a 64 GB machine comfortably. On speed, the 48 GB Mac mini M5 Pro starts at 16 tokens per second base and reaches 58 accelerated; an M1 Max 32 GB still runs it but heats up; the 512 GB M5 Ultra projects a 50 base and up to 180 accelerated on Qwen. On a work machine you must subtract open apps and the macOS default 67 to 75 percent memory ceiling; MLX reserves more than GGUF on long prompts.
The cloud comparison sharpens the economics: a 256 GB M5 Ultra near ten thousand dollars equals four years of a two-hundred-dollar monthly plan. But most work fits twenty to one-hundred-dollar plans, cloud limits and restrictions keep growing, some flagship releases exhaust quickly, and future models may reach consumers late. Moreover the best model is rarely required. Local costs are setup burden: no ready-made memory and connectors like cloud apps, you wire the harness yourself including an agent setup such as Hermes; the math assumes one active chat, four parallel chats mean four times the memory; privacy is a plus but security duty is yours with prompt injection and weaker alignment risk; hallucination scores trail cloud flagships with the Qwen band negative and GLM Flash in single-digit positive. Hence one machine per bucket: Mac mini M5 Pro 48 GB for basic Qwen, 128 GB for Flash, Studio M5 Ultra 256 GB for GLM Flash, October 512 GB for full GLM at an estimated fifteen to eighteen thousand dollars. Comfortable speed is 12 to 30 tokens per second, below 15 is low, above 30 is fast.
The tool itself is free with no signup: an Apple Silicon list from M1 upward, an Nvidia DGX Spark option, custom machine entry, model, quantization and context editing, workhorse assumption and GGUF versus MLX choice, budget filter. The fit matrix shows which model fits which machine at which compromise, the speed table shows per-machine per-model rates, the model-first view maps one model across all Macs, the memory calculator totals selected models. Every cell is clickable: max context, max quantization, memory share, expected speed and Hugging Face sources open up; the speed formula lives on a separate page with wrong-number reporting plus PNG export and shareable settings links. The creator promises open source at one hundred thousand views so the community can extend machine and model data and correct formulas.
AI commentary
"What struck me is how this video flips the question: it is not how much RAM you buy, but which model you run at which context window. Pick the job first, then the machine; the reverse gets expensive fast."
Sources
7 links; 1 of them also cited by 2 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube https://www.youtube.com/watch?v=C9Q1ArLSisw
- @apple https://www.apple.com/newsroom/2026/08/apple-introduces-new-mac-studio-with-m5-max-and-m5-ultra
Also cited by: M5 Ultra Tested: Memory, Storage and Local LLM Speed vs M3 Ultra · 2026 Mac Mini and Mac Studio: From M5 Pro to M5 Ultra — Silent Speed and a 768 GB Memory Pool
- @modelfit.io https://modelfit.io/mac-studio/m5
- @explainx.ai https://explainx.ai/blog/qwen-3-8-27b-open-weight-model-claude-opus-comparison-august-2026
- @openrouter https://openrouter.ai/z-ai/glm-5.3-flash
- @cellcog.ai https://cellcog.ai/blog/glm-5-3-flash
- @muhammadraza.me https://muhammadraza.me/2026/gguf-vs-mlx-decision-guide