The creator says he has used an M5 Max MacBook Pro with 128 GB of unified memory as his daily driver for five months. In his framing it is currently the highest configuration available in the MacBook Pro form, carrying his whole workload on one machine. The video asks a sharp question: how far can this hardware be pushed for local AI, and where does it hit the wall.

The local AI stack is not a single app: LM Studio, MLX-based runners, Ollama and llama.cpp sit side by side. On the model side he has tried various sizes from the Qwen, Gemma and GLM families. Thanks to unified memory, large weight files are shared between CPU and GPU without copying, which keeps the whole setup practical.

Coding agents change the picture: local assistants respond so fast that the experience sometimes feels like being connected to a cloud model. Large-context editing and project-wide navigation run smoothly on the machine. For him, that everyday fluidity is more convincing than raw benchmark numbers.

Heat and noise split by model size: with small-parameter models the machine stays cool and quiet, fans barely spinning. Large models flip the picture, especially the newer Flash-class releases, where fans run constantly and heat output is obvious. Silence is not guaranteed; it depends on the workload.

The fine-tuning segment is one of the video's surprises: he uses Unsloth's MLX framework to fine-tune large language models on his own data, and calls the result shockingly smooth on this laptop. He has also prepared a course on fine-tuning with local models for those who want to go deeper. Training on a laptop sounded like fantasy a few years ago; here it is presented as routine.

Models he names include a 27-billion-parameter dense model from the Qwen family, a Flash Next mixture-of-experts model, and GLM's 5.3 Flash release. He is frank that latency varies with context length, query type and core utilization. Still, his overall verdict is positive: waits vary, but jobs get done and the machine never strands him.

Side by side with his older M1 Max, he reports gaps of up to 8x, crediting the AI acceleration blocks inside the GPU. The gain shows not only in chat replies but in image generation and transformer training. The video itself was edited and rendered on this machine; large projects with many assets and fast exports in Final Cut Pro and Camtasia are highlighted.

Memory bandwidth is quoted at 614 GB/s and presented as the reason this machine stands out. Against his M4 mini the gap is described as several multiples, and the newly announced mini does not come close either. The bandwidth jump from M1 Max to M5 Max shows how wide the gap has grown.

The pricing story is the most concrete part: he bought the machine in April for about NZ$11,599. The later memory-price surge lifted the same configuration by roughly NZ$5,000, a 30-40 percent increase. Buying early gave him a serious edge; the same decision at today's sticker would be much harder.

In the live demo, a 26-billion-parameter mixture-of-experts model from the Gemma 4 family runs with 4-bit quantization and an accelerated speculation module switched on. The model loads fully into memory on first run, Activity Monitor shows around 46 GB in use, and memory pressure stays green. Generation speed measures around 108 tokens per second in Q&A, touching 150 on light prompts.

He also points to earlier experiments: results with Qwen's 27-billion-parameter model, fast SVG image generation, and a full .NET 8 to .NET 10 migration including library updates finishing in minutes. The speculation module visibly shortens such multi-step jobs. A local coding helper running at this speed is the video's strongest demo.

On the plus side: portability, quiet daily use, smooth inference in the 27-32 billion parameter class, and on-laptop fine-tuning; the machine looks future-proof for that class. On the minus side sits the ceiling: current large models can already approach 121-122 GB of virtual memory, and with the OS plus quantization overhead the 128 GB limit is within reach. If a future model needs 192 or 256 GB while promising Claude-class local performance, this machine will not be enough.

He closes by noting Macs once sat on desks for ten years, while memory hunger could date this machine within a couple of years. Apple's plans to run its own AI models on these machines are part of the equation. Even so, he is happy with the purchase, and suggests anyone waiting for an M6 or M6 Ultra decide against this backdrop before thanking viewers and signing off.

To steelman the other side: cloud API advocates say time beats hardware, and paying a few dollars of API fees instead of wrestling a stalled local model is rational for most teams. Independent guides agree, noting that a $4,000 laptop still loses some jobs to a $2,000 GPU desktop. I think that objection narrows rather than refutes the video: for learning, tinkering, and keeping data on-device, the local setup is a genuine option.

There are methodology gaps: this is a single-machine, single-user, uncontrolled experience report. The on-screen 108-150 tokens per second come from a small mixture-of-experts model, while independent leaderboards put dense 70-billion-parameter models on this hardware class at 15-18 tokens per second. The fine-tuning claim has no timings or loss curves behind it, and the heat observations stay at the level of feel rather than thermometer.

The interest and verifiability lens matters too: the presenter bought the machine with his own money and promotes a paid course on the topic, which adds natural optimism to any satisfaction claim. Fortunately the core figures intersect with independent sources: the 614 GB/s bandwidth and the memory-price surge are confirmed by Apple's own statements. Even so, I would re-check token speeds and prices against current listings at decision time, since both shift monthly.

My practical takeaway: for anyone wanting one portable machine that runs the 27-32 billion parameter class smoothly every day, this configuration is a strong pick. Those planning to run dense 70B-plus models locally in production, those priced out by the memory surge, or those staying at a desk will find the 128 GB desktop options or waiting for M6 more sensible. For anyone who wants data to stay on-device, this video offers a real roadmap.

AI commentary

"What I find honest about this video is how it holds praise and warning in the same breath: 128 GB is the sweet spot today, but at this rate of model appetite it could become the ceiling."

AI assessment

To steelman the other side: cloud API advocates say time beats hardware, and paying a few dollars of API fees instead of wrestling a stalled local model is rational for most teams. Independent guides agree, noting that a $4,000 laptop still loses some jobs to a $2,000 GPU desktop. I think that objection narrows rather than refutes the video: for learning, tinkering, and keeping data on-device, the local setup is a genuine option.

There are methodology gaps: this is a single-machine, single-user, uncontrolled experience report. The on-screen 108-150 tokens per second come from a small mixture-of-experts model, while independent leaderboards put dense 70-billion-parameter models on this hardware class at 15-18 tokens per second. The fine-tuning claim has no timings or loss curves behind it, and the heat observations stay at the level of feel rather than thermometer.

The interest and verifiability lens matters too: the presenter bought the machine with his own money and promotes a paid course on the topic, which adds natural optimism to any satisfaction claim. Fortunately the core figures intersect with independent sources: the 614 GB/s bandwidth and the memory-price surge are confirmed by Apple's own statements. Even so, I would re-check token speeds and prices against current listings at decision time, since both shift monthly.

My practical takeaway: for anyone wanting one portable machine that runs the 27-32 billion parameter class smoothly every day, this configuration is a strong pick. Those planning to run dense 70B-plus models locally in production, those priced out by the memory surge, or those staying at a desk will find the 128 GB desktop options or waiting for M6 more sensible. For anyone who wants data to stay on-device, this video offers a real roadmap.

Sources

m5 max · local ai · 128 gb