When Apple’s new M5 Ultra Mac Studio landed next to my M3 Ultra, my first question was simple: do the claims of up to 4.3× local AI, up to 2× storage and 1.2 TB/s bandwidth hold outside the slide? The base system starts at $5,499 , while the 80-core GPU, 256 GB unified memory and 8 TB storage config on my desk reaches $14,299, with a 512 GB option due next month. Apple loaned the M5 unit, the M3 is my own purchase, so I could run both under identical Llama.cpp and MLX builds with identical model files .
CPU and Everyday Performance: First Measured Gaps
On Speedometer the M3 Ultra scored 47 versus 60.2 on M5 Ultra; a single M6 core is faster in isolation, but the Studio chassis still shows this generational lift. For developers, a large .NET build with 100k namespaces and classes fell from 71.7s to 52.4s, and a Python Mandelbrot style interpreted workload dropped from 10.25s to 5.8s, roughly 1.9× . The machine is overkill for pure coding, yet the multicore and interpreter gains are clean and repeatable.
Storage beats the marketing line. Sequential copy with AmorphousDiskMark gave roughly 6.5 GB/s read and 2.7 GB/s write on M3 Ultra versus 14.9 GB/s read and near 20 GB/s write on M5 Ultra; random small-file copies were also well beyond 2×. When you need to pull a 140 GB model from disk into memory , that delta directly shortens load time, and both drives were 8 TB so the comparison is apples-to-apples.
Memory Bandwidth and the AI Architecture Shift
On paper Apple lists 819 GB/s for M3 Ultra and 1.2 TB/s for M5 Ultra. With STREAM Triad , a CPU-centric test, I measured 312.9 GB/s vs 583.6 GB/s; the GPU micro-benchmark , which matters for AI, gave 724 GB/s vs 1,039 GB/s , about 1.4× real bandwidth on M5. That ratio tracks closely with the uplift seen in token generation, which is bandwidth-sensitive.
The core architectural change is that M3-era GPU cores were plain graphics cores, while each M5 GPU core now houses a dedicated neural accelerator for matrix multiply . That is why the same Llama.cpp binary reports `has_tensor = false` on M3 and `has_tensor = true` on M5, unlocking a Metal path tuned for the M5 family . MLX , Apple’s own stack, is often fastest overall, but this hardware hook explains M5’s sharp prompt-processing jump in Llama.cpp. All runs used identical builds and model weights on both machines.
The distinction between prompt processing (PP) and token generation (TG) matters here. PP — reading your prompt — leans on compute and GPU core speed , while TG — writing the answer — leans on memory bandwidth . M5 Ultra lifts both, but PP benefits disproportionately from the per-core accelerator, which is why long prompts pull away.
LLM Numbers: How Much Faster Per Model
On the large mixture-of-experts DeepSeek V4 Flash 284B , TG rose from 37 to 53 tok/s, about 1.5× , while PP leapt from 483 to 1,485 tok/s, roughly 3× . On GPT-OSS 120B , TG moved from 90 to 130 tok/s and PP from about 1,300 to 3,200 tok/s; feeding 32k new tokens on top of 64k already in context took 80s vs 56s, again near 1.4× . The pattern is consistent: TG follows bandwidth, PP follows the new accelerators.
On a smaller but dense Qwen3 27B , where every token touches all weights, the load is heavier yet still revealing. TG went from 37 to 54 tok/s, 1.47× , while PP jumped from 430 to 1,800 tok/s, over 4× . A realistic 14k-token code prompt equivalent to a handful of files saw PP fall from 33s to 8s, 4.2× , confirming Apple’s up-to-4× claim right where it matters — long, file-fed contexts. The curve is prompt-length dependent: about 1.6× at ~350 tokens , 3.4× near 1,700, and the 4× threshold around 4,500 tokens , so chat feels modest while agent workloads shine.
Power and thermals tell a more tempered story. During generation the chip telemetry showed 36W on M3 vs 45W on M5 , so tokens per watt rose only ~12%; at idle it was ~11–12W vs 8–10W, slightly lower on M5. Under sustained full GPU load, M3 peaked near 200W while M5 pushed beyond 400W , running visibly hotter around 43–44°C and audibly louder on mic. Efficiency improved, absolute power and heat did not.
One more note from the lab bench: Llama.cpp’s Metal path now branches for the M5 family while MLX’s per-core tuning keeps evolving, so the same silicon has two different peaks and leadership flips with model, context length and quantization.
Capacity today fit comfortably within 256 GB , but the coming 512 GB option will unlock still larger models and faster multi-model setups. With eight concurrent requests , aggregate throughput rose from 75 tok/s to 134 tok/s, showing the gain persists under concurrency. A deeper dive into MLX versus Llama.cpp and 8-bit variants is next; the first-look takeaway is clear — ~1.5× TG across models and 2–4× PP depending on prompt length , with coding agents as the prime beneficiary.
Token Generation Compared (tok/s)
- DeepSeek M337
- DeepSeek M553
- Qwen M337
- Qwen M554
- GPT-OSS M390
- GPT-OSS M5130
| Finding | Value |
|---|---|
| Bandwidth uplift | 1.4× on GPU (724→1,039 GB/s) |
| Token generation | ~1.5× across MoE and dense |
| Prompt processing | 1.6× short, 4.2× long; threshold ~4.5k tok |
| Metric | M3 Ultra | M5 Ultra |
|---|---|---|
| GPU bandwidth | 724 GB/s | 1,039 GB/s |
| DeepSeek TG | 37 tok/s | 53 tok/s |
| DeepSeek PP | 483 tok/s | 1,485 tok/s |
| Storage read | 6.5 GB/s | 14.9 GB/s |
| Python Mandelbrot | 10.25 s | 5.8 s |
Key moments
- Intro — M5 Ultra on desk next to M3
- Price and config: 256 GB, 80 GPU, 8 TB
- CPU tests: Speedometer and build times
- Storage: sequential and random beyond 2×
- Bandwidth: STREAM and GPU micro bench
- Architecture: per-core neural accelerator and Metal path
- LLMs: PP and TG on DeepSeek and GPT-OSS
- Power and heat: 400W peak, 43°C and fan noise
AI commentary
"My read: M5 Ultra is a serious raw jump, especially for coding agents on long context. For short chat the gain is subtler; to justify a $14k config you need to feed it large contexts and heavy models."
AI assessment
Strength: the video sets a transparent protocol — same builds, same models, same drive size and same power probes. That makes the 1.4× bandwidth → 1.5× token generation correlation read as measurement, not marketing, and showing MoE vs dense on the same day is genuinely useful.
Limits: 1.6× PP on short prompts sits where a daily chat user may barely feel it, and the 400W+ peak with 43–44°C thermal cost is non-trivial; efficiency is up only ~12% while acoustics and heat rise, which can clash with the quiet-studio promise.
Takeaway: M5 Ultra is not just faster, it is Apple rewriting the Metal path for local inference with a per-core neural accelerator . That leaves a mark even in cross-platform stacks like Llama.cpp and pushes Apple, via the coming 512 GB option, deeper into large-model territory for long-context agents.
Practically: if your work is running coding agents on large contexts, frequent 140 GB loads and concurrent serving , you will pocket the gain; for short-prompt chat and light tasks M3 Ultra remains ample and cooler/quieter — let prompt length and model size decide the upgrade.
Sources
8 links; 2 of them also cited by 4 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — M5 Ultra vs M3 Ultra Comparison
- @apple.com https://www.apple.com/newsroom/2026/08/apple-introduces-new-mac-studio-with-m5-max-and-m5-ultra
Also cited by: 2026 Mac Mini and Mac Studio: From M5 Pro to M5 Ultra — Silent Speed and a 768 GB Memory Pool · Buying a Mac for AI: The Only Guide You Need
- @apple.com https://www.apple.com/newsroom/2026/08/apple-introduces-m6-and-m5-ultra-for-a-big-leap-in-performance-and-ai-compute
Also cited by: Apple's Impossible Double — 2-Nanometer in the Cheapest Mac, a Quad-Die Monster in the Priciest · MacBook Pro M6: 2nm Chip, a One-Chip Generation, and the Waiting Question
- @macworld.com https://www.macworld.com/article/3220024/apple-announces-the-m5-ultra-mac-studio-with-up-to-512gb-of-ram.html
- @applemust.com https://www.applemust.com/apple-m5studio/
- @arxiv.org https://arxiv.org/html/2607.19438v1
- @mactech.com https://www.mactech.com/2026/08/25/apple-introduces-new-mac-studio-with-m5-max-and-m5-ultra
- @tomsguide.com https://www.tomsguide.com/computing/cpus/the-usd3-999-m3-ultra-mac-studio-barely-beats-the-usd1-999-m4-max-in-leaked-benchmark
m5 ultra · m3 ultra · mac studio · apple silicon · llm · memory bandwidth · storage