Back to feed

M5 Ultra Tested: Memory, Storage and Local LLM Speed vs M3 Ultra

With M3 Ultra and new M5 Ultra side by side, I tested Apple’s claims of up to 4.3× AI, 2× storage and 1.2 TB/s bandwidth on identical models and stacks; results largely confirm marketing, but the felt gain depends on prompt length.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — c_58D7ixOQI
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

When Apple’s new M5 Ultra Mac Studio landed next to my M3 Ultra, my first question was simple: do the claims of up to 4.3× local AI, up to 2× storage and 1.2 TB/s bandwidth hold outside the slide? The base system starts at $5,499 , while the 80-core GPU, 256 GB unified memory and 8 TB storage config on my desk reaches $14,299, with a 512 GB option due next month. Apple loaned the M5 unit, the M3 is my own purchase, so I could run both under identical Llama.cpp and MLX builds with identical model files .

CPU and Everyday Performance: First Measured Gaps

On Speedometer the M3 Ultra scored 47 versus 60.2 on M5 Ultra; a single M6 core is faster in isolation, but the Studio chassis still shows this generational lift. For developers, a large .NET build with 100k namespaces and classes fell from 71.7s to 52.4s, and a Python Mandelbrot style interpreted workload dropped from 10.25s to 5.8s, roughly 1.9× . The machine is overkill for pure coding, yet the multicore and interpreter gains are clean and repeatable.

Storage beats the marketing line. Sequential copy with AmorphousDiskMark gave roughly 6.5 GB/s read and 2.7 GB/s write on M3 Ultra versus 14.9 GB/s read and near 20 GB/s write on M5 Ultra; random small-file copies were also well beyond 2×. When you need to pull a 140 GB model from disk into memory , that delta directly shortens load time, and both drives were 8 TB so the comparison is apples-to-apples.

Memory Bandwidth and the AI Architecture Shift

On paper Apple lists 819 GB/s for M3 Ultra and 1.2 TB/s for M5 Ultra. With STREAM Triad , a CPU-centric test, I measured 312.9 GB/s vs 583.6 GB/s; the GPU micro-benchmark , which matters for AI, gave 724 GB/s vs 1,039 GB/s , about 1.4× real bandwidth on M5. That ratio tracks closely with the uplift seen in token generation, which is bandwidth-sensitive.

The core architectural change is that M3-era GPU cores were plain graphics cores, while each M5 GPU core now houses a dedicated neural accelerator for matrix multiply . That is why the same Llama.cpp binary reports `has_tensor = false` on M3 and `has_tensor = true` on M5, unlocking a Metal path tuned for the M5 family . MLX , Apple’s own stack, is often fastest overall, but this hardware hook explains M5’s sharp prompt-processing jump in Llama.cpp. All runs used identical builds and model weights on both machines.

The distinction between prompt processing (PP) and token generation (TG) matters here. PP — reading your prompt — leans on compute and GPU core speed , while TG — writing the answer — leans on memory bandwidth . M5 Ultra lifts both, but PP benefits disproportionately from the per-core accelerator, which is why long prompts pull away.

LLM Numbers: How Much Faster Per Model

On the large mixture-of-experts DeepSeek V4 Flash 284B , TG rose from 37 to 53 tok/s, about 1.5× , while PP leapt from 483 to 1,485 tok/s, roughly 3× . On GPT-OSS 120B , TG moved from 90 to 130 tok/s and PP from about 1,300 to 3,200 tok/s; feeding 32k new tokens on top of 64k already in context took 80s vs 56s, again near 1.4× . The pattern is consistent: TG follows bandwidth, PP follows the new accelerators.

On a smaller but dense Qwen3 27B , where every token touches all weights, the load is heavier yet still revealing. TG went from 37 to 54 tok/s, 1.47× , while PP jumped from 430 to 1,800 tok/s, over 4× . A realistic 14k-token code prompt equivalent to a handful of files saw PP fall from 33s to 8s, 4.2× , confirming Apple’s up-to-4× claim right where it matters — long, file-fed contexts. The curve is prompt-length dependent: about 1.6× at ~350 tokens , 3.4× near 1,700, and the 4× threshold around 4,500 tokens , so chat feels modest while agent workloads shine.

Power and thermals tell a more tempered story. During generation the chip telemetry showed 36W on M3 vs 45W on M5 , so tokens per watt rose only ~12%; at idle it was ~11–12W vs 8–10W, slightly lower on M5. Under sustained full GPU load, M3 peaked near 200W while M5 pushed beyond 400W , running visibly hotter around 43–44°C and audibly louder on mic. Efficiency improved, absolute power and heat did not.

One more note from the lab bench: Llama.cpp’s Metal path now branches for the M5 family while MLX’s per-core tuning keeps evolving, so the same silicon has two different peaks and leadership flips with model, context length and quantization.

Capacity today fit comfortably within 256 GB , but the coming 512 GB option will unlock still larger models and faster multi-model setups. With eight concurrent requests , aggregate throughput rose from 75 tok/s to 134 tok/s, showing the gain persists under concurrency. A deeper dive into MLX versus Llama.cpp and 8-bit variants is next; the first-look takeaway is clear — ~1.5× TG across models and 2–4× PP depending on prompt length , with coding agents as the prime beneficiary.

Visualization: nodesdaily AI

Token Generation Compared (tok/s)

  • DeepSeek M337
  • DeepSeek M553
  • Qwen M337
  • Qwen M554
  • GPT-OSS M390
  • GPT-OSS M5130
Token generation on identical models; higher is better.
FindingValue
Bandwidth uplift1.4× on GPU (724→1,039 GB/s)
Token generation~1.5× across MoE and dense
Prompt processing1.6× short, 4.2× long; threshold ~4.5k tok
MetricM3 UltraM5 Ultra
GPU bandwidth724 GB/s1,039 GB/s
DeepSeek TG37 tok/s53 tok/s
DeepSeek PP483 tok/s1,485 tok/s
Storage read6.5 GB/s14.9 GB/s
Python Mandelbrot10.25 s5.8 s

Key moments

  1. Intro — M5 Ultra on desk next to M3
  2. Price and config: 256 GB, 80 GPU, 8 TB
  3. CPU tests: Speedometer and build times
  4. Storage: sequential and random beyond 2×
  5. Bandwidth: STREAM and GPU micro bench
  6. Architecture: per-core neural accelerator and Metal path
  7. LLMs: PP and TG on DeepSeek and GPT-OSS
  8. Power and heat: 400W peak, 43°C and fan noise

AI commentary

"My read: M5 Ultra is a serious raw jump, especially for coding agents on long context. For short chat the gain is subtler; to justify a $14k config you need to feed it large contexts and heavy models."

AI assessment

Strength: the video sets a transparent protocol — same builds, same models, same drive size and same power probes. That makes the 1.4× bandwidth → 1.5× token generation correlation read as measurement, not marketing, and showing MoE vs dense on the same day is genuinely useful.

Limits: 1.6× PP on short prompts sits where a daily chat user may barely feel it, and the 400W+ peak with 43–44°C thermal cost is non-trivial; efficiency is up only ~12% while acoustics and heat rise, which can clash with the quiet-studio promise.

Takeaway: M5 Ultra is not just faster, it is Apple rewriting the Metal path for local inference with a per-core neural accelerator . That leaves a mark even in cross-platform stacks like Llama.cpp and pushes Apple, via the coming 512 GB option, deeper into large-model territory for long-context agents.

Practically: if your work is running coding agents on large contexts, frequent 140 GB loads and concurrent serving , you will pocket the gain; for short-prompt chat and light tasks M3 Ultra remains ample and cooler/quieter — let prompt length and model size decide the upgrade.

Sources

8 links; 2 of them also cited by 4 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

m5 ultra · m3 ultra · mac studio · apple silicon · llm · memory bandwidth · storage

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…