Back to feed

Before Buying the $12,000 Mac Studio, Rent a Test for $2 an Hour

Before buying the $12,000 Mac Studio M5 Ultra, the presenter rents a 96GB RTX Pro 6000 for $2 an hour and puts an open-weights model through three real jobs. Two tests end in draws, the frontier model wins the trap question, and the decision rule sharpens: rent first, measure, then buy.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — jSUPoEOSPqA
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

When a twelve-thousand-dollar box will not arrive before January, the question changes its order: instead of buying first and learning later, rent for two dollars an hour, learn first, and decide later. That is exactly what the presenter does; he makes no secret of wanting the new Mac Studio M5 Ultra, yet before opening his wallet he puts an open-weights model in the 27-32 billion parameter class through three real jobs on a rented 96GB graphics card . The result is the surprising kind: the local model is slow but gets the job done, and the real issue is knowing which job belongs to which model. This piece follows the order of that test, checking every figure against Apple, Nvidia and current store listings.

The $12,000 question asked with rented 96GB

The stage for the experiment is Vast.ai: the presenter rents an RTX Pro 6000 instance for about $1.50-$2 an hour, pausing or deleting it when idle, just like a server on DigitalOcean. According to GPUTable, the same card's spot price sits between $0.29 and $0.67, with on-demand between $0.94 and $1.49; Vast.ai's own listing shows similar bands, rising to $2.09 on RunPod's secure cloud. So the video's two dollars is a reasonable upper band on the safe side. He does not even do the setup himself; through Codex computer use he connects his Vast account and registers the model as a custom provider inside Open Code. The detail matters because it knocks down the technical wall in front of local-model testing: renting, setup and running all collapse into a few steps.

Why 96GB? Because that is the entry ticket for this weight class. The model heard as Quinn in the video is actually Alibaba's Qwen family; according to ResearchAudio, Qwen3 32B needs about 28GB with INT4 quantization at 32K context and 83GB in BF16, while the Q4_K_M build on Ollama comes down to roughly 20GB. So with quantization, 96GB is generous room for this model; even at full precision it fits at the edge. According to Apple, the M5 Ultra starts near $7,300 with 96GB and climbs into the $12,000 band at 256GB. The presenter's choice of the 96GB RTX Pro 6000 is no accident: both candidates carry the same memory, so the duel runs on equal terms.

Three real job tests

The first job is a customer support email: messy helpdesk output, strict rules, a mandatory subject line. The presenter feeds the same input to the local model and to the newest frontier model of the day; according to Anthropic, Opus 5.5 shipped on September 22, 2026, stronger at coding and agentic work and roughly 40 percent cheaper to run than its predecessor on typical workloads. The result is a draw: the local model reaches the same verdict in 3 minutes 12 seconds — a one-month refund or a guided path to the data. The speed gap is wide, the verdict gap is zero. For wait-tolerant queues like support, that table says the rented local model is good enough.

The second job is a Monday-to-Friday 20-hour production plan: leftover budget, sponsor balance, a Friday emergency, shoots plus six LinkedIn posts. The local model builds the same skeleton here too: a 22-hour calendar, a no-list, a contractor message, a list of unknowns. What stands out is that the model asks open questions instead of inventing what it does not know. This is where the presenter's async agent idea is born: if you will not sit and wait, slowness stops being a flaw; for an agent that sorts expenses at 2am, audits a channel, or closes the month, the difference between 3 minutes and 30 seconds is meaningless.

Out of this comes a mental model: interactive shaping work — kneading a product, writing a campaign, live coding — belongs to fast frontier models, while routine, repeated, time-insensitive work belongs to the local model. The presenter names agent harnesses such as Hermes and OpenClaw, and deliberately uses Open Code as the worst showcase: a chat box you sit and stare at, where a slow model looks weakest. Yet the results come out even, which strengthens the thesis: change the showcase to a background agent and the local model's value multiplies. The triple rationale is crisp: cost, privacy, ownership.

What makes the M5 Ultra special is bandwidth, not capacity

According to Apple, the M5 Ultra fuses two M5 Max chips into one package; memory bandwidth reaches 1.2 TB/s, nearly double the M5 Max at 614 GB/s. Unified memory comes in 96, 256 and 512GB options, with the 512GB tier arriving late October. The presenter's excitement sits right there: not more gigabytes, but how much data moves in and out of memory per second. That number decides local-model inference; capacity answers whether it fits, bandwidth answers how fast it flows. For this weight class 96GB is enough; speed is decided by the hose attached to it.

Three hardware paths lead to the same 96GB at different invoices. According to Nvidia, the RTX Pro 6000 tops the desktop class with 96GB of ECC GDDR7, 1,792 GB/s of bandwidth, 600W of power and 4,000 TOPS of AI compute. On the B&H Photo (bhphotovideo) listing the card shows at $10,499; the $15,670 figure quoted in the video should be read as a momentary seller-dependent price, with the current listing lower. The second path pools three 32GB RTX 5090 cards: according to Newegg a single card ranges from $5,300 to $9,800, so the trio passes $18,000, plus a case, networking, and inter-card transfer worries. The third path is the Mac Studio's unified memory: everything in one pool, no transfer headache. Choosing between single-card simplicity and a scalable trio is really one preference: failure surface or transfer speed.

The trap question and the third judge

The third test is a trap: a single-option marketing recommendation is demanded, invented rates are banned, and every missing figure must be written down as unknown. Here the frontier model wins clearly. The presenter hands both outputs to a third model as judge, on four criteria: constraint adherence, mathematical accuracy, reasoning strength, evidence-backed recommendation. The judge's ruling: Opus honors the unknown-means-unknown rule, its experimental logic holds, it strips the Meta-versus-YouTube duplication, while the local model smuggles in a few unsupported assumptions. Score: two draws, one loss. But the presenter's question lands: in which jobs is that quality gap worth the money?

The cost math closes the experiment: the presenter spends about $2.82 of a ten-dollar credit, while the $12,000 box would sit idle through weekends and nights. At a few real hours a day, renting never reaches even a fraction of buying. The rented card doubles as a rehearsal before the purchase decision: without seeing the difference between being squeezed onto one card and expanding in unified memory as model classes grow, twelve thousand dollars stays committed to nothing. Privacy and ownership are the bonus: no data leaves, no token fees, the rehearsal is nearly free.

The verdict: which box, for whom

The presenter does not buy the $12,000 M5 Ultra; his eye is on a roughly $7,000 M4 Mac Studio with 128GB: half the price, more memory, slower — yet slowness carries no verdict in async work — plus a 4TB drive. His advice is a protocol: rent from Vast.ai for an hour, attach a local model with an Open Code or Hermes harness to your real work for $2, see what you can offload, then decide. Image generation, voice-to-text, video editing and computer use can all move to local models too; that is the subject of future episodes. The thesis in one line: rent first, measure, then buy — and when you buy, look at bandwidth and memory math, not prestige.

Visualization: nodesdaily AI
OptionPriceMemoryBandwidth
M5 Ultra 256GB~$12,000256GB unified1.2 TB/s
RTX Pro 6000$10,499 + case96GB GDDR71.79 TB/s
3x RTX 5090 32GB~$18,000+96GB pooledPCIe bottleneck
Vast.ai rental~$2/hr96GBon demand

Key moments

  1. The $12,000 question: M5 Ultra nowhere before January
  2. Vast.ai setup: RTX Pro 6000 at $2 an hour
  3. Codex does the setup, Open Code connects
  4. Test 1: support email ties in 3 minutes
  5. Test 2: 20-hour week plan comes out even
  6. Bandwidth: 1.2 TB/s, twice the speed
  7. Three hardware paths converge on 96GB
  8. Trap question: Opus wins, Codex judges
  9. Verdict: 128GB M4 and the $2 rehearsal rule

AI commentary

"Renting before buying may be the smartest hardware idea this year. The presenter's two-dollar experiment turns twelve-thousand-dollar desire into a measurable decision, and in my view it drags the local-model debate from showroom hype back to math."

AI assessment

The strongest objection is that a rented single card is not the same thing as an owned Mac. A rental instance carries no availability guarantee; spot instances can be reclaimed, peak-hour prices spike, and data sent to a data center trims the privacy gain. The 96GB duel also flatters the comparison by ignoring the 256GB Mac's headroom; when larger weights arrive, today's tie could become tomorrow's bottleneck. The presenter never quite confronts this; the test is run exactly in the model class that fits 96GB.

The measurement side is thin: no tokens per second, no time to first token, no power draw, no fan noise, no total cost of ownership. There is one model class and three text jobs; image generation, voice-to-text, video editing and computer use are waved through as merely possible. The judging is single-blind too: when the third model is a cousin of the same family, the independence claim weakens; an outside judge is named but never scored.

The presenter's interest is transparent: he genuinely wants the box, and he makes the sponsorship joke himself. Read the M4 praise in that light; the value framing warms the audience toward the cheaper alternative. Even so, the experiment money came from his own pocket, the method is repeatable, and he let the third test lose on purpose; that self-check separates the piece from promotion.

The reader's takeaway is a rule: list the jobs you could move to a local model, run the one-hour two-dollar rehearsal on your real work, and score verdict quality rather than speed. The buying threshold is equally crisp: if you have 20-plus hours of uninterrupted local work per week plus a privacy mandate, buy the box; otherwise keep renting. My verdict: this experiment matters because it shrinks twelve-thousand-dollar desire into a two-dollar measurement, pulling the hardware decision from prestige to arithmetic.

Sources

10 links; 1 of them also cited by 10 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

mac studio · rtx pro 6000 · local models · qwen · vast.ai · memory bandwidth

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…