Brains without a cloud bill: Nemotron moves home
The Nemotron Labs episode on the NVIDIA Developer channel opens with a striking claim: some of the most capable open models now run in a box in your living room. The host centers the Nemotron 3 family and walks through turning these models into a live service on DGX Spark and DGX Station. The official playbook on Build.nvidia.com makes the same promise: a production-style inference server on a single DGX Spark behind an OpenAI-compatible API . In other words, a real endpoint you can talk to with curl and wire into agents and IDE tools.
First, the family itself. According to the Nemotron 3 page on Research.nvidia.com, the family was announced on December 15, 2025 and has three members: Nano, Super, and Ultra. All three share a hybrid Mamba-Transformer MoE architecture that promises far higher throughput at matching or better accuracy than classic Transformers. The shared traits are ambitious too: hardware-aware LatentMoE expert design on Super and Ultra, multi-token prediction layers that speed up long-form generation, training in NVFP4 precision, and context windows up to 1 million tokens . Post-training brings multi-environment reinforcement learning plus inference-time reasoning budget control.
Three siblings: Nano, Super, and Ultra
Little sibling Nano reads like a lesson in efficiency. Per the figures on the Research family page, the model uses only 3.2 billion active parameters (3.6 billion with embeddings) for a total size of 31.6 billion. Yet it outscores far larger models such as GPT-OSS-20B and Qwen3-30B-A3B-Thinking-2507. On a single H200, at 8K input and 16K output, it reaches 3.3x the throughput of Qwen3-30B-A3B and 2.2x that of GPT-OSS-20B. It also beats both rivals on RULER across long contexts, and its weights ship openly in FP8 and BF16 alongside training recipes and the Nemotron-CC datasets.
Middle sibling Super, released March 10, 2026, raises the bar. According to the Super page on Research, the model is a 12-billion-active, 120-billion-total parameter MoE: the first in the series to use LatentMoE, carry MTP layers, and pretrain in NVFP4. The numbers are blunt: 2.2x the throughput of GPT-OSS-120B and 7.5x that of Qwen3.5-122B at 8K input and 64K output, while also leading on RULER at 1-million-token context. NVFP4, FP8, and BF16 checkpoints plus training data are open, so the results leave the lab and become downloadable.
A supercomputer on the desk: Spark and Station
The hardware is half the story. According to the DGX Spark page on Nvidia.com, the box delivers 1 petaFLOP of AI performance at FP4 through the GB10 Grace Blackwell superchip, and its 128 GB of unified system memory supports inference on models up to 200 billion parameters plus fine-tuning up to 70 billion. ConnectX networking links up to four Spark systems for scaling to 700 billion parameters . The compact, power-efficient design targets always-on agent workloads.
Big brother DGX Station moves the bar from desktop toward datacenter. Per the Station page on Nvidia.com, the machine pairs the GB300 Grace Blackwell Ultra desktop superchip with 748 GB of coherent memory and 20 petaFLOPS of AI compute, claiming local development, fine-tuning, and inference for models up to 1 trillion parameters . An 800-gigabit ConnectX-8 SuperNIC can link two Stations together. It ships with a preconfigured Ubuntu-based software stack with CUDA-X libraries, plus BMC telemetry and Redfish support so IT teams can manage fleets remotely.
So how does this get installed in practice? The answer lives in the NVIDIA/dgx-spark-playbooks repository on GitHub and the step-by-step recipes on Build.nvidia.com. For Nano the recipe is simple: vLLM inside Docker (v0.20.0 image), local Nemotron-3-Nano Omni weights, and a server on port 8000. For Super there are two paths: vLLM with a reasoning parser or TensorRT-LLM (1.3.0rc9) with trtllm-serve and an extra config file. Both expect a Linux terminal, Docker GPU access, and a Hugging Face token for gated weights; first-time setup takes tens of minutes dominated by downloads, but risk is low since everything stays in user space.
The real message of the episode's second half is agents. Per the June 2026 blog post on Developer.nvidia.com, NVIDIA is shortening the path from unboxing to local agent: with the June 2026 system software, over-the-air updates are no longer forced during first setup, and the open-source NemoClaw plan installs model, agent harness, and the sandboxed OpenShell runtime with one command. Model performance climbs with Qwen3.6, and teams that outgrow one box get a guided multi-node cluster setup. The logic is clear: sensitive context stays on device, per-token costs drop to zero, and control over the agent stays with the user.
The bottom line: the line between renting cloud and buying a desktop supercomputer is getting thinner. Spark fits privacy-sensitive teams, developers building always-on agents, and researchers fine-tuning up to 70 billion parameters; Station targets labs pushing the trillion-parameter frontier. Either way the first bill is the hardware itself plus download waiting time: the models are open, but electricity, disk, and maintenance are on the owner. Local AI is maturing, but the wallet math still matters.
Key moments
AI commentary
"Running large models without a cloud bill or sending data off-site finally looks practical, and the numbers here favor it. Still, read with the NVIDIA badge in mind: the figures impress, but independent verification is a must."
AI assessment
The strongest counterargument sits in the cloud: an hourly-rented H200 runs the same Nano model more flexibly, with no purchase cost, electricity, or maintenance burden, and upgrades are one click away when hardware ages. The 3.3x and 7.5x throughput multipliers the host cites are also NVIDIA's own measurements; independent reproductions across varied workloads remain limited to the selected comparisons on the Research pages. What 4-bit formats like NVFP4 take away in accuracy varies by task and deserves caution, especially at long context.
Gaps exist too. The episode never mentions desktop realities like power draw and fan noise; how loud an always-on agent box gets in a living room is a mystery. Pricing is absent from official pages and reseller listings swing. Gated weights needing a Hugging Face token and license terms get skipped; not every model downloads in one click. And this article is itself a web synthesis: the episode's spoken wording was unreachable, so the content is compiled from official recipes and product pages.
The speaker's interest is also plain: this is an NVIDIA Developer production whose goal is plainly ecosystem invitation, selling DGX and growing the NIM and CUDA stack. That does not falsify the numbers, but it shapes the frame: rival hardware, AMD- and Apple-based local rigs, even lightweight llama.cpp builds stay extras in this story. NVIDIA Research figures are company technical reports, not peer-reviewed papers.
The practical takeaway for readers: if your team already builds agent infrastructure, keeps data off the cloud, and works in the 200-billion-parameter band, Spark is a serious candidate with lower trial-and-error cost than cloud. But for anyone running big models a few times a month, renting still wins. Before deciding, benchmark your own workload, see and hear power and noise in person, then open the wallet.
Sources
8 links; 1 of them also cited by 1 other story. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — NVIDIA Developer
- @build.nvidia.com Build.nvidia.com — Nemotron on DGX Spark playbook
- @developer.nvidia.com NVIDIA Developer Blog — local AI agents on DGX Spark
- @research.nvidia.com NVIDIA Research — Nemotron 3 family
- @research.nvidia.com NVIDIA Research — Nemotron 3 Super
- @nvidia.com Nvidia.com — DGX Spark
Also cited by: Decisions as Retrieval: CLM-8B, the 13x Faster Architecture
- @nvidia.com Nvidia.com — DGX Station
- @github.com GitHub — dgx-spark-playbooks
nemotron · dgx spark · dgx station · local ai · open models · nemoclaw