Back to feed

Beyond a Single Card: How Giant Language Models Run on Hundreds of GPUs

Today's frontier AI models no longer fit on one GPU; this article explains how trillion-parameter models run across hundreds of cards with data, pipeline, tensor, and expert parallelism, and why production systems move prefill and decode into separate pools.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — qZBibWYcKH4
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

The largest AI models today carry over a trillion parameters, needing close to 2 terabytes of memory just to store their weights. Yet the biggest GPU on the market offers about 288 GB — nowhere near enough for these frontier giants. So how do the chatbots used by millions of people every day serve these models reliably? The answer is disarmingly simple: there is no single mega-computer hosting them. They run on a coordinated system of dozens, sometimes hundreds, of GPUs. On the IBM Technology channel, Grace Ableidinger opens up how an LLM inference workload stays alive at that scale, step by step.

Serving a model in production means solving three constraints at once. First, the model memory footprint: every weight must sit in memory before work begins. Second, the working memory known as the KV cache : as the model generates its reply it keeps the conversation context there, and the area grows with every token produced. Third, request throughput: when hundreds or thousands of people hit the model at once, each request queues on the card and everyone's wait grows as the line lengthens. The KV cache is the sneakiest of the three because it never sits still — it swells for as long as generation continues. According to Nvidia's Dynamo work, that swelling context can move off GPU memory into cheaper tiers such as CPU memory and fast disks; validations with Vast and WEKA measured 35 GB/s into a single H100 and 270 GB/s across eight cards.

Meet traffic first: data parallelism

When the model and its cache fit on one card but the user count grows out of control, the simplest fix is copying the whole model onto multiple cards. In data parallelism , each card holds an identical replica and incoming requests route by current load — even by which card already holds the relevant context in its cache. The replicas need no coordination; the job is simply admitting each request through the right door. Per the vLLM team's distributed-inference post of February 17, 2025, low-bit quantization alone stops being enough past hundreds of billions of parameters, which is why the engine offers tensor parallelism inside a card and pipeline parallelism across cards. Traffic, then, is solved with replica counts; fit is solved with splitting techniques.

Copying stops working for truly huge models because the model itself will not fit on one card. A model is really a sequence of transformer layers holding knowledge, and inference means pushing the input through them one after another. Pipeline parallelism cuts the model into layer groups: early layers live on the first card, later ones on the second. Send a single request and most cards idle while waiting their turn, so requests stream through continuously like an assembly line, keeping every station busy. Baseten's engineering notes put a billion parameters at roughly a gigabyte of memory in FP8 precision; the 671-billion-parameter DeepSeek-V3.1 throws an out-of-memory error on a single B200, and even 720 GB across four B200s covers the weights while leaving no room for the KV cache — which usually claims 80 percent or more of whatever space remains after weights. The result: real traffic at this scale wants a full eight-card node.

Two ways to split a model: pipeline and tensor

Instead of cutting the model vertically, it can be cut horizontally. Tensor parallelism splits the math inside each layer rather than the layers themselves: every card takes a slice of the same layer, computes its share, then the cards combine partial results. That combining step is a collective operation, a blocking barrier at every layer — the cards must chatter constantly mid-computation. Such dense conversation only works over a high-bandwidth, low-latency link inside one server; over a slow network the chatter becomes the bottleneck instead of the optimization. DigitalOcean's guide notes a dense 70-billion-parameter model in BF16 needs about 140 GB for weights alone; a model comfortable at 4K context may refuse to fit at 64K or 128K under realistic production traffic. The message is plain: the speed of the road between cards matters as much as the splitting technique.

Then there is the model family that works like a community of specialists: mixture of experts. Instead of running one weight block for every input, these models hold small sub-networks tuned to different sides of generation — one leaning toward code samples, another toward punctuation, a third toward numbers. Expert parallelism spreads those experts across cards so no card holds the entire model. At every layer a router dispatches each token to the handful of experts it picked; current models select about 8 out of 256, and the outputs merge afterward. Compute per token drops sharply, paid for in heavy cross-card dispatch traffic. As Modular's handbook stresses, each data-parallel replica may itself use tensor or pipeline techniques internally, and the router draws a failure boundary by retiring replicas that miss their health checks.

Pool hardware by job type: prefill and decode

The two phases of LLM inference stress hardware in completely different ways. In prefill , the model reads and digests the input, running dense parallel math over fixed weights at compute speed while building the KV cache. In decode , the reply emerges token by token, pulling the whole grown cache out of memory each step at memory-bandwidth speed. Park both phases in one pool and decode's swelling appetite squeezes prefill's territory. The fix moves each phase onto its own GPU pool tuned to its bottleneck — but only pays off if the cache flies between pools almost instantly. Plain TCP over standard Ethernet cannot keep up; the link must be dedicated, low-latency, RDMA-capable fabric. Amazon's SageMaker HyperPod account connects the split pools with EFA and RDMA, tunes time to first token and inter-token latency independently, and keeps long inputs from stalling live generation. Ray's guide adds that the same split scales each phase independently on cheaper, heterogeneous node types, with vLLM's NIXL and LMCache transfer backends carrying the load.

Production systems stack the methods: tensor splits confined to one machine, pipeline stages spanning machines, data replicas absorbing traffic, experts sharded for mixture models, plus pools divided by serving phase. That stacking is called multi-dimensional parallelism. Above everything runs a coordination tier that sends each request to its proper GPU pool, evens out load, and absorbs failed cards without dropping requests. As the presenter sums up, the rule is plain: traffic overflowing, add replicas; model not fitting, split it layer by layer or slice by slice; memory swelling from two different jobs, give each job its own pool. Distributed inference is not stubbornness about cramming giant models into one machine — it is the discipline of cutting work into the right pieces and running each on the right hardware.

Visualization: nodesdaily AI
TechniqueWhat it does
Data parallelismFull model replicas share traffic
PipelineLayer groups split across cards
Tensor and expertIn-layer slices and MoE routing

Key moments

  1. Three bottlenecks: footprint, cache, traffic
  2. Meeting traffic with data parallelism
  3. Pipeline: splitting layers across cards
  4. Tensor parallelism: slicing inside layers
  5. Expert parallelism: 8 experts of 256
  6. Prefill and decode pools
  7. Multi-dimensional parallelism at scale

AI commentary

"The presenter builds distributed inference from production bottlenecks instead of textbook definitions, turning each technique into an engineering decision. The strongest move is tying every method to a single question: did memory fill up, did traffic overflow, or did the nature of the work change?"

AI assessment

The objection comes first: not every model deserves this engineering. A mid-size LLM serves comfortably on one or two cards with reduced precision, while a distributed setup adds communication overhead, observability cost, and failure surface. The per-layer collective in tensor parallelism slows things down instead of speeding them up when the interconnect is weak. The scaling decision should follow the model's size and the traffic's real profile; the video never draws that line with numbers.

The video contains a single measurement: no latency figures, no tokens per second, no invoice numbers. How much each technique actually buys is therefore left to the viewer's intuition. Energy use and card costs go unmentioned, although an eight-card node recommendation floats without a budget line. Production pains like queue overflow and card failure get a passing sentence each.

The presenter's frame is instructive but not neutral: IBM Technology speaks to an enterprise audience, and the chosen examples keep flattering the picture that demands more infrastructure. There is no comparison of inference engines and no pros-and-cons table of open options. That is the genre's nature more than a flaw — yet viewers should listen knowing the story is told in a product ecosystem's language.

The practical order for readers: try data parallelism on a ready-made inference engine first, and measure. If the model will not fit the card, move to pipeline or tensor splits; prefer pipeline when the cards sit on different servers. When the cache overflows on long context, bring up offloading and pool separation. Measure the traffic before anything else; splitting is an expensive answer to an unmeasured problem.

Sources

8 links; 1 of them also cited by 1 other story. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

llm · distributed inference · gpu · kv cache · moe

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…