Back to feed

Eight B300s Meet Kimi K3: Real Numbers From a 128-User Server Test

BIZON benchmarked its eight-B300 X9000 server with Kimi K3: 100 tokens a second for one user, 64 at 32 users, 48 at 64, and 38 at 128; the 64-user point at 3,000 tokens overall takes the sweet spot.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — 5rGBUCFJJrk
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

The BIZON team put one of six customer-bound servers on the test bench before shipping it. The machine is the X9000: a DGX-class 5U box carrying eight B300s. Systems like this are both extremely expensive and poorly covered, so the video fills a genuine gap. Its value is not commentary but meters on screen: power, temperature, and per-user speed stream side by side.

The tested configuration pairs two 128-core Intel Xeons with 2 TB of memory, and the host jokes that RAM prices sting these days. Eight B300s handle the compute. The vendor page lists this machine from 472,233 dollars with 12 units in stock, so a single box costs roughly a mid-size company's annual tech budget.

The 5U chassis splits into two decks. Up top sit the fan wall and GPUs on rails; slide the cover and each card's enormous heatsink block appears. CPUs and memory live below, NVMe bays line the sides. The rear carries the InfiniBand ports for linking servers, with 800 Gb/s XDR per card and 1.8 TB/s GPU-to-GPU NVLink bandwidth per the vendor spec.

Some context on the B300: Blackwell Ultra silicon, 288 GB of HBM3e per card, and a 1,100 W envelope in this HGX form. The video quotes around 280 GB usable per card, which squares with the official figure. One distinction matters: the 1,400 W number belongs to the liquid-cooled GB300 rack, not to an air-cooled HGX box like this one. Mixing the two wrecks any power and cooling math from the start.

The workload is Kimi K3, Moonshot's open-weight flagship: a 2.8-trillion-parameter mixture-of-experts design with a million-token context window. It loads sharded across all eight GPUs at about 266 GB per card. So a model file around 1.6 TB fits in VRAM together with the full million-token window. That is the video's core claim: running a model of this class at full size takes a box from this league.

Under the hood runs vLLM, the open-source serving engine Nvidia itself recommends for multi-user inference. The team bolted a plain test UI on top: fire a prompt and watch generation speed, utilization, power, and temperature on one screen. The simplicity matters, because every claim here rests on observation, not on a synthetic scoreboard.

A lone user gets 100 tokens a second. The crew mentions briefly touching 280 with extra tuning but calls that setup unstable, so they report the raw 100 baseline, which strikes me as the honest call. In that run the GPUs sit at full utilization drawing 596 W, hovering near 52 degrees. Longer pulls push some cards toward 75-76 degrees, a sensible curve for air cooling from a 44-degree idle.

Then it gets serious: 32 simultaneous users, as if dozens of people hammered the same chatbot at once. Per-user speed eases to 64 tokens a second with nobody stalling; even user 32 keeps streaming. Aggregate output climbs to 2,058 tokens a second at 614 W. Calling that still good is no exaggeration: serving 32 times the users while keeping two-thirds of the speed is a strong result.

At 64 users the rig delivers 48 tokens a second each and 3,000 a second overall, with power nearing 700 W. The engineer calls this the sweet spot, and I agree: 48 tokens a second feels perfectly fluid in chat. Note the power curve rising step by step with headcount, from 596 W to 700 W across these runs.

At 128 users each still gets 38 tokens a second, inside the usable band. The 256-user attempt saturates the cards, and the video's key metric steps forward: P95 latency, the true threshold of how many people one box can carry at once. The bottom line is clear: an eight-B300 box like this is not a single-user machine but a box for dozens of people leaning on the same model. Configurable CPU and memory options sit on the vendor's page.

Visualization: nodesdaily AI

Per-user speed falls as headcount grows

  • 1 user100 t/s
  • 32 users64 t/s
  • 64 users48 t/s
  • 128 users38 t/s
Per-user generation speed from one to 128 users.
UsersEachTotalPower
1100 t/s100 t/s596 W
3264 t/s2,058 t/s614 W
6448 t/s3,000 t/s700 W
12838 t/snot sharednot shared

AI commentary

"This is the most naked look I have seen at a B300 server's meters, and 48 tokens each at 64 users convinced me: this league is about carrying crowds, not sprinting solo."

AI assessment

Steel first: for the price of one 472,000-dollar box you could rent API calls for years. Kimi K3's own tariff is 3 dollars per million input tokens, falling to 0.30 dollars on cache hits, roughly a tenth of rival closed models. Unless you are a regulated shop whose data cannot leave the building, this box's math is tough. Open weights are genuinely attractive, but the bill lands up front.

The methodology has gaps: one short prompt, one model, brief runs. The 256-user result never appears on screen; a P95 graph shows up but the threshold number is never stated. There is nothing on long-queue stability, thermals under multi-day load, or the power and cooling bill. Without those, the numbers sit somewhere between a sales demo and field data.

On verifiability: this is a vendor video, filmed by the team selling the box. The 280-token tuning is waved away as unstable with no measurement detail; the 100-token baseline looks honest but wants independent replication. The 280-versus-288 GB gap is not an issue: one is usable, the other is nameplate capacity. Read every figure through that lens.

My practical call: if you serve one model to dozens of people at once, this video is the evidence you needed. Forty-eight tokens at 64 users and 38 at 128 slot straight into a capacity plan. For a lone user or an occasional tinkerer the box is oversized; the same money rents cloud and still runs smaller open models in-house. I would decide on the P95 threshold, not the headline speed.

Sources

6 links; 1 of them also cited by 1 other story. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

b300 · blackwell ultra · kimi k3 · gpu server · vllm · ai inference

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…