Back to feed

One GPU, Four Virtual Machines: Sharing the Card Without Giving It Away

Experimental virtio-nvgpu keeps one RTX 3060 with the host while opening graphics access to four Linux guests at once; the card gave 102.9 fps alone and 103.69 fps combined when split four ways.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — AknK5ixgRSs
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

Hand a whole graphics card to one virtual machine and the host desktop plus every other guest loses access. The host opens with that blunt trade and asks why giving a card to one guest must take it away from everything else. The answer sits inside whole-device passthrough: the device is detached from its host driver for assignment, one guest takes it, and the host plus other guests cannot use it at the same time.

The mechanics match the PCI passthrough guide on archlinux.org: detach from the host driver, assign through IOMMU groups, lose host-side display. Anyone who tried this on a single-GPU laptop knows the pain; the host desktop does not come back until the guest shuts down. The host frames that exclusive ownership as the problem and places the experiment exactly there: access without ownership.

Forward the request instead of the card

The experimental virtio-nvgpu project reverses the path. Its GitHub description says it in one line: a virtio device for near-native NVIDIA access in KVM guests. The card stays with the host while guests forward driver requests to the host driver. What crosses the boundary is not the hardware but operations aimed at the software driving it. That opens separate request paths for a second, third and fourth guest without buying another card.

The most instructive passage follows a single buffer request. An app inside a Linux guest needs a graphics buffer to hold frame data. The NVIDIA user-mode driver inside the guest stays unchanged; a small forwarding module in the guest kernel moves the device-control request across a virtqueue to a host backend. Readers of the QEMU documentation will recognize the shape: frontend in the guest, backend on the host, queue discipline between them.

Before reaching the host driver, the backend translates references. Memory pointers, resource handles and file descriptors may mean different things on each side, so the backend converts them into counterparts the host driver understands. Then a shared memory window opens two views onto the same buffer, one from the guest and one from the host, with no duplicate copy. Once mapped, the documented rendering path submits work through that access; every draw call needs no separate forwarded message.

Keeping the user-mode driver unchanged is the clever part. Applications talk to the driver they know while the forwarding module and backend do the paperwork. A second guest can attach its own forwarding path to the same card. The contrast with the Nvidia-documented vGPU and MIG arrangements matters here: those use time slicing or hardware partitions, while this uses forwarded requests plus mapped memory. The host stresses the separation; guarantees cannot be borrowed from another system.

Four guests, one card: the measured result

The numbers come from an RTX 3060 running the same synthetic load. One guest produced 102.9 frames per second. Four guests running that load each landed near 26 frames per second, totaling 103.69. The TechPowerUp database lists this card at 3584 CUDA cores with 12 GB of GDDR6; in this test the same compute and memory did all the work, nothing multiplied. Four doorbells, one kitchen fits: more ways to ask, same kitchen.

Test conditions are explicit: an 8-second warm-up followed by a 30-second run, the identical job in all four guests. As the Linuxiac summary notes, these are project-reported results with no independent confirmation. More importantly, the outcome promises no fixed quotas, no fairness and no protected slices of graphics memory . Four is the reported guest count, not a measured maximum. Four different demanding applications would be another test entirely.

The second cost exists even with one guest: forwarding overhead. In one synthetic case a 2-millisecond host frame slowed by 1.7% in the guest; in a much shorter case a 0.05-millisecond host frame slowed by 40.8%. The authors attribute much of the short-frame gap to waking the guest after the hardware finishes. A small extra weight looks huge on a tiny job, so percentages must always travel with their original durations. No universal near-native promise follows from two points.

What it fits, what it does not

The stated target is headless streaming: the guest produces an image and serves it as a video stream with H.264 encoding , viewed on a remote screen. Physical display output is out of scope. The readme points to the Vulkan video interface for presentation and encoding inside guests; four lightweight encoding sessions ran at once. That points toward several remote Linux environments with graphics access.

The not-list matters as much. Untrusted guests need secure isolation; the sandboxed helper and multi-tenant infrastructure are unfinished, so joint use proves no separation. CUDA stands at device enumeration only; no completed computation, local inference or training run is shown. Compatibility hinges on the forwarded interface, the driver ABI ; keeping the user driver unchanged does not make every version mix work. A plain single-desktop user gains nothing demonstrated here; adding guests to a saturated card creates no new resources.

Visualization: nodesdaily AI

Key moments

  1. Two guests, one card problem
  2. A buffer request crosses over
  3. One buffer, two views
  4. 103.69 fps across four guests
  5. 40.8% cost on tiny jobs
  6. Who should care

AI commentary

"The measurement discipline carries this story: the host states the numbers plainly and separates what was shown from what was not, without inflating either."

AI assessment

The strongest counterpoint is simple: sharing fixes access, not capacity. Four guests splitting the same synthetic job held total output at 103.69 frames per second while each guest dropped to about 26; that is far from the single-guest 102.9 experience. How four different heavy applications would behave together was never measured, so no capacity lesson can be borrowed from this test.

The missing list is long and openly stated: no fixed quotas or fairness guarantees, no proven protected memory slices, unfinished isolation infrastructure and sandboxed helpers, CUDA support only at device enumeration, Windows and multi-user gaming out of scope. Nvidia documentation describes MIG hardware partitioning with strong isolation, which this experiment does not offer. The QEMU documentation already describes built-in virtio-gpu paths, so dependence on one more driver interface is its own risk.

The speaker's likely interest is guiding experimenters rather than selling setup: no install advice for single-desktop users, a watch list for Linux virtual-machine enthusiasts. That balanced tone builds trust. Still, the project documentation contradicts itself on multi-guest support, which suggests release discipline on GitHub is unsettled; the Linuxiac summary likewise stresses the experimental status.

The practical takeaway is narrow: if several Linux guests need graphics acceleration at once and lightweight remote sessions are enough, this project is worth watching; with 12 GB of memory the RTX 3060 is a reasonable base for such sharing tests according to the TechPowerUp database. If one workload already saturates the card, the classic single-guest passthrough described in the archlinux.org guide remains the predictable route. Before skipping a hardware purchase, ask for measurements of your own application, per guest and combined.

Sources

7 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

virtio-nvgpu · gpu sharing · kvm · rtx 3060 · virtual machines

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…