Can a single RTX 3090 run a 27-billion-parameter model well enough to replace its full-precision twin? That is the question Digital Spaceport puts on the bench, using PrismML's ternary Bonsai 27B. Viewers keep asking what the best model is for one card or two small cards, and the episode tests that directly. The headline claim frames the test: does a nine-times smaller build preserve almost all of the Qwen3 27B capability?
What Ternary Means — Three Values, Nine Times Smaller
Ternary sounds simple: each weight lives in one of three states, minus one, zero, plus one, instead of two. That third state trims quantization error per scale group. In the numbers PrismML publishes, a 27B model that occupies about 54 GB in 16-bit drops to about 5.9 GB in ternary and 3.9 GB in binary. Roughly a ninefold shrink, which puts the full model comfortably inside a 24 GB card's budget.
The shrink is not a thin wrapper. Bonsai builds on Qwen3 27B, shifts around three-quarters of its attention layers to linear form, and keeps a 262K token window workable on device. The language stack is compressed end to end — embeddings, attention projections, MLPs, and the LM head share the same low-bit representation with no high-precision bypass. The vision tower ships separately in 4-bit HQQ so screenshots and documents travel with text. A DSpark speculative decoder sits on top as a lossless draft-and-verify boost.
The Rig — Proxmox LXC and a Single 3090
The setup is a familiar self-hosting rig: a Proxmox 9 LXC container numbered 118 on the Prox 2 host. The Hugging Face download is quick because the artifact is small enough to skip the key. Launch goes through PrismML's llama server starter from the demo repo. A key detail is that mainline llama.cpp does not yet carry these low-bit kernels upstream, so you must pull PrismML's fork for correct outputs. Environment wiring binds BONSAI_HOST to 0.0.0.0, leaves contact size on auto, and keeps speculative and KV-cache squeeze off for this run.
Once loaded, the card shows about 16.772 GB occupied, leaving headroom inside 24 GB. The server becomes reachable on 192.168.1.61:8080. Early chat throughput sits near 68 to 69 tokens per second and holds there for a while. For reference, the prior full-precision Qwen3 27B run on four RTX 3090s with vLLM landed near 40 tokens per second. Context is set at 131K; bumping it makes sense for Hermes Agent style multi-step loops. The vision path stays enabled.
Arcade Test — Same Inputs, Different Outputs
The core test is a fair head-to-head: the same arcade package previously generated with full-precision Qwen3 27B is generated again with Bonsai. The Hermes-style agent engages, thinking traces appear, prompt processing kicks around 1,250 tokens per second. One generation runs about 25,000 tokens, prompt work is negligible, and decode settles near 64.4 tokens per second. Results are slated for side-by-side play at arcade.digitalspaceport.com.
Each of the three mini games tells a different slice. The shooter moves only sideways, enemy flow feels excessive, score sits tiny at about 16,000 in the upper left, and the boss fires the wrong way. Grades there land in the C-minus to D-plus range. Space Racer feels stiff, lane changes are harsh, the tiny font does not help, and controls feel close to unplayable. Pixel Breaker lands better: the fence and cat vignette is decent, music feels corny yet gameplay is smoother, not full-screen but the breakable stars react as expected.
Is It 98 Percent — The Field Note
The final grade in the video is blunt: two fails and a D-plus, with quality clearly behind the full-precision run. The host says it does not feel like 98 percent, and the quirks in creative layers — level logic, music, small score differences — support that. The story lines up with lab numbers too: about 80.49 versus 85.0 on a 15-benchmark average, roughly 94.6 percent retained. Lab average and single creative generation are not the same; the latter is harsher.
Where Speed and Tool Use Shine
The bright spot is speed and tool use. Chat feels fluid, non-code tool calls land well, and no missed calls were seen in this run. Single-stream decode near 68 tokens per second on a 3090 translates to up to 134 tokens per second ternary and 163 in 1-bit on a 5090, and around 58 and 87 on an M5 Max in the published figures. DSpark's lossless 1.34x lift is part of that picture. It is enough to make a single card feel viable for a persistent single-thread assistant.
The boundary is drawn at sustained coding agents. The episode states it is fine for chat and single-thread agent flows, but not yet a good time for agentic code development. The reason is the quality drop seen in creative and structural code generation; even with long context and linear attention, coherence across levels slips. There is room to improve as the forked kernels land upstream and the context budget grows, and the two 3060 12 GB setup gets its own speed check.
Practical take ties to hardware prices. The channel's live board shows used RTX 3090s near 1,400 dollars and climbing, while a 3060 12 GB sits near 300 dollars. The dual 3060 configuration is tested separately for speed. A single 3090 already sits at the practical threshold for memory and throughput. Picking this build makes sense if your plan is an assistant, document reading, and light automation on one card; making fully autonomous code agents or vision-heavy game generation the main workload is early.
Average Score Comparison
- Qwen3 27B FP1685.0
- Ternary Bonsai80.49
- 1-bit Bonsai76.11
| Variant | Size | Average | Retained |
|---|---|---|---|
| Qwen3 27B FP16 | 54 GB | 85.0 | 100% |
| Ternary Bonsai | 5.9 GB | 80.49 | 94.6% |
| 1-bit Bonsai | 3.9 GB | 76.11 | 89.5% |
Key moments
AI commentary
"What stands out to me is how Bonsai pushes density from paper to a single card; still, you should weigh the creative generation drop before trusting it for sustained coding agents."
AI assessment
Steel-manning PrismML, the lab claim is strong: about 80.49 versus 85.0 on a 15-benchmark average for ternary, roughly 94.6 percent retained, and 76.11 around 89.5 percent for 1-bit. Math near 93.4 versus 95.3, coding near 86.0 versus 88.7, and STEM knowledge near 77.0 versus 83.1 keep the gap narrow on the exact skills agentic workloads need. The narrative of hybrid attention and end-to-end compression holds: about three quarters linear layers plus 4-bit KV cache keep a 262K window workable on device. In that frame, telling a ninefold shrink without dramatic loss is defensible.
The arcade field test shows the lab average does not map one-to-one to creative generation. Across three games, level logic, collision, and font scaling slip; the boss shoots the wrong way, score placement stays tiny, and play turns repetitive. One 25K token generation is a thin sample for broad claims, yet two weak grades and a D-plus suggest the honest marketing frame is 94 to 95 percent rather than 98. The run also used a forked llama.cpp build and 131K context; upstream landing and a larger window could close the gap, but the fork dependency is a cost today.
On verifiability, the core numbers line up across sources: 54 GB to 5.9 GB, a 16.77 GB occupancy reading, about 68 to 69 tokens per second on a 3090, and published ceilings near 134 ternary and 163 1-bit on a 5090, 58 and 87 on an M5 Max. These are repeatable in independent rigs, not carved in stone. A second moving piece is hardware pricing: used 3090s near 1,400 dollars have ticked up in recent days, which changes the single-card 27B math tomorrow. Method gaps remain too: no multilingual, long-document, or extended tool-chain test in this episode, so the arcade alone does not prove agentic competence.
Practically, the pick hinges on the job. For daily chat, document Q&A, vision-aware light automation, and single-thread Hermes-style agent loops, Bonsai 27B is already usable: it runs locally, data stays on device, marginal cost is zero, and speed clears the bar. If sustained coding agents, long-horizon level generation, or high-fidelity creative output is the main workload, full-precision Qwen3 27B or a frontier cloud model is still the safer bet. The dual 3060 12 GB path is a budget-friendly plan B; just do not expect an out-of-the-box ride without accounting for drivers and the forked kernels.
Sources
7 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — Is Bonsai 27B REALLY 98% of Qwen 3 27B fp16?
- @prismml.com https://prismml.com/news/bonsai-27b
- @docs.prismml.com https://docs.prismml.com/models/bonsai-27b
- @huggingface.co https://huggingface.co/prism-ml/Ternary-Bonsai-27B-gguf
- @huggingface.co https://huggingface.co/prism-ml/Bonsai-27B-gguf
- @gate.com https://www.gate.com/news/detail/prismml-releases-bonsai-27b-39-gb-ai-model-runs-on-iphone-22620494
- @digitalspaceport.com https://digitalspaceport.com/gpus
bonsai 27b · ternary 1.58-bit · qwen3 27b · rtx 3090 · llama.cpp