Ternary Bonsai 2 makes large language models more accessible on entry-level hardware

Dragos Roua showcases how highly compressed Ternary Bonsai 2 enables running large AI models on affordable laptops, signalling a significant step forward for on-device AI, despite performance caveats.

Dragos Roua has begun a weekly video series testing local AI models on an ordinary computer, with the aim of showing how far private inference can go on entry-level hardware. The latest instalment focuses on Ternary Bonsai 2, a heavily compressed version of Qwen 3.8 27B that the creator says is small enough to run on a 16GB M1 MacBook Pro. According to the model notes published by CanItRun.dev and Andrew OOO’s guide, the ternary build reduces the model to roughly 6-7.2GB, compared with around 50GB for the full 27B variant, which makes it a far more realistic fit for consumer machines.

That size reduction is the key point. Local AI models usually have to fit within available memory, with some headroom left for the operating system and the runtime. LocalClaw’s model listing says Bonsai 27B has a minimum RAM target of 16GB in its ternary configuration, while practical performance still depends on context length, backend choice and available system headroom. In other words, the model may load on a 16GB laptop, but the usable experience can still vary sharply depending on the task.

Roua said PrismML released Ternary Bonsai 2 only hours before he recorded the video, and he initially regarded it as a striking example of how quickly local AI is improving. But the longer he used it, the more he concluded that the marketing claim of near-total quality retention should be treated cautiously. That scepticism is consistent with the broader benchmark material: while several recent write-ups say the model preserves around 98% of the reference performance, those same sources also frame that figure as dependent on the test set and the exact quantisation method.

Other summaries suggest the model is indeed technically notable. A report from Nowline says PrismML has reduced Qwen3.8 27B from 53.8GB in FP16 to 5.93GB in ternary form, while still targeting roughly 98.2% of baseline performance and support for either a 16GB laptop or a single 24GB GPU. That is why the model is attracting attention beyond this single demo: it illustrates how aggressively quantised weights are making large language models usable on machines that would previously have been ruled out.

Even so, the practical limits remain obvious. Roua says one Three.js generation task ran for hours without producing a useful result, which underlines a broader point about local AI: shrinking a model is not the same as making it reliable for every workload. The result is still meaningful, though. On the evidence assembled across the model pages and recent guides, Ternary Bonsai 2 represents a genuine step forward for sovereign, on-device AI, even if its real-world performance is less flawless than its promotional claims suggest.

Disclaimer: This content is intended for informational purposes only. Readers are advised to exercise their own judgement, conduct due diligence, or consult a qualified expert before acting on any information provided.