Updated
Updated · TechCrunch · Sep 17
PrismML Shrinks 27B LLM to 5.9 GB for On-Device AI
Updated
Updated · TechCrunch · Sep 17

PrismML Shrinks 27B LLM to 5.9 GB for On-Device AI

3 articles · Updated · TechCrunch · Sep 17

Summary

  • Bonsai 2 27B compresses Alibaba’s Qwen3.8 27B into a 5.9 GB model, small enough for PCs and potentially high-end smartphones.
  • A ternary-weight method cuts each model weight from 16 bits to three values — +1, -1 or 0 — delivering a 9x to 10x memory reduction.
  • PrismML says the new model retains 98% of Qwen’s aggregate benchmark performance, up from 95% for the first Bonsai release in March.
  • That earlier model has been downloaded more than 11 million times, while PrismML’s smaller models added another 2.6 million downloads, suggesting demand for local AI.
  • Backed by a $22.25 million seed round, PrismML plans several-hundred-billion-parameter compressed models within months as it pushes private, cloud-free AI on user devices.

Insights

If a 54GB model shrinks to 5.9GB, could this compression trick finally untether trillion-parameter AI from massive cloud servers?
PrismML claims 98% performance retention, but what critical reasoning abilities are secretly hidden within that missing two percent?
Local AI promises privacy, but with smartphones being less energy-efficient, will this ternary breakthrough quietly destroy your battery life?