Yovao News · The World, In Focus. From Local to Global, Never Miss a Beat

PrismML Shrink Large Language Models to Fit on Personal Devices

PrismML Shrink Large Language Models to Fit on Personal Devices

PrismML, a startup founded by California Institute of Technology researchers, is challenging the assumption that capable reasoning language models must be massive. Backed by a $22.25 million seed round, the company is developing compression technologies that allow high-performing models to run on personal computers and smartphones.

On Thursday, PrismML released Bonsai 2 27B, its latest model in a growing family. The new release compresses Qwen3.8 27B, an open-source model from Alibaba, down to just 5.9 GB—a reduction of nine to ten times the original memory footprint. This size makes the model viable for PCs and potentially high-end mobile devices.

The company’s unique approach involves ternary weights, which reduce the standard 16 bits required for each weight to just three possible values: +1, -1, or 0. CEO Babak Hassibi, a Caltech professor specializing in compression, stated that this method preserves nearly all of the model’s original intelligence. Bonsai 2 matches 98% of Qwen’s aggregate benchmark scores, an improvement over the first Bonsai release in March, which achieved 95% parity.

Download numbers reflect strong interest in the technology. The initial Bonsai model has been downloaded over 11 million times, with an additional 2.6 million downloads recorded for the startup’s smaller variants. Hassibi noted that while perfect benchmark parity may remain theoretical, the marginal performance difference is unlikely to impact real-world utility, especially as surrounding software infrastructure continues to improve.

PrismML counts former Databricks co-founder Ion Stoica and Caltech among its advisors and backers, alongside Khosla Ventures and Cerberus Capital. Stoica highlighted the potential for local AI to offer both privacy and accessibility, allowing users to run advanced models on hardware they already own without sending data to the cloud.

Looking ahead, Hassibi expects the team to apply these compression techniques to models with several hundred billion parameters within the next few months. He suggested that larger models may actually yield better compression results with fewer intelligence losses. While rumors suggest PrismML is in discussions with Apple, Hassibi declined to comment on specific partnership talks.

Leave a Reply

Your email address will not be published. Required fields are marked *