What if a 54GB AI model could be compressed to just 4GB without any loss in efficiency?
That is the guarantee behind PrismML, an AI startup that has attracted attention for a model-compression technique designed to make large language models (LLMs) small enough to run directly on smartphones.
Founded in March 2026, PrismML has developed what it describes as an advanced mathematical approach that reduces the size of AI models while preserving their capabilities.
The company holds exclusive commercial rights to patents developed by the California Institute of Technology (Caltech).
To demonstrate the technology, PrismML compressed Alibaba’s 27-billion-parameter Qwen model from about 54GB to roughly 4GB, allowing it to run locally on an iPhone during company demonstrations.
The development has reportedly caught Apple’s attention and the iPhone maker is said to be in the early stages of discussions with PrismML as it explores ways to bring more advanced AI processing directly onto its devices instead of relying on cloud servers.
What It Could Mean for Future iPhones
Technology companies are working towards AI features that can run entirely on users’ devices, even without an internet connection.
The challenge is that today’s large AI models require far more memory and computing power than smartphones typically offer.
For example, even a premium smartphone with 12GB of RAM cannot natively run a 54GB AI model. Memory limitations, battery consumption and bandwidth all present significant technical hurdles.
PrismML believes its compression technology can help overcome those barriers by shrinking AI models to a fraction of their original size.
According to the company, this makes it possible for models that would normally require much larger hardware to run on devices with as little as 6GB of RAM while maintaining usable performance.
The startup also says its Bonsai family of models can deliver processing speeds five to eight times faster while reducing energy consumption by as much as 80%.
It attributes these gains to a technique that enables hardware to process complex decimal values more efficiently, reducing the amount of computation required.
Although Apple has not announced any agreement with PrismML, the reported discussions suggest the company is continuing to invest in more capable on-device AI. Given that the talks are still at an early stage, the technology is unlikely to appear in the upcoming iPhone 18 lineup. Even so, it points to the direction Apple could take in future generations of its devices.
Looking Beyond PrismML
PrismML is not the only company pursuing more efficient AI models. Google recently introduced TurboQuant, while Microsoft has continued to advance its BitNet framework. Several startups are also competing in the growing field of AI model compression and efficient on-device computing.
Whether Apple ultimately reaches an agreement with PrismML is still being discussed. If the partnership moves forward, however, it could strengthen Apple’s goal to deliver more powerful AI experiences directly on the iPhone while reducing dependence on cloud-based processing.



