THE SINGULARITY LABS • AI NEWS

Apple in talks with PrismML, a startup that shrinks AI models to run on an iPhone

CNBC Technology
QUICK SUMMARY

The startup says it compressed Alibaba's Qwen3.6 27B model from roughly 54GB to under 4GB, letting its new Bonsai 27B model run fully on-device on Apple hardware via MLX at around 11 tokens per second, and its CEO says…

DETAILED EXPLANATION

6 27B model from roughly 54GB to under 4GB, letting its new Bonsai 27B model run fully on-device on Apple hardware via MLX at around 11 tokens per second, and its CEO says Apple is evaluating the technology. The compressed model reportedly retains more than 95% of full-precision benchmark performance while using up to 15 times less memory.

0 license, and a deal could accelerate Apple's push to run powerful AI locally on its devices.