PrismML: Bonsai 27B Revolutionizes AI on iPhone with 4GB

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
PrismML: Bonsai 27B Revolutionizes AI on iPhone with 4 GB
Bonsai 27B is a comprehensive reasoning model that fits on an iPhone.
PrismML has launched Bonsai 27B, an AI model that, according to them, can operate on a smartphone without sacrificing reasoning or agent capabilities. Apple is reportedly testing this technology.
PrismML, a company founded by a team of researchers from Caltech, has developed Bonsai 27B, an AI model with 27 billion parameters designed to run directly on an iPhone. Bonsai 27B is based on Alibaba's Qwen3.6-27B model. PrismML claims it supports multi-step reasoning, tool usage, image understanding, and agent-based tasks.
PrismML argues that modern AI applications increasingly require powerful models to function locally. An agent can make hundreds of sequential calls to the model, each carrying context, producing structured output, and feeding into the next step. In the cloud, token costs accumulate, each call adds network latency, and intermediate results, tool calls, and private data, such as screen content or documents, all leave the device.
However, running the model on-device reduces the marginal cost of these loops to zero and keeps user data localized. PrismML sees this as the foundation for always-active agents, offline assistants, and hybrid systems. Simple and privacy-sensitive tasks remain on the device, while only the most challenging steps are sent to advanced models in the cloud.
According to a report from CNBC, PrismML is already in talks with Apple regarding the compression technology behind Bonsai. PrismML CEO Babak Hassibi confirmed that Apple and other companies are testing the models for speed, energy consumption, and performance. Discussions are at a "very early" stage, but "things are progressing well."
Two Versions of the Model for Laptops and Smartphones
A model of this size typically occupies around 54 GB of storage. Even with standard compression, it still requires about 18 GB. PrismML offers two much smaller versions: the quality-focused variant occupies about 5.9 GB and is intended for laptops, although the packages currently shipped may be larger depending on execution. The technical document indicates about 7.2 GB for the llama.cpp version and 8.49 GB for the MLX version.
The smaller variant is about 3.9 GB, small enough to fit within the limited storage of an iPhone 17 Pro Max. According to PrismML, an iPhone with 12 GB of RAM actually makes only about 6 GB available for a single application, split between the model and cache.
Instead of storing each neural network weight in 16 bits, PrismML uses only one or just under two bits. In the most aggressive variant, each weight has only two states. In the slightly larger version, three. This approach is applied across the entire language model. As an example of common labeling issues, PrismML cites the construction Qwen3.6-27B-IQ2_XXS compared in the technical document, which averages 2.8 bits per weight despite its "2 bits" label.
The 1 bit Bonsai variant stands out for its efficiency with an intelligence density of 0.530 per GB, far ahead of ternary models and FP16.
PrismML Claims Compression Has Limited Impact on Quality
In PrismML's evaluation across 15 benchmarks, the larger variant retains 95% of the original model's performance. The smaller variant retains 90%. Mathematics and programming remained "virtually unchanged," according to PrismML.
The largest drops were observed with more aggressive compression, particularly in image understanding, instruction tracking, and agent-based tool usage. A conventionally compressed Qwen3.6-27B model at 9.4 GB scores only 72.7 points, while the smaller Bonsai variant at 3.9 GB scores 76.1.
Compressed Bonsai models retain up to 95% of the original Qwen3.6-27B performance. The 1 bit variant shows a more pronounced drop in vision and instruction tracking.
According to the technical document, the smaller variant generates about 11 tokens per second on an iPhone 17 Pro Max. A battery test yielded about 672 tokens generated per percentage point of battery charge, extrapolating to around 67,000 tokens on a full charge. The chip slightly slowed down after just over five minutes.
Apple Could Use This Technology to Catch Up in Local AI
The model weights are available under the Apache 2.0 license. Bonsai 27B runs on Apple devices via Apple's MLX framework and on NVIDIA GPUs. PrismML also offers a limited-time free Developer Preview API and a live demo on HuggingFace.
The company was founded with support from Khosla Ventures, Cerberus, and Google, with ongoing support from Samsung. PrismML plans to apply its compression technology to Google's Gemma model series, with smaller versions already running on smartphones.
Licensed compression technology would also be significant for Apple, whose own models have so far lagged behind competitors in benchmarks. At WWDC 2026, Apple unveiled a revamped Siri based on foundational models developed with Google using Gemini technology. The most powerful on-device model already requires an iPhone with at least 12 GB of RAM, and complex queries run on Nvidia GPUs in Apple's cloud.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.