Overview
- PrismML publicly released compressed versions of Alibaba’s Qwen on Tuesday, July 14, 2026, saying it reduced a 27‑billion‑parameter model from about 54 GB to under 4 GB so the full model can fit on iPhone 15 and newer devices.
- The company uses extreme quantization that stores each model weight as one or three possible values rather than 16‑bit numbers to cut memory needs and speed up on‑device inference.
- PrismML claims the compressed models use 10–15 times less memory, respond six to eight times faster, and consume three to six times less energy while losing a few percentage points of accuracy and showing weaker factual recall.
- CEO Babak Hassibi told reporters that Apple and other firms are evaluating the technology and holding early talks but that no partnership or licensing deal has been announced.
- Independent experts say large‑scale testing is needed to confirm real‑world performance on long prompts, multitasking and battery use, and that validated on‑device models could shift chip demand from datacenters to phones and change how companies build assistants like Siri.