Overview
- The model launched in late August and this week was publicly identified as Zhipu AI’s MIT‑licensed GLM‑5.3‑Flash, which has been widely used through gateways such as OpenRouter.
- Publishers report per‑token prices roughly an order of magnitude below frontier list rates — an international per‑token ratio often cited as about 1/40th of Claude Opus 4.8’s list price — making long contexts and many agent loops much cheaper on paper.
- Independent tests this week show Flash can match flagship accuracy on several tasks but often uses far more tokens and more time on hard problems, so actual cost per completed task depends on workload, retries, and standard (not promotional) rates.
- Adopters will need engineering work to capture savings, including prompt and result caching, selective routing between models, validation layers, and budgeting against the end of promotional discounts.
- The open MIT release and reported large‑scale serving on Chinese silicon raise practical concerns about export controls, self‑hosting, infrastructure choices, and the governance of widely distributable, low‑cost models.