Overview
- MiMo has permanently lowered prices for the MiMo‑V2.5 API by as much as 99 percent and removed separate charges for different context window lengths.
- Xiaomi attributed the move to inference‑engine and cache optimizations, saying its framework now supports hierarchical KV caching targeted at SWA and overlapping cache reads between attention modules.
- Company tests reportedly show cached‑token capacity rose about fivefold, which Xiaomi equates to roughly an 80 percent reduction in cache cost and input/output price cuts of about 60 to 80 percent.
- MiMo said its production inference fleet is nearing full utilization at the new prices and that the business can essentially break even given the stated architecture and infra efficiencies.
- Xiaomi framed the cuts as passing structural cost savings to developers and warned that similar price moves are only sustainable for providers with comparable model sparsity and inference infrastructure; the team will publish a detailed technical blog to substantiate the claims.