Particle.news

Xiaomi Cuts MiMo‑V2.5 API Prices Up to 99%

Xiaomi says hierarchical KV caching and a 1:7 Full:SWA sparsity design sharply lower inference costs, with a technical blog promised to show details.

Overview

  • MiMo has permanently lowered prices for the MiMo‑V2.5 API by as much as 99 percent and removed separate charges for different context window lengths.
  • Xiaomi attributed the move to inference‑engine and cache optimizations, saying its framework now supports hierarchical KV caching targeted at SWA and overlapping cache reads between attention modules.
  • Company tests reportedly show cached‑token capacity rose about fivefold, which Xiaomi equates to roughly an 80 percent reduction in cache cost and input/output price cuts of about 60 to 80 percent.
  • MiMo said its production inference fleet is nearing full utilization at the new prices and that the business can essentially break even given the stated architecture and infra efficiencies.
  • Xiaomi framed the cuts as passing structural cost savings to developers and warned that similar price moves are only sustainable for providers with comparable model sparsity and inference infrastructure; the team will publish a detailed technical blog to substantiate the claims.