Particle.news

Zhipu Open‑Sources GLM‑5.3‑Flash After Ox Alpha Preview

Zhipu says the release is meant to lower long‑context serving costs.

Overview

  • Zhipu confirmed that the anonymously surfaced Ox Alpha is its new GLM model and published the GLM‑5.3‑Flash weights while opening API access and limited experience keys.
  • GLM‑5.3‑Flash is described as a native multimodal model with 320 billion parameters and 18 billion activations designed for lower memory use during inference.
  • The company says the model uses a mix of sparse and linear attention plus manifold‑constrained hyper‑connections to cut activation footprint and reduce costs for long contexts.
  • Ox Alpha’s preview on aggregator platforms registered record call volumes, and independent tokenizer‑fingerprint checks gave preliminary support for attribution but full community benchmarking is still pending.
  • Zhipu reported an AA index score of 57 that it says matches a leading proprietary model, claimed all inference ran on domestic chips, and announced aggressive pricing to speed third‑party use while broader validation continues.