Particle.news

Alibaba Open‑Sources Qwen3.8‑Flash and Qwen3.8‑Flash‑Next

The released weights include FP8 quantization and a sparse‑expert prototype that Alibaba says previews the next Qwen4 architecture.

Overview

  • Alibaba’s Tongyi Qianwen team confirmed late Wednesday that it published Qwen3.8‑Flash and the Qwen3.8‑Flash‑Next weights on Hugging Face and ModelScope, with FP8‑quantized versions available for download.
  • Qwen3.8‑Flash is a multimodal MoE model with a 125B main backbone plus a 51B N‑gram embedding and activates about 6B parameters per token, which Alibaba says cuts training cost to roughly one‑ninth of the previous Qwen3.7‑Plus.
  • The release highlights several architecture and system innovations — Gated DeltaNet plus Qwen Sparse Attention for long context, Gated Residuals, Muon Optimizer, asynchronous host memory for large embeddings, and FP8 storage to reduce memory and IO use.
  • The model natively supports 262,144 tokens of context and includes tools to scale to 1 million tokens, and Alibaba plans to offer the model through its Qwen API with published token pricing for input and output.
  • By open‑sourcing a Next‑architecture MoE prototype, Alibaba lets researchers and engineers reproduce results and adapt tooling before Qwen4 arrives, while external validation of the company’s cost and benchmark claims will be required by the community.