Particle.news

MoonshotAI Open-Sources Kimi K3, a 2.8 Trillion-Parameter Multimodal Model

MoonshotAI says the release will let researchers reproduce a 2.8T model built for long‑context engineering and multimodal reasoning

Overview

  • MoonshotAI publicly released Kimi K3 on Monday, July 27, publishing model weights, a technical report, and the training infrastructure needed to run it.
  • The open release includes three infra projects—MoonEP, AgentEnv, and FlashKDA—with FlashKDA previously available and MoonEP plus AgentEnv newly posted to Hugging Face and GitHub.
  • MoonshotAI describes Kimi K3 as a 2.8 trillion‑parameter, multimodal model with a 1,000,000‑token context window and architecture changes called KDA and Attention Residuals.
  • The company says K3 expands sparse Mixture‑of‑Experts to 896 experts with 16 active experts to save compute, a technique that routes tokens to a few specialist sub‑networks rather than running every parameter on every token.
  • Vendors moved quickly to adapt the model: Huawei announced same‑day Ascend training and inference support with specific optimizations and native mxFP4 quantized weight support, while independent benchmarks and community verification of MoonshotAI’s performance claims remain pending.