Overview
- Meituan has published LongCat‑2.0 weights and code for a 1.6 trillion‑parameter model that the company says is built for Agentic Coding and native 1M‑token contexts.
- The model uses several architecture changes to support long context and efficiency, including a learned sparse attention (LSA), N‑gram embeddings, a sparse cross‑layer MoE (ScMoE) and dynamic expert activation.
- Meituan released multi‑precision variants (BF16, FP8 and INT8) and open‑sourced inference optimizations targeting memory, bandwidth and scheduling limits of domestic accelerator cards.
- Moore Threads announced Day‑0 full‑stack adaptation on its MTT S5000 GPU and MUSA software, saying it completed model loading, operator and engine optimizations plus deployment and accuracy validation.
- If the companies' claims hold up under independent tests, the combination of open code and vendor adaptations could cut the time and hardware cost to run very large, long‑context models on China’s existing compute stack.