Particle.news

ByteDance Reportedly Discussing a Foundation Model Exceeding 5 Trillion Parameters

The proposal would aim to outpace China’s largest known models by pursuing extreme parameter scale at very high projected cost.

Overview

  • A LatePost scoop relayed by IT之家 says ByteDance is discussing training a foundation model with more than 5 trillion parameters, but the plan is reported as early-stage and unconfirmed.
  • If built at the reported size, the model would exceed Alibaba’s Qwen 3.8‑Max (about 2.4 trillion parameters) and K3 (about 2.8 trillion parameters) to become the largest known model in China.
  • The effort is said to be led by Seed Foundation head Xiang Liang working with pretraining‑data lead Shen Ke, both long‑time ByteDance AI engineers with research backgrounds.
  • Multiple unnamed Seed team staff told reporters the jump in parameter scale is controversial inside ByteDance, with one person calling it “like a gamble,” and founder Zhang Yiming reportedly opposing distillation as a shortcut.
  • Realizing a >5 trillion model would require vast compute, data and engineering effort, so key near‑term signs to watch for are internal approvals, budget and compute commitments, and any formal confirmation from ByteDance.