Overview
- The Financial Times reported Friday that ByteDance is training a model that could reach about 10 trillion parameters, but the FT said it could not independently verify the claim and ByteDance has not commented.
- Sources cited by the coverage say the system is in the pre-training phase, a months-long stage that typically takes three to six months before a model is fine-tuned or released.
- At the scale reported, the model would approach industry estimates for Anthropic’s Mythos class (roughly 8 trillion parameters) and be far larger than recent Chinese models such as Moonshot AI’s 2.8 trillion-parameter Kimi K3.
- Parameter counts are a rough shorthand for scale and do not directly measure performance or safety, and many leading firms do not publish precise parameter totals, which complicates direct comparisons.
- If the report proves accurate, expect scrutiny on the compute, energy and hardware needed to train such systems, potential effects on product offerings, and renewed debate about governance and access to frontier models.