The TikTok parent is pretraining a model of up to 10 trillion parameters, the Financial Times reports — a size that would hand ByteDance the largest AI system ever built in China, and by a margin of three to one.
Moonshot’s Kimi K3 holds the domestic record today. Should ByteDance’s figure hold up, its run would be three times that.
How it stacks up against the American labs
At ten trillion, ByteDance lands in the same territory as Anthropic’s top system, Mythos 5, which outside industry estimates peg at roughly eight trillion parameters. One caveat is worth holding onto: Anthropic has never published its own figures, so those numbers are guesswork from beyond the company’s walls.
Nor is ByteDance the only lab operating at that scale. Elon Musk has said xAI is training Grok variants at six trillion and 10 trillion parameters on the Colossus 2 cluster.
Parameter counts are the least interesting number here
What parameters govern is how much a model can hold, not how well it actually performs. That comes down largely to the quality of the data and the training method — the reason a smaller model can beat a larger one, and often does.
The sharper detail in the FT’s account concerns technique. According to one source, ByteDance has spent more than a year without resorting to distillation, meaning it has avoided training on the outputs of rival companies’ models. That is the harder path for a Chinese lab to take, and it is the sort of claim that only gets settled by benchmarks months down the line.
What Zhang Yiming told the team
Internally, founder Zhang Yiming has instructed the 2,000-strong Seed team to pursue world-leading model capabilities over the long run. That is a directive to staff, not a release date.
Three insiders told the FT the model is currently in pretraining, a phase that generally runs three to six months. In practice, that means no one outside ByteDance will be able to judge what 10 trillion parameters actually bought until well into that stretch.


















STAY ALWAYS UP TO DATE