ByteDance is in early-stage discussions to train a model with more than 5 trillion parameters, a scale that would make it the largest known AI model in China by a wide margin.
ByteDance is in early-stage discussions to train a model with more than 5 trillion parameters, a scale that would make it the largest known AI model in China by a wide margin.

ByteDance is in early-stage talks to train a model with more than 5 trillion parameters, surpassing Alibaba's Qwen 3.8-Max and Moonshot AI's K3 to become China's largest known AI model by parameter count.
The plan, reported by Chinese media outlet LatePost citing sources familiar with the matter, would be led by Xiang Liang, head of ByteDance's Seed Foundation, in collaboration with Shen Ke, who oversees large language model pre-training data.
The proposed model would more than double the parameter count of Moonshot AI's K3, which launched with 2.8 trillion parameters, and exceed Alibaba's Qwen 3.8-Max at 2.4 trillion. Seed is reorganizing its structure, clarifying responsibilities, and allocating resources for the initiative. The plan remains at an early stage and does not guarantee the final model will be released.
A 5-trillion-parameter model would represent a significant escalation in China's frontier AI ambitions, potentially reshaping the competitive balance among ByteDance, Alibaba, and Moonshot AI. It also raises the stakes for compute procurement and training costs, which scale non-linearly with parameter count.
Training a model at this scale carries enormous infrastructure demands. Anthropic, which trains models with roughly 4.2 trillion code tokens alone, operates more GPU capacity than all Chinese manufacturers combined, according to a senior employee at a large Chinese tech company cited by Digital Intelligence Frontline. ByteDance would need to secure comparable compute resources, a significant challenge given US export controls on advanced chips.
The parameter race is not purely academic. Larger models generally exhibit higher intelligence, though the relationship between scale and capability has diminishing returns. Moonshot AI's K3, released earlier this year, demonstrated that a 2.8-trillion-parameter model can achieve meaningful gains in coding and reasoning benchmarks. The company's daily API sales increased at least sixfold after K3's launch, according to third-party industry data.
The move would intensify competition across China's AI industry. Alibaba's Qwen family and Moonshot AI's Kimi have been locked in a race for model supremacy, with DeepSeek's R1 release in 2025 resetting industry expectations for what Chinese models can achieve. Zhipu AI has also emerged as a contender, with its GLM-5.1 model achieving coding capability roughly equivalent to Anthropic's Claude Code.
ByteDance's Seed Foundation has been expanding its AI footprint beyond consumer applications. The company's Trae coding assistant reached approximately 8 million users by June, and its Seedance 2.0 video generation model launched in February. A 5-trillion-parameter foundation model would position ByteDance to compete directly with Alibaba and Moonshot AI at the frontier of Chinese AI research.
The plan's early-stage status means the model may never ship. But the organizational restructuring at Seed suggests ByteDance is committing serious resources to the effort. For investors tracking China's AI sector, the key question is whether ByteDance can secure the compute infrastructure needed to train at this scale, and whether the resulting model justifies the enormous training cost. Hong Kong-listed tech names and AI infrastructure providers could see sentiment shifts if ByteDance accelerates compute procurement, though the speculative nature of the plan limits near-term fundamental impact.
This article is for informational purposes only and does not constitute investment advice.