ByteDance is discussing training a large language model with more than 5 trillion parameters, a scale that would leapfrog Alibaba's Qwen 3.8-Max at 2.4 trillion and Moonshot AI's Kimi K3 at 2.8 trillion, according to people familiar with the plan. The move marks an aggressive bet that raw scale, not imitation, can close the gap in coding and reasoning where ByteDance's Seed 2.0 drew a muted response after its February release.
"Distillation can improve model performance in the short term, but it is essentially copying capabilities Claude already has — along that path you can only approach the leader, never truly surpass it," Zhang Yiming, ByteDance's founder, told the Seed research team at an all-hands meeting two weeks ago, according to people who attended. He said the company could accept temporarily falling behind rivals rather than rely on the technique, which the White House has accused Chinese labs of using at industrial scale.
The new model would be led by Xiang Liang, head of the Seed Foundation team, working with Shen Ke, who oversees large-language-model pretraining data, and would be the largest known model under development in China. ByteDance is reorganizing Seed to consolidate resources and eliminate competing internal projects, a shift from its long-standing "raw power works wonders" approach of running multiple parallel bets. The plan remains early-stage and may not result in a release.
The stakes are commercial as well as technical. Volcano Engine, ByteDance's cloud arm, is China's largest model API seller with roughly half the market, generating about 15 billion yuan in 2025 revenue against an internal 2026 target above 40 billion yuan. Yet Doubao token consumption reached 180 trillion in June, below the 250 trillion to 300 trillion goal, with video and image models Seedance and Seedream accounting for more than half of usage — a concentration that exposes growth to a slowdown in short-drama demand.
Why scale, and why now
ByteDance's bet echoes the path that made Seedance 2.0 the first video model to fully adopt a mixture-of-experts architecture at 200 billion parameters, launched in February and widely rated the strongest video generator globally. The model, with gross margins of 70 percent to 90 percent, became the revenue base of Volcano Engine's model-as-a-service business and proved that users flow to the most capable product, one Seed insider said.
The language-model gap is the problem. Zhipu's open-source GLM-5 is widely seen as the first Chinese model comparable to Anthropic's Opus series, while Moonshot's Kimi K3 has approached overseas closed-source flagships in third-party benchmarks. Zhipu and Moonshot have pushed annual recurring revenue past $1 billion and $300 million respectively, powered by coding strength. ByteDance recruited DeepSeek core researcher Guo Daya to lead coding-specific training, consolidating related resources under him.
Zhang's anti-distillation stance lands as Washington escalates scrutiny. White House Office of Science and Technology Policy director Michael Kratsios accused Moonshot of covertly distilling Anthropic's Fable model to build Kimi K3, with Anthropic tracing 3.4 million Claude conversations to fake accounts. Treasury Secretary Scott Bessent has floated sanctions if Chinese companies adopt distillation at scale, even as Stanford's Institute for Human-Centered AI finds top US and Chinese models roughly equal despite the US outspending China on AI by 23 to one.
The enterprise payoff
ByteDance's broader restructuring ties the model bet to its office software. The Feishu product team merged into the Doubao team under Zhao Qi, with Feishu's Xie Xin reporting to him, while the go-to-market unit folded into Volcano Engine. Feishu generated more than 3 billion yuan in 2025 revenue, with second-quarter 2026 growth above 100 percent year on year, as Doubao becomes the task entry point and Feishu supplies identity, permissions, and workflow context.
For investors, the question is whether ByteDance can convert a 5-trillion-parameter training run into revenue before rivals do. Volcano Engine's token growth already missed its internal target, and the model's success depends on solving the systems-engineering challenge of pretraining at a scale no Chinese lab has attempted. ByteDance's private valuation and its capital expenditure — the largest of any Chinese AI player — mean the bet carries no public-market check, but a breakthrough would pressure Alibaba's Qwen, Moonshot, and Zhipu, which are racing to monetize coding and reasoning demand.
This article is for informational purposes only and does not constitute investment advice.