Zhang Yiming’s Strategic Retreat to Advance

08/07 2026 399

Leveraging the 'Resource Gap' to Close the 'Technology Gap'

Author|Qingyun

Editor|Xiaobai

Produced by|Qiangdiao Next

On August 6, ByteDance CEO Liang Rubo admitted at an all-company meeting that Seed’s large language model was falling further behind leading international models. He urged the team to accept a temporary lag and persist in independent research and long-term optimization. Earlier, at an internal meeting, Zhang Yiming explicitly opposed the practice of distilling competitor models.

On the same day, LatePost reported that ByteDance was considering training a large model with over 5 trillion parameters. For context, Alibaba’s Qwen 3.8-Max and Yuezhi’anmian’s Kimi K3 have parameter counts of 2.4 trillion and 2.8 trillion, respectively. The plan is still in its preliminary discussion phase and may not materialize.

These two signals may appear contradictory, but in reality, the former resets the timeline for catching up, while the latter defines the approach: ByteDance is not seeking to rapidly close the gap by relying on competitor models’ outputs. Instead, it aims to focus computational power, data, and organizational resources on a larger-scale pre-training effort.

This is a costlier and less predictable path. Parameter scale merely indicates the direction of investment, not the final model capabilities. What ByteDance truly aims to test is whether it can convert its resource advantages into a training system and organizational capability capable of independently generating cutting-edge results.

───

01

 A Bold Bet on 'Leapfrogging Competitors Several Times Over' ■

According to reports, the over-5-trillion-parameter model will be led by Xiang Liang, head of Seed Foundation, in collaboration with Shen Ke, who oversees pre-training data for large language models. Both have backgrounds in ByteDance’s search, advertising, and recommendation systems. To support this initiative, Seed is redefining roles and reallocating resources.

The direct impetus for this plan was Seed’s underperformance in the first half of the year. The February release of Seed 2.0 drew limited market attention, while domestic models like GLM-5 and Kimi K3 made significant progress in programming, tool invocation, and complex tasks.

Rather than trying to win while catching up, ByteDance is opting for a bolder strategy.

The gamble is first evident in engineering complexity. For models using a Mixture of Experts (MoE) architecture, total parameter count, per-token activated parameters, training compute, and inference costs are distinct metrics. While 5 trillion parameters can expand model capacity, they do not automatically resolve issues like data quality, training stability, or emergent capabilities. Without simultaneous improvements in architecture, data, and training methods, the result may simply be a more expensive model.

Thus, the 5-trillion-parameter target is more about resource allocation than a guaranteed technical breakthrough. Xiang and Shen’s backgrounds in search, advertising, and recommendation systems also suggest that this task involves not just algorithmic research but also large-scale training systems, data engineering, and cross-team collaboration. The larger the model, the more any inefficiency in any link is magnified.

───

02

 Lagging Behind Is Already Impacting Revenue ■

Another driver of ByteDance’s urgency is its financial performance.

First, consider the technological gap. In mid-February, Seed 2.0, a key model launched under Wu Yonghui’s leadership, was officially released to limited market response. Around the same time, Zhipu’s open-source GLM-5 was praised as the first domestic model comparable to Anthropic’s Opus series. Kimi K3, released in July, was deemed by multiple third-party evaluations to approach overseas closed-source flagship models.

The most critical gap is in coding. During the Lunar New Year, Anthropic quickly gained traction among programmers and B-end markets with the coding capabilities of its Opus series. According to public reports, its Annual Recurring Revenue (ARR) soon approached and surpassed OpenAI’s. Zhipu and Yuezhi’anmian, leveraging their improving coding capabilities, achieved ARR breakthroughs of $1 billion and $300 million, respectively. Chinese tech giants then realized they had missed the coding window.

Now, consider ByteDance’s revenue structure. Volcano Engine is currently the largest seller of model APIs in China, holding roughly half the market share, with revenue of about 15 billion yuan in 2025 and an internal target exceeding 40 billion yuan this year. However, according to LatePost, over half of the token consumption for Doubao’s large model comes from two multimodal generation models, Seedance and Seedream, with language models accounting for a smaller share. As the short-drama industry’s growth peaks, Seedance’s token consumption and revenue growth have slowed. Doubao’s token consumption was 120 trillion in March and 180 trillion in June, falling short of the original target of 250-300 trillion.

In short, ByteDance’s AI revenue is imbalanced. Multimodal generation earns from content consumption, with demand tied to industry cycles like short dramas and e-commerce materials. Language models, especially coding, earn from productivity, serving developers and enterprises with higher stickiness and average revenue per user. The 40-billion-yuan revenue target cannot be met with video generation alone; the language model gap must be closed.

───

03

 No Distillation: Shouldering the Catch-Up Costs Alone ■

Two weeks ago, Zhang Yiming and Seed head Wu Yonghui attended a Seed all-hands meeting. According to media reports, Zhang acknowledged the difficulty of training large models and stated that the team could accept a period of lag. He recognized coding as the current key direction but also cautioned against letting all research resources be driven by a single scenario.

His stance on distillation was clearer: Seed should not rely on the outputs of competitor models to improve.

Here, two approaches must be distinguished. Distillation and synthetic data are mature model training techniques, commonly used by leading labs to generate data with their own models for training subsequent versions. What ByteDance opposes is using closed-source models like Claude or open-weight models like Kimi K3 as “teachers” to batch obtain outputs and then use this data to fill gaps in reasoning, coding, and tool invocation capabilities.

This principle did not emerge overnight. As early as April 2023, ByteDance mandated that GPT-generated data not be included in its model training sets and conducted API call checks and output similarity sampling to detect violations. Subsequently, multiple internal discussions at Seed on leveraging external model distillation were rejected.

Rejecting competitor distillation means ByteDance abandons a faster catch-up path. Discussing 5 trillion parameters means it is willing to bear higher costs in compute, data, and time. Both decisions align with the same goal: not to approach competitors’ capability curves but to leap to the next capability platform and wait for them there.

This strategy has succeeded once for ByteDance in video models. In 2025, the prevailing view in video generation was that further scaling pre-training along the DiT architecture offered limited gains, leading most teams to focus on post-training. Kuaishou Kling attempted to train larger models but retreated. Seedance took the opposite approach: pushing pre-training to the extreme, it created the first video generation model fully adopting MoE architecture, with 200 billion parameters. Launched in February 2026, it was widely recognized as the world’s highest-performing video model and became the revenue backbone of Volcano Engine’s MaaS business.

However, this case does not directly prove that a 5-trillion-parameter language model will succeed. Seedance bet on a direction where consensus was fading, with fewer competitors and a core algorithm team of just over ten people. Scaling language models is a route pursued by OpenAI, Anthropic, xAI, and leading Chinese labs. Without a clear information advantage, ByteDance can only gamble on whether its resource gap can translate into training efficiency and model capabilities.

───

04

 'Brute Force Miracle 2.0': From Racing to High-Stakes Betting ■

To wage this battle, ByteDance is overhauling its famed organizational approach.

ByteDance’s “brute force miracle” strategy previously relied on racing: deploying multiple teams in parallel to explore the same direction and then concentrating resources based on market feedback. Its early batch launches of news apps and later simultaneous entries into short video with Volcano, Douyin, and Xigua exemplify this mechanism. The essence of racing is trading team quantity for success probability, but it requires low trial costs and rapid market feedback.

Cutting-edge large models do not work this way. Single training runs cost hundreds of millions of yuan, with feedback cycles spanning years. Having ten teams pursue ten directions equates to underinvestment in each.

According to media reports, ByteDance leadership is pushing Seed to eliminate internal racing in the same direction, consolidate R&D resources, clarify team roles, and break down silos. To bolster coding capabilities, Zhang Yiming personally invited Guo Daya, a core researcher from DeepSeek, to join Seed and lead specialized training. Related R&D resources were also unified under him. It was under this strategy that resources from Volcano Engine, Feishu, and Doubao were integrated.

This organizational adjustment addresses scattered investment but does not eliminate research uncertainty. With fewer races, each route receives more resources, making the cost of misjudgment higher. For ByteDance, the 5-trillion-parameter model is both a technical and organizational trial: it requires a product company accustomed to rapid trial-and-error to adapt to longer-cycle, slower-feedback foundational research.

Zhang Yiming said ByteDance “should be willing to sacrifice some short-term gains for long-term goals.” The true test of this statement lies not in model rankings but in revenue statements.

Liang Rubo said at the all-hands meeting: The market potential for large models is vast, so economic returns are not a concern, “but revenue still matters.” The 40-billion-yuan annual target, below-expectation token growth, and over-reliance on multimodal revenue make accepting technological lag easy but accepting the resulting revenue gap hard.

For now, all that is certain is that ByteDance has chosen a more expensive catch-up method. Whether 5 trillion parameters can unlock a new capability platform will be answered jointly by the model, its clients, and revenue.

- END -

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.