Zhang Yiming: Strategic Retreat as a Path Forward

08/07 2026 371

Leveraging the 'Resource Gap' to Outmaneuver the 'Technology Gap'

Author | Qingyun

Editor | Xiaobai

Produced by | Qiangdiao Next

On August 6, ByteDance CEO Liang Rubo admitted during a company-wide meeting that Seed's large language model is falling further behind leading overseas models. He encouraged the team to accept a temporary lag and persist in independent research and long-term optimization. Earlier, at an internal meeting, Zhang Yiming explicitly rejected the idea of distilling competitor models.

On the same day, LatePost reported that ByteDance is considering developing a large model with over 5 trillion parameters. In comparison, Alibaba's Qwen 3.8-Max and Yuezhi'anmian's Kimi K3, as cited in the report, have parameter counts of 2.4 trillion and 2.8 trillion, respectively. This plan is still in the early discussion phase and may not materialize.

These two signals, seemingly contradictory, actually serve different purposes: the former resets the timeline for catching up, while the latter defines the strategy. ByteDance is not looking to rapidly close the gap by relying on competitor models' outputs. Instead, it aims to focus computational power, data, and organizational resources on a larger-scale pre-training effort.

This approach is costlier and more unpredictable. The parameter scale merely indicates the direction of investment, not the final model capabilities. ByteDance's true goal is to verify whether it can convert its resource advantages into a training system and organizational capabilities that can independently generate cutting-edge results.

───

01

 'Leapfrogging Peers by Multiples' - The Bold Bet ■

According to reports, the over-5-trillion-parameter model will be led by Xiang Liang, head of Seed Foundation, in collaboration with Shen Ke, who oversees pre-training data for large language models. Both come from ByteDance's search, advertising, and recommendation systems. To support this initiative, Seed is redefining roles and reallocating resources.

The direct impetus for this plan is Seed's underperformance in the first half of the year. The February launch of Seed 2.0 received a lukewarm market response, while domestic models like GLM-5 and Kimi K3 made significant advancements in programming, tool invocation, and complex tasks.

Rather than trying to win by simply catching up, ByteDance is taking a bolder gamble.

The gamble is first evident in engineering complexity. For models using a Mixture of Experts (MoE) architecture, total parameter count, per-activation parameter count, training computation, and inference costs are distinct metrics. While 5 trillion parameters can enhance model capacity, they do not automatically resolve issues like data quality, training stability, and emergent capabilities. Without concurrent upgrades in architecture, data, and training methods, the outcome may simply be a more expensive model.

Thus, the 5-trillion-parameter figure is more about resource allocation than a guaranteed technological breakthrough. Xiang and Shen's backgrounds in 'search, advertising, and recommendation' also highlight that this task involves not just algorithmic research but also large-scale training systems, data engineering, and cross-team collaboration. The larger the model, the more any inefficiency in any link is magnified.

───

02

 Lag Begins to Affect Revenue ■

Another driving factor behind ByteDance's urgency is its financial situation.

First, consider the technological gap. In mid-February, Seed 2.0, a key model launched under Wu Yonghui's leadership, was officially released but received limited market response. Around the same time, Zhipu's open-source GLM-5 was praised as the first domestic model comparable to Anthropic's Opus series. Kimi K3, released in July, was deemed by multiple third-party evaluations to be approaching overseas closed-source flagship models.

The most significant gap is in coding capabilities. During the Spring Festival, Anthropic quickly captured the programmer and B-end market with Opus's coding abilities, with its Annual Recurring Revenue (ARR) soon approaching and surpassing OpenAI's, according to public reports. Zhipu and Yuezhi'anmian, with their continuously improving coding capabilities, saw ARRs exceed $1 billion and $300 million, respectively. Chinese tech giants then realized they had missed the window for coding.

Now, let's examine ByteDance's revenue structure. Volcano Engine is currently the largest seller of model APIs in China, holding roughly half the market share, with revenue of about 15 billion yuan in 2025 and an internal target of over 40 billion yuan this year. However, according to LatePost, over half of the token consumption for Doubao's large model comes from two multimodal generation models, Seedance and Seedream, with language models accounting for a smaller share. As the short drama industry's growth peaks, Seedance's token consumption and revenue growth slow. Doubao's token consumption was 120 trillion in March and 180 trillion in June, falling short of the original target of 250-300 trillion.

In short, ByteDance's AI revenue is imbalanced. Multimodal generation earns from content consumption, with demand tied to industry cycles like short dramas and e-commerce materials. Language models, especially coding, earn from productivity, serving developers and enterprises with higher stickiness and average revenue per user. The 40-billion-yuan revenue target cannot be met by video generation alone; the language model gap must be closed.

───

03

 No Distillation: Shouldering the Catch-Up Costs Alone ■

Two weeks ago, Zhang Yiming and Seed head Wu Yonghui attended a Seed all-hands meeting. According to media reports, Zhang stated that training large models is inherently difficult, and the team can accept a period of lag. He acknowledged coding as the current key direction but cautioned against letting all research resources be driven by a single scenario, as programming is just one of several hot areas.

His stance on distillation was clearer: Seed should not rely on outputs from competitor models to improve.

Here, two approaches must be distinguished. Distillation and synthetic data are mature model training techniques, with leading labs commonly using their own models to generate data for training subsequent versions. What ByteDance opposes is using closed-source models like Claude or open-weight models like Kimi K3 as 'teachers' to batch obtain outputs, then using this data to supplement its own reasoning, programming, and tool invocation capabilities.

This principle did not form overnight. As early as April 2023, ByteDance mandated that GPT-generated data not be added to its model training sets and conducted API call checks and output similarity sampling to root out non-compliant use. Subsequently, multiple internal discussions at Seed on leveraging external model distillation were rejected.

Rejecting distillation of competitor models means ByteDance abandons a faster catch-up path. Discussing 5 trillion parameters means it is willing to bear these costs with more computational power, data, and time. Both decisions share the same goal: not to approach competitors' capability curves but to leap to the next capability platform and wait for them there.

This strategy has succeeded once for ByteDance in video models. In 2025, the mainstream view in video generation was that further scaling pre-training along the DiT architecture offered limited improvement, leading most teams to focus on post-training. Kuaishou Kling attempted to train larger models but retreated. Seedance did the opposite: pushing pre-training to the extreme, it created the first video generation model fully adopting MoE architecture, with 200 billion parameters. Launched in February 2026, it was widely recognized as the world's top-performing video model and became the revenue backbone of Volcano Engine's MaaS business.

However, this case does not directly prove that a 5-trillion-parameter language model will succeed. Seedance bet on a direction where consensus was fading, with fewer competitors and a core algorithm team of just over ten people. Scaling language models is a route pursued by OpenAI, Anthropic, xAI, and leading Chinese labs. ByteDance lacks a clear information advantage and can only bet that its resource gap will translate into training efficiency and model capabilities.

───

04

 'Brute Force Miracle 2.0': From Racing to Heavy Betting ■

To wage this battle, ByteDance is revamping its renowned organizational approach.

ByteDance's 'brute force miracle' strategy previously relied on racing: deploying multiple teams in parallel to explore the same direction, then concentrating resources based on market feedback. Its early batch launches of news apps and simultaneous entries into short video with Volcano, Douyin, and Xigua exemplify this mechanism. The essence of racing is trading team quantity for success probability, but it requires low trial costs and rapid market feedback.

Cutting-edge large models do not work this way. Single training runs cost hundreds of millions of yuan, with feedback cycles spanning years. Having ten teams explore ten directions means underinvestment in each.

According to media reports, ByteDance leadership is pushing Seed to eliminate internal racing in the same direction, consolidate R&D resources, clarify team roles, and break down departmental silos. To bolster coding capabilities, Zhang Yiming personally invited DeepSeek core researcher Guo Daya to join Seed, leading specialized training. Related R&D resources were also unified under him. It was under this strategy that resources from Volcano Engine, Feishu, and Doubao were integrated.

This organizational adjustment addresses scattered investment but does not eliminate research uncertainty. With less racing, each route receives more resources, making the cost of misjudging the direction even higher. For ByteDance, the 5-trillion-parameter model is both a technical and organizational trial: it requires a product company accustomed to rapid trial-and-error to adapt to longer-cycle, slower-feedback foundational research.

Zhang Yiming said ByteDance 'should be willing to sacrifice some short-term gains for long-term goals.' The true test of this statement lies not in model rankings but in revenue statements.

Liang Rubo said at the all-hands meeting: The market potential for large models is vast, so economic returns are not a concern, 'but revenue still matters.' The 40-billion-yuan annual target, below-expected token growth, and over-reliance on multimodal revenue make accepting technological lag easy but accepting the resulting revenue gap hard.

For now, all that can be confirmed is that ByteDance has chosen a more expensive catch-up method. Whether 5 trillion parameters can unlock a new capability platform awaits answers from models, clients, and revenue.

- END -

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.