Seed has its own sense of mission

08/14 2026 454

A seed's mission is to learn how to grow in the soil before its season arrives.

Earlier this year, Seedance 2.0 was released without much hype. Within hours, videos generated using it flooded social media platforms both domestically and internationally: cinematic camera movements, native audio-visual synchronization, and multi-shot storytelling. Even Elon Musk shared a demo on X, leaving the comment: “It's happening fast.”

Prior to this, a man named Wu Yonghui had become a focal point in the industry. A 2001 undergraduate in computer science from Nanjing University, he spent 17 years at Google, became a Google Fellow, and served as one of the chief technical leads for the Gemini large model applications. A year ago, he left Mountain View, returned to Beijing, and took over ByteDance's large model research team, Seed.

When Wu Yonghui took over, the situation was described by the media as follows: a research team of over a thousand people, with over 10 billion yuan invested in catching up over two years, finally developed a foundational model ranked in China's first tier. However, it was quickly surpassed by a model developed by DeepSeek, which had only a hundred people and fewer resources. The department head admitted mistakes, and the company CEO called out the team by name in an all-hands meeting, stating that they could have done better.

A year and a half later, the daily token usage of the Doubao large model surged to 180 trillion, growing over 1500 times in two years. Volcano Engine captured the top spot in China's public cloud MaaS market with a 49.5% share. Seedance 2.0 became a benchmark for global video generation models, with a gross margin estimated by multiple sources to be over 70%, contributing more than half of Volcano Engine's MaaS revenue.

Doubao 2.1 Pro, released in June this year, matched Claude Opus 4.7 in programming evaluations like Terminal Bench 2.1, with a comprehensive usage cost less than 20% of its competitor's. In August, the native audio-video full-duplex large model SeedRealtime was released and fully launched on the Doubao App the same day, enabling 380 million users to engage in video calls with AI, “watching, listening, and speaking” simultaneously overnight.

Behind these numbers lies a quiet organizational transformation. An internet company that rose to prominence through recommendation algorithms, AB testing, and a tournament mechanism spent three years learning to do something it was not initially good at: foundational research. This process involved misjudgments, personnel turmoil, strategic disputes, and moments of being pushed by external circumstances.

The story of Seed is about technology, about people, and even more so about how a company redefines its boundaries during a paradigm shift.

01 The Old Map and the New War

ByteDance's entry into AI large models was not early.

After GPT-4 was released in 2023, the company established a dedicated large model team, led by Li Hang, then head of AI Lab. In August of the same year, the “Yunque” large model completed its filing, five months later than Baidu's Wenxin Yiyan and three months later than Alibaba's Tongyi Qianwen. Zhu Wenjia was transferred from the search business line to lead the team that would become Seed.

This technical executive, described by ByteDance's recommendation algorithm lead Yang Zhenyuan as “one of the top three in Toutiao's algorithm,” had a resume steeped in ByteDance's culture:

He was a principal architect in Baidu's search department, joined ByteDance in 2015 to work on search, took over Toutiao in 2019 and reported directly to Zhang Yiming three months later, grew Toutiao's DAU to 140 million during the 2020 Spring Festival, and was transferred to Singapore in 2021 to support TikTok's technology.

He understood large-scale systems, recommendation algorithms, and how to turn technology into products used by hundreds of millions.

2024 was ByteDance AI's year of application factories. In May, the Doubao large model family was officially released, and Volcano Engine slashed API prices to 0.0008 yuan per thousand tokens, 99.3% lower than the industry average, sparking the first shot in the large model price war.

The Doubao App, relying on ByteDance's mature product methodology, traffic ecosystem, and growth hacking tactics, became China's top AI application by DAU within a year, reaching nearly 60 million monthly active users by November 2024, nearly 50 million ahead of the second-place contender.

This playbook was all too familiar to ByteDance. Over the past decade, from Toutiao to Douyin, from Xigua Video to CapCut, the same logic had been repeatedly validated: rapid iteration, data-driven decisions, AB testing, a tournament mechanism, and saturated investment. Whenever it entered a content or tool segment, ByteDance could always catch up through more refined operations and heavier resource allocation.

Last year, when DeepSeek released R1, the capability curve of reasoning models took a steep leap that winter, but ByteDance didn't even have a corresponding reasoning model product. The user advantage accumulated at the application layer proved fragile in the face of foundational capability gaps.

Thus, Liang Rubo's second round of reflection took place at that year's all-hands meeting. The previous time, he had said the company had become “sluggish,” neglecting Transformer-based language models. This time, he admitted that the follow-up to OpenAI's o1 had been too slow, with the team believing that a month's delay didn't matter.

He made a statement that would later be repeatedly quoted: Being a tech company is not enough; we must be an innovative tech company, not just applying new technologies well but also exploring and inventing them.

This was a crucial cognitive turn. Success during the internet era largely relied on engineering catch-up: identifying a direction and rapidly closing the gap with stronger engineering capabilities and resource investment. However, in the large model era, feedback cycles for recommendation systems are measured in hours, while model capability feedback cycles span months or even quarters. The former can rely on AB testing for incremental progress, while the latter risks months of investment being wasted if architectural choices or training routes are incorrect.

That year, Liang Rubo set a new goal for Seed: pursue the upper limit of intelligence. Intelligence itself became the most important objective, rather than the DAU of any particular product. This reordering of priorities was more fundamental than any organizational restructuring.

02 The Successor and the Implantation of Research Culture

Wu Yonghui officially joined ByteDance in February last year. He became the company's second external executive directly parachuted in at the CEO-1 level, following CFO Julia Guo.

An undergraduate from Nanjing University and a PhD from UCR, he joined Google in 2008 and stayed for over a decade. In the search ranking team, he internalized the habit of serving hundreds of millions of users with large-scale machine learning systems. In 2015, he transferred to Google Brain and was a core author of the GNMT neural machine translation system. Later, he made key contributions to the Conformer speech architecture, GLaM sparse expert model, and CoCa image-text model.

After Google Brain merged with DeepMind, Wu Yonghui was promoted to Vice President of Research and Google Fellow. In the Gemini team, he was one of the chief technical leads for applications, spearheading the implementation of a million-token context window. Google Scholar shows his papers have been cited over 73,000 times.

DeepSeek-R1 was released a month before his joining. The shockwaves had not yet subsided within ByteDance when Liang Rubo completed his reflection at the all-hands meeting, and Wu Yonghui stepped into the spotlight.

In fact, ByteDance's AI organizational adjustments had already begun before Wu Yonghui's arrival: in January 2025, Seed Edge was officially established, targeting long-term research over five years, with evaluation cycles extended to three years and performance reviews allowing “retroactive compensation” for significant outcomes. This was nearly anti-OKR for a company that operated on six-month OKR rhythms.

The five directions—next-generation reasoning, next-generation perception (world models), software-hardware co-designed models, next-generation paradigms (beyond backpropagation and Transformers), and next-generation scaling (Multi-Agent and test-time training)—were not expected to yield clear outputs within a year.

After Wu Yonghui's arrival, he further structured the virtual organization into three layers: Edge for long-term exploration beyond five years, Focus for tackling core bottlenecks in next-generation models, and Base for engineering delivery of current-generation models. Personnel and projects could flow between the layers, with Edge outcomes flowing down and Focus identifying long-term projects for Edge.

He encouraged communication, guided exchanges, and made data and codebases fully transparent internally, breaking down the barriers where teams had previously kept their documentation hidden from each other.

Zhang Yiming invested more personal energy in this than many expected. The founder began one-on-one visits with authors of important AI papers in late 2023, including doctoral students yet to graduate. After Wu Yonghui joined, Zhang Yiming increased his frequency of reviewing technology on the front lines, often attending Seed's key technical meetings and occasionally discussing details with researchers.

However, implanting a research culture did not happen overnight. Wu Yonghui's first major lesson with the team came when training Doubao 2.0, the largest model Seed had attempted, with one trillion parameters and a Gemini-like multimodal architecture.

As parameter scale increased, training stability became an issue. In the words of one researcher, “It's like building a house with an unstable foundation—adding bricks on top leads to problems.” Multiple teams collaborated and spent three months patching the model architecture and training data before finally completing the model before the 2026 Spring Festival.

OpenAI's RL Infra lead Weng Jiayi once said: Every model team's Infra has bugs, and what fundamentally distinguishes model companies is the speed at which Infra fixes bugs, as this determines how many ideas can be validated per unit of time—and ideas can be addressed by increasing talent density.

For ByteDance, personnel turnover and cross-pollination began early. During Seed's formation, the company aggressively recruited globally, bringing in Jiang Lu, head of Google's VideoPoet; Zhou Chang, former technical lead of Alibaba's Tongyi Qianwen; and Huang Wenhao, co-founder of 01.AI.

In the year Wu Yonghui arrived, Zhu Wenjia's reporting line shifted from Liang Rubo to Wu Yonghui, transitioning from a dual-leadership structure to Wu Yonghui dominating foundational research and Zhu Wenjia overseeing model applications. The visual generation lines also underwent adjustments, with text-to-image Seedream, text-to-video Seedance, and 3D generation Seed3D all placed under Zhou Chang's management.

Personnel strengthening continued into 2026. Guo Daya, a core author of DeepSeek R1, joined Seed to lead Coding and Agent directions, with relevant R&D resources consolidated under him. In June, the Seed Robotics team was also merged under Zhou Chang, further concentrating responsibilities for multimodal and embodied AI.

A noteworthy detail is the training decision for Seedance 2.0. During a team dinner in late 2025, Zeng Yan, the video model lead and a 2021 campus recruit, proposed training a 200B-parameter model directly.

There was internal disagreement, with some arguing for a more cautious 100B-parameter approach due to limited training resources. After discussions, Wu Yonghui and Zhou Chang supported her judgment. This seemingly aggressive choice later proved critical to Seedance 2.0's generational leap.

03 Seed Without Distillation

At a recent Seed all-hands meeting, Zhang Yiming spoke at length, delivering a core message: Seed would not pursue catch-up through distillation, neither from closed-source nor open-source models, and could accept temporary setbacks.

According to subsequent media reports, this decision underwent three rounds of internal debate.

The first occurred after DeepSeek-R1's release in January 2025, when researchers proposed using data generated by leading models to supplement reasoning capabilities. The second came in late 2025, after NVIDIA's Blackwell GPU deployment widened the Sino-US computing power gap, with voices suggesting “distillation as a last resort” resurfacing.

The third peaked in July 2026 after Kimi K3's release, when Moonshot AI advanced open-source models into the global first tier with a much smaller team. Internal wavering at Seed reached its height, even leading to a compromise proposal of “at least distilling open-source models.”

Zhang Yiming rejected all three proposals. From the timeline, ByteDance's wariness of distillation was not short-lived. In April 2023, the team introduced GPT API call specification checks, explicitly prohibiting GPT-generated data from being added to training sets and conducting internal audits.

Zhang Yiming's exact words at the meeting were, “To achieve long-term goals, we should be willing to sacrifice some short-term interests.” The underlying logic is straightforward. Distillation is an effective shortcut, allowing models to close benchmark gaps within months. However, it fails to explain how these capabilities arise or provide direction when current state-of-the-art models hit ceilings.

A model can enter the frontier through distillation, but a lab cannot become a frontier lab through distillation alone.

Accepting short-term setbacks meant delivering sufficiently impactful results in other areas to sustain confidence. Seedance 2.0 filled this role.

Video generation was the track where ByteDance and overseas competitors started closest. After Kuaishou Kling adopted a DiT native video architecture, it led for nearly a year in 2024. Internally, ByteDance's early PixelDance followed a 2D UNet extended to 3D route, which later proved limited. Zeng Yan brought AI Lab's video team into Seed, switching the technical route from UNet to DiT while consolidating teams and directions.

Multiple practitioners attribute Seedance 2.0's success to “a data victory.” ByteDance allocated a thousand-person-scale data evaluation team for this model, while many video startup evaluation teams numbered only in the dozens.

Algorithm teams had dedicated personnel liaising with data teams, formulating clear requirements like data product managers, with algorithmists also participating in data construction and cleaning. However, training data rarely used Douyin content; instead, they heavily procured film-grade materials, decomposed them into scripts and storyboards using language models, and conducted targeted training for motion, indoor spaces, gaming footage, and other scenarios.

Commercial returns came quickly. During the 2026 Spring Festival, Seedance 2.0 was fully integrated into Doubao, CapCut, and Jimeng, with consumer-side queues stretching to ten hours at peak. Jimeng generated its first revenue through member priority queues, earning approximately 140 million yuan in March, rising to 210-220 million yuan in April.

Volcano Engine promptly opened enterprise APIs, priced at roughly 1 yuan per second for 720P output. Seedance was indeed more expensive than domestic peers but delivered on quality, with leading short-drama companies like China Literature and Jiuzhou Culture pre-purchasing 50 million yuan in credits.

Tan Dai confirmed in interviews that over half of Volcano Engine's 2026 MaaS revenue came from Seedance. The video model's high pricing, high margins, and relatively uncontested market allowed model SOTA to directly translate into revenue.

More importantly, it forms a closed loop with ByteDance's existing businesses: models are provided to content companies, which use Seedance to create AI short dramas. Hongguo and Douyin handle content distribution and support producers in advertising. A flywheel, spanning from model capabilities to content production and then to advertising revenue, has started spinning in the video modality.

At the Volcano Engine FORCE Conference in June, Doubao 2.1 Pro was released. This represents Seed's progress in catching up in the Coding domain, with research and development resources consolidated under Guo Daya's leadership after his joining. Terminal Bench 2.1 scored 71.0, closely trailing Claude Opus 4.7's 71.7; SciCode scored 59.8, surpassing GPT-5.5 and Opus 4.7; NL2Repo's repository-level code generation scored 47.0, significantly ahead of GPT-5.5.

Tan Dai remarked that Seed has now "joined the table," noting that few domestic companies have truly done so.

In August, SeedRealtime was released. This native audio-video full-duplex model replaces the traditional "speech recognition + language model + speech synthesis" cascade approach with a unified architecture, reducing end-to-end latency to a level that allows natural interruptions and real-time dialogue. It was fully launched on the Doubao App the same day. From video generation to full-duplex interaction, the accumulation of capabilities in multimodal has begun to enter the product harvest phase.

Liang Rubo recently admitted at an all-hands meeting that large language models have fallen behind overseas SOTA, but emphasized, "Don't distort our actions." He advocated accepting a period of lag and persisting in self-research and long-term optimization. He summarized the company's strategy as highly prioritized, with a thick trunk and optimized for the long term—AI, information platforms, and transaction services being the three core businesses, with AI being the thickest of them.

Looking back at Seed's three years, the path has not been linear. Starting late in 2023, covering up foundational gaps with application-layer successes in 2024, learning lessons from DeepSeek in 2025 to catch up, achieving leadership in video and multimodal in 2026 while still chasing in language models.

The organizational structure has undergone several rounds of adjustments, with core personnel coming and going. The technical roadmap has taken detours, but there have also been moments when the persistence of young researchers paid off.

As they internally agree: "AI is a marathon, and we're still just at the first 500 meters." This statement carries additional significance for ByteDance.

This company has always quickly succeeded in certain tracks using certain methodologies, but foundational research is inherently uncertain. You need to tolerate failure, long feedback cycles, investments that may not yield returns for years, and even periods of clear lag.

The name Seed itself carries a certain metaphor. It is more like a seed still taking root in the soil than a fully formed product. The mission of a seed is not to bloom in the first quarter. It needs to root deeply first, learning to grow in the soil before its season arrives.

This article is original to Xinmou. For reprint authorization or business cooperation, please contact.

— END —

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.