MiniMax H3 Closes In, Leaving Seedance with Limited Safe Space

08/18 2026 484

At a Seed all-hands meeting in late July, Zhang Yiming rarely discussed model distillation.

He made it clear to the team: Do not rely on the outputs of competitors' models to shorten the catch-up time, even if it means incurring some short-term losses. He did not urge the team to catch up quickly on the next leaderboard; instead, ByteDance is willing to allocate time for longer-term model capabilities.

This statement clarified ByteDance's trade-off in model competition—short-term rankings can be conceded, but foundational capabilities cannot be supplemented by copying competitors. Seedance exactly provides a multimodal sample that has already yielded results.

Seedance 2.0, released in February this year, already integrates text, images, videos, and audio into a single audiovisual generation model. Seedance 2.5, announced in June, further extends single video duration to 30 seconds and enhances complex reference and editing capabilities. 2.0 proved that this approach could lead, while 2.5 began to address how long such a lead could be sustained.

A little over a month later, MiniMax followed a similar path. On July 31, MiniMax released the video model H3. Text, images, videos, and audio can simultaneously enter the model, allowing users to reference a video's camera movements, characters from another image, and audio from a third source, then complete generation or editing through natural language. MiniMax also announced that it would open the model weights within days, hoping developers would deploy and customize it themselves.

The two companies are making video models increasingly similar, no longer satisfied with "generating a video from a single sentence." But as the models become more alike, the methods they use to maintain their leads diverge. MiniMax chose to push H3 toward an open-weight approach, hoping developers, chip companies, and third-party tools would expand its deployment; ByteDance keeps Seedance within its own product ecosystem, using Jimeng, Doubao, and a growing number of internal services to support the model.

One borrows speed externally. The other actively buys time at the most fundamental level.

In late 2025, the Seedance model team had a dinner gathering where Seedance head Zeng Yan proposed to her superiors that she still wanted to attempt a larger model, with parameters reaching at least 200B. There was internal opposition at the time. Some argued that directly scaling the model to 200B or even 300B was too risky and that starting with a 100B-scale model would be more prudent given limited training resources.

A senior executive happened to attend the dinner. After understanding Zeng Yan's idea, he offered to help coordinate resources. At the subsequent formal review meeting, Seed head Wu Yonghui and visual multimodal generation head Zhou Chang ultimately supported the larger model proposal.

Seedance had taken detours before. PixelDance initially extended from an image diffusion approach, while Kuaishou's Kling bet earlier on native video DiT, once holding a lead. By the end of 2024, ByteDance gradually consolidated its video generation resources under Seed and readjusted its technical approach.

The turning point came with Seedance 2.0. Zeng Yan and her team insisted on further expanding pre-training scale, making the data and model structure more substantial. Multiple sources close to the team said that Seedance 2.0's data evaluation team numbered in the thousands; the core algorithm team itself had only a dozen members, but behind each algorithm engineer stood a larger data team handling standard-setting, sourcing, cleaning, and evaluation. Seedance even relied little on Douyin videos as primary training data, instead procuring large amounts of high-quality film and television material.

Many video industry practitioners later summarized Seedance 2.0's improvement in one sentence: This was a victory of data. In February, the results were in. Seedance 2.0 was no longer just a text-to-video model. It could simultaneously read text, images, audio, and existing videos, combining subject reference, camera movements, audio, and editing into a single task. ByteDance's April technical report revealed that the model supported 4–15 second native audiovisual generation and reference from up to 9 images, 3 videos, and 3 audio clips.

This model definition is also rapidly becoming an industry consensus. In H3's official technical introduction, MiniMax specifically reviewed the fragmented state of previous video models: text-to-video, image-to-video, motion transfer, audio reference, and video editing were often split into different tasks. H3 aims to reunify these capabilities within a single context, using natural language to describe relationships between materials and between materials and the final video.

MiniMax even voluntarily abandoned some aspects of its previous architectural design, believing that "task generalization" was more important than continue (continuing) a particular architectural advantage. H3 ultimately supports joint input of text, images, videos, and audio, also integrating video generation and editing into a single model.

Both sides ultimately arrived at similar positions, but for Seedance, this does not mean its previous technical investments were meaningless. To date, Seedance 2.0 remains in the first tier of mainstream third-party video evaluations.

In the past, a model excelling at native sound, multi-angle shots, and subject consistency was enough to create a clear gap with competitors; now, these capabilities are beginning to appear simultaneously in flagship models from different companies. H3's emergence makes this even more concrete—native sound, multimodal reference, motion control, and video editing are shifting from differentiated capabilities of a few models to foundational capabilities contested by all flagship video models.

Model capabilities remain Seedance's first ticket to users, but a single model lead is increasingly insufficient to explain long-term leadership alone.

When H3 was released, MiniMax devoted considerable space in its official blog to explaining why it chose to open weights. The first reason was the developer community, the second was domestic chip adaptation, and the third was enabling users to customize models according to their needs. MiniMax also stated that H3 had considered compatibility with multiple domestic chips from the early design stage. On July 31, Reuters confirmed this open-weight plan.

This is both a technical and a distribution choice. MiniMax has Hailuo, MiniMax Hub, and API services, already serving a large number of individual and enterprise users. However, in the video content industry, it lacks platforms like Douyin and TikTok or ByteDance's massive ecosystem for model-driven creation, distribution, and advertising. MiniMax's official website currently discloses that its models and AI products have "cumulatively served" over 300 million individual users, more than 1 million enterprise clients, and developers.

If the open-weight plan proceeds as planned, developers and hardware vendors will become new distribution nodes for H3. Each chip vendor's adaptation, each inference platform's addition of H3, and each video workflow's inclusion of it as a default model will provide MiniMax with reach paths not entirely dependent on its own sales team.

The open route does not mean MiniMax abandons official services. H3 still offers full capabilities via API; MiniMax claims its 2K version's per-unit-time price is less than one-third of mainstream models'. It is expanding model deployability while retaining official inference and commercial services.

ByteDance faces a different situation. During this year's Spring Festival, after Seedance 2.0 was widely integrated into products like Doubao, Jianying, and Jimeng, it briefly experienced long queues due to overwhelming demand. By late July, Seedance 2.5 was incorporated into Gauth. Business Insider reported that this ByteDance-owned education app has 86 million monthly active users and began using Seedance 2.5 on August 4 to convert history lessons into AI videos with narration.

An interesting shift also occurred within Seed. Known for rapid trial-and-error, data feedback, and high-frequency iteration, the company is now proactively slowing down evaluations in foundational model research. In January 2025, ByteDance established Seed Edge. Officially defined as a research initiative supporting longer-term, high-uncertainty research topics, it provides independent computing power for selected directions and adopts longer-cycle evaluation methods. Zhang Yiming also re-engaged in AI R&D discussions. According to The Paper, citing insiders, since the second half of 2024, Zhang has regularly participated in Seed's core technical team reviews and discussions.

By this summer, he brought a long-standing internal Seed issue to the forefront. On August 6, Reuters, citing The Paper, reported that Zhang urged the team not to rely on distilling competitors' models to improve capabilities, even if it meant sacrificing some short-term gains.

A clarification is needed here to avoid confusion: ByteDance does not restrict distillation itself. Seedance 1.0 used multi-stage distillation to improve inference speed, and Seed's current hiring directions still include model distillation. What Zhang opposes is using competitors' cutting-edge models as "teachers" to gain short-term rankings by leveraging their established capabilities.

In this sense, ByteDance has not truly chosen to slow down. It is willing to allocate longer cycles to foundational research, but once models are developed, product-side operations continue at another pace. MiniMax borrows speed externally—letting more developers find use cases for H3; ByteDance buys time internally—allowing model research to take longer, then relying on its existing products to make up the difference.

This may be closer to the true divergence between the two companies than "open-source versus closed-source."

Seedance 2.0 first proved that in the current video model market, sufficiently pronounced capability leadership can quickly translate into revenue.

By June, Seedance had contributed more than half of Volcano MaaS's revenue. In the video model market, where large language models have already entered price competition, Seedance 2.0's release was followed by months of Volcano securing large model orders from advertising, short drama, comic drama, and other content production clients.

But once models entered production environments, questions shifted. Clients began considering not just which model was best but which model was worth committing several months or even a year's production capacity to.

Zhou Zhipeng is one such example.

One afternoon in March, Zhou Zhipeng, co-founder of ShuiMu Intelligence, was hiking on Beijing's Baiwang Mountain when he received six consecutive sales calls from Volcano Engine, all promoting Seedance 2.0. According to subsequent reports, he decided to sign the contract the next day. What truly gave him pause was not the 10 million yuan itself but what it meant: a significant portion of his video production capacity would be tied to Seedance for the next year.

Thus, clients were not just purchasing Tokens but also betting on how long Seedance could maintain its lead. This was already reflected in the revenue structure. In June, when interviewed by media, Tan Dai, head of Volcano Engine, said Seedance contributed more than half of Volcano MaaS's annual revenue. When interviewed again in June, Tan confirmed that Seedance still contributed over half of Volcano's MaaS revenue at that time.

For the first time, model leadership had become a large enough business. The question then shifted: How long could this lead last?

After H3's emergence, the risks in these one-year contracts became clearer: Clients buy a year's worth of capacity, yet the gap between models could shift within months. In H3's July 31 official blog, MiniMax included "participating in the entire content production process" as a model direction. It aimed to address not just generating footage but also handling revisions and complex tasks through natural language.

As models become more capable, new questions arise. When Seedance 2.0 could only handle a dozen reference inputs at a time, users first cared whether the model could understand the materials; when next-gen models can manage more characters, scenes, sounds, and timelines, users' questions will gradually shift to: Who organizes these materials? How are characters bound? How is a mistake at second 17 separately reworked? How do multiple 30-second videos maintain consistent characters and narrative rhythm?

The more complex tasks a model can handle, the more work exists beyond the model. After Seedance 2.5 increased single-generation duration to 30 seconds and reference materials to 50, material management, character binding, partial rework, and segment connection (transitions) became new engineering challenges. Workflow products like LibTV began addressing these links. The closer underlying models become, the more clients' reasons to stay must extend beyond the model itself.

However, ByteDance still has two assets MiniMax cannot easily replicate in the short term. One is the data, pre-training, and Infra systems behind Seedance, forming a relatively heavy model production ecosystem. Seed's current hiring directions still cover foundational models, multimodality, Long-Horizon Tasks, and training infrastructure, indicating ByteDance continues to strengthen this layer.

The other layer sits above the model: ByteDance owns deployment scenarios. Jimeng serves creators, Volcano Engine connects advertising, e-commerce, and content production clients, and Gauth brings Seedance into education, allowing the same model to enter diverse real-world tasks more quickly.

ByteDance's product advantages primarily lie in scenarios and distribution, rather than simply being equated to 'training models with Douyin data.' Previous investigations have provided contrary details: The core training of Seedance 2.0 heavily relies on the procurement of high-quality film and television materials, with almost no direct use of Douyin data.

Jimeng, Volcano Engine, and Gauth have respectively brought Seedance into creative fields, enterprise services, and education. After the model is developed, ByteDance rarely lacks a first-use scenario. Where creators frequently need to rework, what capabilities enterprise clients are willing to pay for, and which capabilities are worth prioritizing for iteration can all be validated earlier on the product side.

MiniMax's approach is to enable more external developers and hardware manufacturers to deploy, adapt, and develop around H3. After open-source weight implementation, enterprises can deploy it themselves, chip manufacturers can adapt around it, and workflow teams can continue to integrate H3 into their products.

On one hand, ByteDance keeps the closed loop within the company; on the other hand, it entrusts the closed loop to the ecosystem. Which approach is more effective remains unanswered. For Seedance, the real danger lies in when clients are willing to sign a one-year contract, they are essentially betting on how long Seedance can maintain its lead. After the emergence of H3, this bet becomes even harder to make.

A single first-place ranking on a leaderboard can bring orders, but sustained leadership requires a different set of capabilities.

Zhang Yiming's discussion of 'long-termism' at the Seed all-hands meeting ultimately boils down to a resource allocation choice. Not using outputs from competitor models to fill capability gaps means Seed needs to prepare more data, conduct more experiments, and bear the costs of directional failures and short-term setbacks. Seed Edge's adoption of longer evaluation cycles also means the company must allow a portion of its computing power and researchers to operate without direct business results for extended periods.

Seedance 2.0 has already shown ByteDance a return on investment. By the end of 2025, the additional training resources Zeng Yan advocated for during a team dinner ultimately resulted in a video model that reached the global first tier; a few months later, Volcano Engine sales began using it to approach enterprise clients.

A model's initial lead can come from a single correct bet, but sustained leadership requires a mechanism capable of repeatedly producing the next generation of advantages. H3 has already begun to cover a range of generation, reference, and editing capabilities similar to Seedance. Next time, there will be Kling, Wan, and new models. For ByteDance, long-termism must ultimately prove not whether it can endure a few months of setbacks, but whether the additional time, data, and computing power invested can ultimately yield a next step that competitors find harder to replicate.

MiniMax has taken a different path. It chooses to push H3 toward an open ecosystem, hoping more developers and hardware manufacturers will participate in the model's deployment, adaptation, and application development. One company places speed externally, while the other retains time internally.

However, ByteDance hasn't truly slowed down. Once a model is ready, Jimeng, Doubao, Volcano Engine, and Gauth will still demand its rapid deployment to users. It is simply beginning to experiment with splitting a company into two speeds: allowing slowness in foundational research while maintaining speed in product development above the model level.

After H3 catches up, what Seedance truly needs to prove next is whether the additional time it invests can yield a next step that competitors find even harder to match.

*The featured image and illustrations in the text are sourced from the internet.

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.