09/24 2026
564
Two days ago, reports emerged that Doubao's largest-ever General Session team was downsizing, with its size expected to halve. At the same time, the Product Posttrain-Chat team responsible for post-training C-end dialogue on the model side is also shrinking, with some employees shifting to Horizon RL and Product Posttrain-Work—the latter being one of the new first-level departments established by Seed Foundation Model in August, responsible for agentic capabilities in office and B-end scenarios.
Liu Xing provided a different set of figures: The General Session team has fewer than 50 members, and this restructuring affects 11 individuals, with 3 departures, not a "50% cut." However, his explanation confirmed that some functions previously handled by Session have been transferred to Transactions and Doubao at Work, with personnel moving accordingly.
In the past, the General Session team was responsible for the most fundamental aspects of Doubao's product experience: what users ask, when to search, the length of answers, how to follow up, and whether a conversation round is successful. Now, some of this work is shifting to specific scenarios. Conversations in Transactions will continue to be optimized by the Transactions team; conversations in work scenarios will fall under Doubao at Work.
These adjustments occurred after Doubao had already amassed hundreds of millions of users. According to QuestMobile data, Doubao's MAU reached 382 million in June, ranking first among AI-native apps in China. In the past, Doubao even adjusted its response styles based on user retention; now, some functions previously concentrated in general dialogue are being integrated into more specific transactional and work scenarios.
On August 25, Doubao at Work was launched. The official description is simple: Break down goals into tasks, utilize tools, and push complex processes to results. The headline is even more straightforward: "Let AI go from generating content to completing work." With 382 million MAU, ByteDance no longer desperately needs another chat entry point. It's starting to look for revenue beyond the chatbox.
For ByteDance's most successful content platform in the past, more usage often meant more monetization opportunities. Every additional user visit and minute spent increases signals for the recommendation system, providing more attention to monetize through ads, livestreaming, and transactions. The value function of an office agent is different. A good agent might even aim to reduce user participation.
Thus, a company that grew by competing for user time is now starting a business that gives users their time back. For Doubao at Work, the success of a task might even hinge on whether users can leave earlier—the less time people spend in the product, the more valuable it might become.
On June 24, Doubao introduced its first professional version. The standard package costs 68 RMB for a continuous monthly subscription, the enhanced version is 200 RMB, and the premium version is 500 RMB. One of the core capabilities for paying users is "office tasks": AI can operate local computers and browsers, process Office files, invoke skills, and return task results. The professional version uses Doubao 2.1 Pro, while free users can also experience office tasks powered by 2.1 Turbo within certain limits.
By then, Doubao already had a user base that most AI applications would envy. QuestMobile statistics show that Doubao's MAU reached 382 million in June, with the App gaining 13.78 million new active users compared to May. At least from that month's results, the professional version's launch did not interrupt user growth.
On the day of the professional version's launch, 36Kr tested four office tasks, three of which were not fully completed. For example, after cleaning the C drive, available space changed from 3GB to 2.82GB, and two spreadsheet tasks got stuck during file operations. The test primarily ran on 2.1 Turbo, which does not represent the upper limit of 2.1 Pro, but it exposed a problem in advance: as tasks moved from answering questions to file operations, software interactions, and multi-step execution, failure points increased.
By August, Doubao had transformed this capability into "Doubao at Work."
The product capabilities continued to iterate, but the problems did not disappear. The 382 million users gave Doubao at Work a very low threshold for first-time use. Doubao does not need to spend years telling users who it is, like a new agent startup would. Its existing user base, enterprise scenarios, and account and context capabilities allow many people to quickly try office tasks once.
However, this is still different from forming trust. When a user first asks AI to write an industry report, they likely just want to see how smart it is; the second week, when they delegate the same weekly report again, the capability starts to enter their workflow; three months later, when they no longer remember how the report was manually completed in the past, the agent has truly taken over a task.
So, the funnel for Doubao at Work has changed: Usage → Payment → Repeated Tasks → Long-Term Trust. With 382 million MAU, the first step is much easier than for most agents; 68 RMB sets a price for the second step. What remains unanswered is how many trials will turn into payments and how many payments will turn into repeated trust.
McKinsey's global survey of 1,719 respondents this year provides a good reference. While 80% of respondents believe AI has improved personal productivity, only 37% think AI has positively contributed to their company's EBIT; just 6% meet McKinsey's criteria for "AI high-performing enterprises," meaning at least a 5% EBIT contribution and significant perceived impact.
The gap between these numbers lies in "completing work." A person asking AI ten questions a day can increase DAU, message volume, and tokens together. But whether those ten questions replace ten tasks is another matter. Someone might have revised PPT titles twenty times manually before; now, they converse with AI twenty rounds. Message counts rise, but employees' off-work times do not change.
In the consumer internet era, this kind of growth rarely needed to be distinguished separately; for agents, it must be calculated separately. Chat products fear users leaving after two messages; for office agents, people staying glued to the screen means the task hasn't truly been handed over.
Xinlichang also wants to know how far Doubao at Work can go when "completing work" truly becomes the acceptance criterion.
We prepared an August operations package for Doubao at Work from a fictional company. It contained 153 orders with two date formats, including refunds, cancellations, cross-month and cross-quarter records, and intentionally duplicated exported orders. The same order number retained two versions with amounts of 15,000 RMB and 18,000 RMB, neither directly verifiable as genuine.
The file included business guidelines, with standard answers pre-calculated manually and not provided to Doubao. We only gave a requirement close to what a real manager would say: Help me organize August's operations and output an Excel file directly usable for Monday's operations meeting, a 5-page PPT, and a 500-word management summary. Verify or annotate important anomalies, and do not make assumptions for uncertain areas.
Under the standard package and "auto/high" settings, Doubao at Work took 11 minutes and 6 seconds to deliver all three files. It did not ask us to supplement business information, require manual error correction, or reinitiate tasks. The task card showed a quota consumption of 0.92%.
If the acceptance criterion is "whether files were generated," this task was highly successful. The Excel file contained four sheets, ranging from an operations overview and departmental breakdowns to an anomaly list and cleaned-up orders. August net sales were 227,340 RMB, up 70.2% from July's 133,540 RMB, matching the standard answer. It correctly removed five duplicate records, did not include canceled orders in sales, and did not arbitrarily choose a version for the conflicting order amount.
The problem emerged in the next step. The number of valid orders in Excel was 76 for August and 57 for July. On the second page of the PPT and in the management summary, these numbers became 83 and 60, respectively. Sales figures were correct, but another key metric for operational judgment was wrong.
A similar issue occurred with the order pending verification. Both Excel and the fourth page of the PPT correctly stated: If confirmed, this order might increase August net sales by 15,000–18,000 RMB. However, in the fifth-page action table and final response, the impact was written as "±3,000 RMB." The difference between "two candidate amounts" and "how much revenue increases after confirming the entire order" are clearly two different things.
There was also a 7,200 RMB order placed in August and refunded in September. Following the guidelines in the file, Doubao correctly did not deduct it from August's operating results but further wrote in the PPT that "September net sales will decrease by 7,200 RMB." According to the given business rules, refunds should be allocated back to the original order month. A September refund is a disclosure timing issue; counting it in September sales follows a different logic.
As a stress test, this does not represent Doubao at Work's average completion rate but does verify one thing: When tasks simultaneously involve data cleaning, business guidelines, and cross-file delivery, manual verification costs may resurface after production.
In this test, the real areas requiring human intervention in Doubao at Work emerged after delivery. Generating three files in 11 minutes does not mean completing a task that can be confidently handed over. While the model reduced production costs, verification costs rose.
Doubao rarely needed to design products around "letting users leave earlier" before. The "pleasing" adjustment before this year's Spring Festival provides a counterexample.
According to LatePost, when Doubao previously launched new versions, it had a strict product constraint: Retention must not easily decline. The product team defined response strategies, and the post-training team adjusted reward models based on user feedback, gradually shaping Doubao's responses to be more user-friendly.
After one training version pushed user preference weights too high, the model became better at agreeing with users, affecting accuracy. The team knew there was an issue, but the challenge was making changes. Cooling down responses might make users less chatty, so managers ultimately decided to accept a short-term retention drop of less than 1%, allowing the team to spend over a month retraining.
This follows a familiar growth logic for content platforms: Give users what they like and improve matching accuracy; every additional day and minute spent increases future monetization opportunities. Doubao at Work now needs to make "less participation" a good product metric for the first time. Its goal is not to reduce user-submitted tasks but ideally to have users delegate more work to it while spending less time in the product themselves.
There's a second reversal here: As human time decreases, machine time may increase. After receiving a prompt, an agent must plan, search, read context, open software, invoke tools, and repeatedly verify. If it fails once, it tries again.
Microsoft Research analyzed the execution trajectories of eight cutting-edge models on SWE-bench Verified in April this year and found that agentic coding tasks could consume roughly 1,000 times more tokens than ordinary code reasoning and code chat; token consumption for the same task repeated could vary by up to 30 times. Higher token consumption did not reliably translate to higher accuracy.
While programming tasks are not directly equivalent to office work, they reveal an important point: A very short human instruction can hide a very long machine work time behind it. Even as models become cheaper, this problem does not automatically resolve.
Cheaper models do not necessarily mean cheaper tasks. As capabilities improve, people will delegate longer, more complex tasks to agents: from revising an email to processing ten files to running a complete monthly workflow. While the price per million tokens may drop, the computational volume consumed by a task could continue to rise.
Moreover, machine costs are not the only costs. When McKinsey dissected the economics of agent workflows this year, it found that in a bank customer service case, tokens accounted for only about 20–25% of an AI agent's variable operating costs; human supervision accounted for 70–75%. Professionals still handle final checks, exceptions, and risks—often the most expensive part of a task.
This explains the operations analysis test in Chapter Two. The model compressed much of the production process, but as long as a person still needs to recheck order counts, amount definitions, and anomalies, these human costs remain. Improving a model's accuracy from 90% to 95% has value, but if the final 5% of errors forces employees to recheck 100% of results, that 5% redefines the task's economics.
Thus, enterprises will not only count tokens. They will calculate the total cost of completing a task. After a first execution fails, should they keep trying with the most expensive model or return the task to humans? To raise success rates from 95% to 99%, should they run three more verification rounds? If a task that originally took humans 30 minutes now takes an agent one hour but only occupies humans for 3 minutes, is it more or less expensive?
The September controversy over the Session team provides a more concrete cross-section of these changes. According to Doubao's own response, the team was not "halved," but some functions previously belonging to General Session were indeed transferred to Transactions and Doubao at Work.
These changes are not limited to product teams. A month earlier, Seed's model organization had already made a clearer division. Seed Foundation Model established a new Product Posttrain-Work department responsible for B-end applications, integration and release of agentic models, and agentic capabilities in office scenarios. The original Application team was renamed Product Posttrain-Chat, continuing to handle C-end dialogue.
One is called Chat, the other Work—these names clearly define Doubao's current challenges. In the past, many capabilities of large model products converged into a single chatbox, with model performance and user willingness to keep chatting providing direct feedback. Once agents enter work scenarios, they must not only answer questions well but also complete tasks.
Doubao hasn't stopped engaging in chat; it's just that the chat box no longer needs to bear all of Doubao's commercial aspirations. When an AI application already has hundreds of millions of users, the next question shifts from 'how to bring more people back' to include another layer: after these users arrive, which behaviors can ultimately be transformed into work, transactions, and revenue.
When this happens, ByteDance itself is also bearing increasingly heavy AI costs. The Wall Street Journal reported in September that ByteDance's revenue in the first half of 2026 grew by more than 30% year-on-year, reaching approximately $120 billion; net profit during the same period was about $20 billion, with AI investments weighing down on profits.
In the past, ByteDance had a comfortable premise for growth: the longer users stayed, the greater the commercialization potential for advertising, live streaming, and transactions, with traffic itself being an asset. Doubao Work puts another segment of time, previously considered unimportant, onto the balance sheet. However, after people's working time is saved, how much time the machines use to make up for it begins to determine the cost of this business.
Over the past decade or so, ByteDance has created several of the most successful 'time products' in the history of the Chinese internet. For example, Toutiao and Douyin, whose underlying recommendation engines revolve around human attention.
Doubao Work faces a different situation. Once a task is assigned, it's best if people don't have to come back. At least, not come back so frequently. In this case, usage duration is no longer sufficient. Metrics like repeat task rate, first-time completion rate, human intervention rate, and the total cost of a single task will be closer to the business outcomes than 'how long users stay inside'.
The repeat task rate shows whether users have turned initial curiosity into a habit, the first-time completion rate indicates whether the Agent can handle the work, the human intervention rate reveals how much time people have reclaimed, and the total cost determines whether it's worth the machine's time to exchange. Until these calculations are clear, many advancements in office Agents can only demonstrate one thing: AI is getting better at working, but it hasn't yet proven that replacing human work is a sufficiently good business.
200 million daily active users remain one of ByteDance's most important user assets in the AI era. However, with Doubao Work, it only solves the question of 'where people come from' but not 'when people dare to leave'. Over the past decade, ByteDance has been best at calculating how long people are willing to stay. With Doubao Work, it now needs to calculate how much earlier people can leave.
*The featured image and illustrations in the text are sourced from the internet.