08/20 2026
569
By 2026, Silicon Valley is reshaping office work—shifting from “people operating software” to “people assigning tasks while AI agents operate software.”
Microsoft is expanding its AI agent ecosystem around Microsoft 365, integrating Copilot Cowork, Office’s agent capabilities, and Agent 365. Google’s Workspace Studio transforms enterprise standard operating procedures (SOPs) into reusable agent workflows. OpenAI launched ChatGPT Work in July, enabling sustained cross-file and application-connected workflows lasting hours.
While product designs vary, the strategic direction converges: agents aren’t rebuilding office software but competing to dominate the task execution layer above existing tools. Domestically, the landscape is unique—the first wave of truly scaled products includes independent agents like WorkBuddy and Baidu’s WorkBuddy, aligning more closely with Claude Code’s approach.

On July 23, Tencent introduced WorkBuddy Bench, integrating seven leading models into two agent frameworks—CodeBuddy Code and Claude Code—to complete 260 tasks across coding, web operations, office automation, and security. After model and framework optimizations, office tasks demonstrated the highest stability. In CodeBuddy Code, seven models scored between 77.47–82.37; when switching frameworks, five models showed changes of <2 points, with a median absolute change of just 0.86.
This suggests that for standardized office tasks, differences in models and frameworks haven’t yet created significant performance gaps, unlike coding or security tasks. “Getting things done” is rapidly becoming a universal baseline capability.
Agents accelerate task completion, but human workloads don’t shrink proportionally. A researcher might delegate three years of financial data from five companies to an agent to identify anomalies, trace causes, and generate a one-page PPT—compressing the most time-consuming steps. However, if the researcher must then reopen financial reports to verify numbers, dates, and sources, some of the AI-saved time becomes verification time.
This poses a tougher challenge for office agents: After “getting things done,” how to make people trust the output directly.
Baidu is catching up at this critical juncture, but its assets differ from Tencent’s. Q2 earnings revealed Baidu’s AI applications generated RMB 2.5 billion in quarterly revenue, up 3% YoY. In June, the enterprise version of WorkBuddy launched; simultaneously, Baidu Wenku and Netdisk AI DAU penetration rose 27.4% YoY. These revenues and users don’t all come from office agents, but they mean Baidu isn’t starting from scratch—it has search, Wenku, Netdisk, and other information/file systems with years of operation.

As “getting things done” nears baseline capability, office agent competition shifts to two ends: Why users entrust them with the first real task, and why they trust the output enough not to recheck it.
Tencent and Baidu’s legacy assets align with these two questions.

Agents require more guidance than chatbots.
When ChatGPT debuted, users needed no product education—a chatbox and a question like “What’s the weather in Beijing tomorrow?” sufficed. Agents are far more complex: they read local files, modify Excel, operate browsers, invoke skills, execute multi-hour tasks, run on schedules, and connect external software. Greater capability raises a concrete user question: What exactly should I delegate now?
For agents, the most effective product demo is often something others have already done. On July 1, three Microsoft researchers published a study on Claude Code and GitHub Copilot CLI adoption among Microsoft engineers, analyzing real usage data from early 2026. Separating “first use” from “sustained use,” they found different drivers: continued use correlated more with engineers’ programming activities, while initial trials were significantly influenced by colleagues.
The paper summarized it bluntly: “First use spread primarily through social networks.” It recommended enterprises incorporate “visible peer use”—letting employees see colleagues actually using agents—into their promotion strategies.
The study focused on Microsoft engineers and coding agents, so it doesn’t directly imply all office users will follow the same path. But it explains a unique propagation mechanism for agent products: functional descriptions are less effective than task demonstrations.
Tencent has an ideal network for this propagation. WorkBuddy lets users assign tasks remotely to computers via WeChat, Enterprise WeChat, QQ, DingTalk, Feishu, and other communication platforms—a network where agent usage can easily become visible among colleagues, friends, and group chats.
Traditional internet metrics emphasize Customer Acquisition Cost (CAC), but the agent era may require a new metric: the cost of acquiring the first real task. Ads can buy installations; free tokens can drive registrations. But neither guarantees users know what to do with the product the next day—or that the results will be satisfactory.
This is why WorkBuddy's current >11 million monthly active PC users warrant observation but can’t be simply explained as “Tencent has more traffic, so Tencent wins.” Product capabilities, subsidies, launch timing, and marketing all influence numbers. No evidence yet proves a singular causal link between WeChat’s social graph and WorkBuddy’s growth. A more precise statement: Tencent possesses an asset ideally suited for agent cold-start challenges.
Baidu’s WorkBuddy growth shows the market is far from settled. Baidu officially launched DuMate in March, positioning it as a desktop-grade office agent from day one, with rapid iteration since. By August 7, the personal version had reached v1.0.67, adding web publishing, embedded browsers, remote tasks, and automation templates. On August 12, Baidu disclosed that daily queries on WorkBuddy had grown 60x since launch, doubling again in the past month.

But if competition remains stuck on “who can help users discover agent capabilities faster,” Baidu faces a disadvantage. Tencent can leverage human relationships to propagate tasks through colleagues, friends, and group chats. Search brings demand; Wenku and Netdisk bring content; Maps bring locations; Cloud connects enterprises. But Baidu can’t easily replicate WeChat’s social graph.
Thus, Baidu needs a different approach—not just proving its agent can “get things done,” but that users can trust the output enough to skip rechecking. This is where search, professional content, and knowledge systems could regain value.

The more agents do, the more critical human verification costs become—a key difference from chatbots.
Asking a model to write copy with one error requires a single correction. Asking an agent to conduct industry research means it searches, filters materials, extracts data, calculates, forms judgments, and generates a file. Longer chains and complete task execution amplify early errors.
Integrating search doesn’t fully solve this. Multiple ACL 2026 studies focused on hallucination detection in RAG scenarios. One method aligns generated content with retrieved evidence; another, building a long-context RAG benchmark, found current hallucination detectors still fall short of optimal performance. In other words, retrieving materials doesn’t guarantee every generated sentence is material-supported.

Thus, “having search” and “being trustworthy” remain separated by a layer: evidence.
Model unfaithfulness to evidence is one issue; worse, evidence selection is growing harder. GEO (Generative Engine Optimization) complicates this. Previously, SEO optimized for search rankings. With GEO, content creators study how to make pages more likely to be selected, cited, and included in final answers by ChatGPT, Gemini, Perplexity, and other generative search engines.
On July 15, a preprint review summarizing 45 GEO studies from 2023–2026 tightened conclusions. The authors argued that while adjusting content can alter citation/adoption probability once a page enters retrieval context, no method yet proves long-term, cross-platform control over organic exposure. Different engines also show low source overlap and fluctuating results for repeated queries on the same issue.
In other words, future office agents will face an internet where information isn’t just increasing—it’s actively optimizing to be seen by agents.
This revives the importance of a “old” internet capability: determining which websites are trustworthy, which is the original source, which content is republished, which figures are outdated, and which source to trust when two conflict. Search engines have spent two decades solving parts of this problem—Baidu’s unique strength in this competition.
When launching WorkBuddy in March, Baidu built search capabilities directly into its skills. On April 27, Baidu Qianfan packaged search, Baidu Baike, deep research, and smart PPT generation into standard skills for agent invocation. Baidu Search handles real-time web and news retrieval; Baidu Baike provides structured encyclopedic information. Baidu Wenku adds professional documents, which, combined with Netdisk, let Baidu cover public webpages, professional materials, and users’ private files simultaneously.
Of course, Baidu isn’t building this evidence chain from scratch. Its search division already incorporates multi-source comparison, source identity verification, “screen-then-use” workflows, and multi-source cross-validation into information filtering. Qianfan’s deep research agent can directly invoke search components to identify authoritative sources; Baike’s multi-round review and structured knowledge system provide another layer of curated knowledge for agents.

But two decades of search doesn’t automatically make today’s agents trustworthy. Baidu Search operates in an internet shaped by GEO, marketing content, webpage updates, and low-quality information; Baike and Wenku have their data boundaries. More critically, the ACL studies prove models may still deviate from materials during final generation even when correct materials are found.
Thus, Baidu’s first challenge isn’t informational but proving the relationship between final statements and supporting information.
This requires product innovation. For example, in an industry report generated by WorkBuddy, a market size figure shouldn’t just cite “Source: X media” but link directly to the specific location in the original report. A listed company’s gross margin should trace back to the annual report, not a financial website’s republished version. A regulatory policy should connect to the government document, showing its release date and current validity.
When two sources disagree, the agent shouldn’t silently choose one for the PPT but expose the discrepancy. A truly useful evidence chain doesn’t just attach a dozen links at the report’s end—it lets every key judgment trace back to specific supporting evidence.
The second challenge stems from Baidu’s own assets: search, Baike, Wenku, and Netdisk represent different data types. In an agent report, public webpages, Baike entries, third-party research reports, and users’ internal company files have varying credibility. If the model blends them into indistinguishable natural language, users won’t perceive Baidu’s decades of data accumulation.
This requires WorkBuddy to productize source types—distinguishing internal company files from public internet content, corporate announcements from media reports, single-source data from multi-source confirmed facts—all visible in task results. Assets only become product advantages when users see them.
Truly usable “trustworthiness” will not involve re-verifying all tasks uniformly but may be stratified like security permissions: ordinary creation prioritizes speed, industry research adds citations, and financial, legal, and critical decision-making tasks enter a stricter evidence-based mode. In the search era, Baidu addressed the question: “Where is the information?” In the agent era, it has the opportunity to continue answering a more pressing question: “Why does the machine trust this information, and why can people trust the machine?”

It is still too early to determine the winner in the office agent market. In July, the monthly active users of AI office products on PCs in China just exceeded 30 million. WorkBuddy has over 11 million, Baidu Dazi has around 6.7 million, and products like TraeWork are rapidly iterating behind. Compared to the user scale of traditional office software, this more resembles the first batch of users beginning to choose their tools rather than a market landscape that has already taken shape.
However, the first phase has already begun to reveal the strengths of different companies. Tencent’s easiest advantage to capitalize on is relationships, while Baidu’s strength stems from products like search, encyclopedia, document libraries, and cloud storage—none of which were designed for the agent era.
Over the past few years, these have often been seen as legacy assets from the previous internet era, but with agents starting to perform real work, the questions have changed. Generating text with models has become cheap enough, and retrieving information has become increasingly easy. What hasn’t disappeared in sync is the final step: someone must still judge whether the results are correct.
Baidu thus has the opportunity to redefine an old capability. This time, however, it cannot stop at “I can find more things.” It needs to transform search rankings into source selection, turn encyclopedias and document libraries into different levels of evidence, distinguish private contexts from cloud storage from internet materials, and then explicitly connect all evidence back to every critical judgment made by the agent.
Should this capability eventually evolve into a service that professional offices are genuinely prepared to invest in, then the most significant legacy Baidu inherits from the search era may transcend being merely a gateway for traffic. It could also encompass the power—concealed within the search box for more than two decades—to discern the authentic origin of an answer.
*The featured image and accompanying illustrations in this text are sourced from the internet.