A Week with QianWen Office: The Highs and Lows

08/12 2026 362

A delightful discovery within the Alibaba ecosystem.

"A pleasant surprise."

During last week's public beta of QianWen Office, nearly everyone on our team gave it a try, and the feedback was surprisingly uniform: it's surprisingly good.

Overall, QianWen Office boasts a high degree of completion, primarily catering to 'professionals.' It eschews flashy features and doesn't demand prior skill creation or environment configuration. With templates and pre-packaged 'expert skills,' it's user-friendly with a low barrier to entry. Daily tasks such as cross-software data extraction, Excel processing, PPT generation, and database management can generally be accomplished seamlessly.

Many are familiar with the origins of this product. Led by Chen Yusen, Alibaba's youngest-ever business division CEO, QianWen Office merges three AI office products—QoderWork, Wukong, and MuleRun—which were previously quite similar in form.

In just 46 days, from assuming the role of DingTalk CEO to product launch, this pace seems somewhat atypical for a company of Alibaba's scale. Yet, it underscores the internal emphasis on AI office solutions.

The reasons are clear. Over the past two months, Tencent's WorkBuddy has launched standalone apps across three platforms, ByteDance's Feishu and Doubao product teams have recently merged, and all three major companies have shifted focus from internal competition to the AI office market.

However, our focus this time is on Chen Yusen's insights. In a recent interview, he stated:

"Currently, the functions of various Agents and the office tasks they can perform are highly similar, making it challenging to create significant differences in product experience alone. Long-term success hinges on enterprise resources, B-end customer channels, and strategic direction choices, rather than just interface or basic capabilities."

In simpler terms, this means prioritizing a user-friendly experience, stable and comprehensive task execution, building barriers through differentiated tools rather than relying solely on large models, possessing enterprise collaboration attributes, and truly undertaking high-value, high-frequency office tasks. This is also his, or Alibaba's, vision for AI office products.

Over the next few days, we selected multiple tasks commonly encountered in daily work and life to test QianWen Office. These scenarios covered structured information extraction, data cleaning and analysis, fact-checking, content generation, multimodal file processing, constrained planning, and more—areas where large models often struggle in practical applications.

We aimed to see how far QianWen Office had come in bridging the gap between Chen Yusen's description of a 'good product' and a 'great AI office product.'

01 Real-World Office Scenarios

As media professionals, when 99% of peers get their hands on Agent-level office software, the first scenario they almost always test is topic selection.

From Manus to various Work assistants and even early large model applications, the initial step is to sift through recent major companies, hot sectors, and industry leaders' new moves and statements, compile them into a fact list, suggest several topic directions, and if inspired, even draft an initial article.

Before using QianWen Office, they internally launched a web application called 'QianWen Night Shift,' created by QianWen Office. Essentially, it's a real-time updated information website that automatically fetches, translates, and organizes blog transcripts, overseas updates on key figures, and industry news into a page, with continuous updates.

This showcases QianWen Office's website generation and continuous operation capabilities. It doesn't just provide a static result but runs 24/7 in the cloud, constantly consuming computing power and refreshing content, with links that can be directly shared with colleagues.

If applied in the media industry, this would mean everyone could have their own 24/7 information assistant.

However, we didn't delve deeply into this scenario this time. Information aggregation tasks are difficult to truly assess a product's capabilities. Instead, we chose scenarios where 'failure could cause significant issues.'

For example, contract review.

We provided it with a rental contract containing hidden pitfalls, asking it to identify unfavorable terms from the tenant's perspective, focusing on deposits and penalties, and to explain in plain language without legal jargon.

This task seems simple but is actually high-risk. Many AIs might flag standard clauses as risks, fail to identify real traps, or even invent 'issues' not present in the contract.

QianWen Office performed admirably. It identified all pre-buried traps—non-refundable deposits, shifting all maintenance responsibilities to the tenant, unilateral termination rights for the landlord, and exorbitant late fees—without overinterpreting. Instead, it provided direct modification suggestions for negotiation with the agent.

If contract review tests 'comprehension,' then expense report processing tests 'accuracy.'

We uploaded six PDF invoices, including clear electronic invoices and blurry scans, along with a statement, asking it to sort by invoice date, merge PDFs, extract invoice numbers, amounts, and dates into Excel, and calculate the total.

OCR recognition accuracy exceeded expectations. Blurry scans were flagged with 'recognition unclear, verification recommended,' and date sorting worked flawlessly, correctly interpreting inconsistent formats like '2026-07-15' and 'July 15, 2026.' Most critically, the total amount calculation was precise.

E-commerce operational data review is another typical scenario.

We provided it with July's order details, asking it to exclude refunded orders, calculate Top 10 SKUs, identify abnormal orders, and generate a summarized table with charts.

Many AIs mechanically sum all orders without excluding refunds. QianWen Office demonstrated clear business rule understanding, matching row counts before and after cleaning, and providing judgment bases for anomalies rather than random flags. The final Excel charts were directly usable in weekly meetings.

Writing weekly reports reveals an AI's understanding of workplace dynamics.

We provided a project manager with progress on three tasks and a risk point, asking for a <400-word report for the director.

QianWen's report was comprehensive and risk-aware, stating, 'Currently, a 2-person-day gap in frontend resources exists. If not supplemented this week, delivery is expected to delay by 3 working days.' It clearly outlined problems and impacts, avoiding valueless pleasantries.

Sales visit summary organization is another area prone to pitfalls. Salespeople often scribble shorthand notes during client visits, requiring half an hour to organize into structured summaries.

Boundary awareness is crucial here. A client saying 'we'll consider it further' is polite; mentioning a competitor casually is informative. Many AIs naively treat polite remarks as commitments and casual mentions as explicit demands, leading sales to follow up on ineffective summaries.

QianWen Office demonstrated rare sophistication in this scenario. It distinguished between explicit client demands, potential concerns, and off-topic remarks. The follow-up actions it organized were specific, filtering out empty talk and leaving actionable instructions. It even accurately mapped mentioned competitor information to specific companies.

PPT generation is a veteran scenario in AI office solutions but also one with the deepest waters.

Many AI-generated PPTs appear lengthy but are actually full-page image collages, uneditable in practice and useless at work. We tested two PPT scenarios: a Xiaohongshu proposal for a sports brand by an agency and an HR training slide deck converted from a new hire manual.

In both scenarios, QianWen Office generated natively editable PPTs with text boxes and charts directly modifiable. The proposal logic was sound, covering background, strategy, influencer matrix, timeline, and pricing, with placeholder text in the pricing section.

The new hire manual conversion tested summarization ability most critically. QianWen Office broke long paragraphs into key points, retained tables and flowcharts, and preserved all necessary information from institutional clauses.

File organization tests an AI's sense of propriety.

We dumped over a dozen disorganized annual meeting files into a folder, asking it to categorize by type, rename supplier files by date, identify duplicates without deleting, and emphasized, 'Provide a plan first, then act after my confirmation.'

QianWen Office first outlined a detailed organization plan: which files belonged in which folders, suspected duplicates, and renaming rules. It only proceeded after our confirmation, merely flagging duplicates without deletion. This 'request permission before acting' boundary awareness is crucial in AI agent scenarios.

02 Behind the Product Lies Alibaba's Ecological Strength

While the previous scenarios tested QianWen Office's 'individual combat capabilities,' what truly differentiates it from other products is its underlying ecosystem. This became increasingly evident during trial use.

Take e-commerce first. Alibaba's ecosystem in this area cannot be replicated by other companies shortly. QianWen Office integrates a vast array of e-commerce-oriented skill packages and expert suites, covering nearly all daily operations for an e-commerce merchant, from product selection analysis, product design, marketing copywriting, to legal compliance.

For example, people in e-commerce know that product selection is the most challenging task. They need to check the supply sources on 1688, compare prices, review sales volumes, and analyze competitors' pricing strategies and customer feedback. The entire process can easily consume half a day.

However, with QianWen Office, you can directly utilize its product selection capabilities. It can pull supply data from 1688, view best-selling lists on Taobao and Tmall, scrape user reviews of competing products, and finally generate a product selection recommendation report, even calculating the approximate profit margins.

For instance, we asked it to focus on the pet smart collar market and complete the entire process from product selection to uploading promotional materials. The specific tasks involved: cross-platform opportunity scanning, capturing data on this product category from major platforms over the past 90 days, organizing sales volumes, review counts, and negative feedback keywords by price range, and outputting an opportunity matrix; dissecting competitors, analyzing main image selling points, prices, sales estimates, negative review pain points, etc., and finally providing detailed product selection recommendations combined with cost and differentiation analysis, along with uploading a package of main image materials.

The final deliverables were somewhat surprising: three documents, namely a pet smart collar product selection research report, a materials package, and an Excel sheet for the product selection library. The overall work was comprehensive and practical.

It's worth mentioning that these instructions were not manually designed. QianWen Office has a large number of built-in task templates, and users can simply modify keywords to invoke them with one click.

To put it bluntly, this kind of ecosystem-level advantage is far more challenging to replicate than pure model capabilities. Large models are available for purchase by anyone, and Agent frameworks are gradually becoming more and more alike. However, for a product to directly access a company's chat records, delve into an e-commerce platform's operational data, and seamlessly integrate approval and attendance processes, it necessitates years of accumulated enterprise service ecosystems and a wealth of merchant resources.

This is precisely why Chen Yusen emphasized that long-term success hinges on enterprise resources and B-end channels. The product experience merely serves as an entry ticket; the real competitive edge lies in the ecosystem.

03 What's Missing Beyond "Usefulness"

After exploring these scenarios, our overall assessment of Qianwen Office is that it is a highly refined product. In numerous scenarios, it has reached a level where it is "ready to use straight out of the box," rather than being a half-baked product that requires you to tidy up after it.

It has essentially met all the criteria set forth by Chen Yusen.

The interface is straightforward and uncluttered, with no confusing function entries, making it immediately user-friendly. The task execution is highly complete, with most scenarios being handled end-to-end without any need for your intervention. The engineering fallback capabilities demonstrate that significant effort has been invested, and it possesses a fundamental understanding of office rules across various industries. It's not just a rigid tool that blindly follows instructions without comprehending human subtleties.

However, if we measure it against the benchmark of a "truly great AI office product," there is still room for improvement.

For example, in multi-round interactions, all the scenarios we tested involved single-round instructions—you give it a clear task, and it delivers a result. But in real office settings, few things can be clarified in a single attempt. Qianwen Office currently excels in single-round tasks but still has scope for optimization in terms of understanding long-dialogue contexts and facilitating iterative modifications.

Then there's the issue of cross-scenario information flow. Each scenario in our test was treated as an independent entity, but in reality, tasks are often interconnected. For instance, after concluding a sales meeting and organizing the minutes, you might immediately need to update the to-do items in the minutes to the weekly project report and compile competitor information mentioned by clients into an analysis table.

At present, Qianwen Office can handle each individual task competently but falls short in fully achieving cross-task information flow, such as proactively asking if you want to sync the to-do items to the weekly report right after organizing the minutes.

This is also the challenge that Chen Yusen pointed out regarding the need for "enterprise-level Agent collaboration" to address.

Another area for improvement is handling vague requests. In our tests this time, we provided very clear instructions. However, in reality, user requests are often ambiguous.

"Help me check if this contract has any issues," "Summarize the recent industry trends for me," "Help me think of where to go next week"—these types of requests without clear boundaries are how most people interact with AI in their daily lives.

Qianwen Office performs admirably with clear instructions, but when users are unsure of what they want, whether it can proactively ask questions, guide them, and offer suggestions—we haven't witnessed particularly outstanding performance in this regard yet.

But having said that, these are all advanced requirements. As of August 2026, Qianwen Office is already one of the most reliable desktop AI office products we've encountered.

The war in the AI office space has indeed commenced, as Chen Yusen predicted. From Devin sparking the concept in 2024 to everyone scrambling to create demos in 2025, and then to the major players consolidating and competing on implementation capabilities in the second half of 2026, this field is rapidly shedding its hype and bubble.

From this vantage point, Qianwen Office has made a solid start. Whether it can evolve into the "truly great AI office product" that Chen Yusen envisions is still too early to determine. But at least, it's heading in the right direction.

This article is an original piece by Xinmou. For authorization to reprint or business cooperation, please contact us.

— END —

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.