10/08 2026
341

What a model can ultimately deliver depends not only on its own capabilities but also on the work environment it is placed in.
Author|Xu Ting
Editor|Dr. Qi
The Flash model is increasingly resembling a "basic model" in the AI world.
It is fast, inexpensive, and more than adequate for handling daily tasks. As a result, model evaluations have developed a fixed routine: assign the same tasks to several models, see which one provides better answers, and then rank them.
This Flash model evaluation initially followed this established approach.
We used DeepSeek-V4.1-Flash, Yunzhisheng U2-Flash, and Qwen-3.8-Flash to perform two tasks: a deep research report on the mobile phone industry and the development of a match-three puzzle game. U2Flash and DeepSeekFlash excelled in rigor and insight, respectively, while Qwen exhibited significant shortcomings in report fact-checking and game delivery.
However, when we reviewed the entire testing process, we suddenly realized that the test was unfair from the start: the three models did not operate in the same work environment.
DeepSeek Flash and U2Flash ran entirely within WorkBuddy, with access to internet searches, file reading and writing, and tool invocation; Qwen3.8 Flash, on the other hand, was completed in the web experience version of the Qwen AI platform, with only an isolated dialog box throughout the process.
Since we were unable to directly invoke the Qwen3.8 Flash model in the Qwen office desktop version during the initial test, we conducted an additional round of retesting with strictly controlled variables, using the exact same tasks and materials to run again in WorkBuddy.
The results were unexpected. The initially underperforming Qwen3.8 Flash showed a marked improvement in quality during the retest, with numerous data errors in the first round being rechecked and corrected; the previously stuck match-three puzzle game also completed the full loop from locating bugs, modifying code, to actual verification.
The fact that the same model could produce such vastly different delivery results when placed in a different work environment led us to ponder: what exactly are we testing in these model evaluations? Is it the model itself, or the actual outcomes it can deliver once integrated into the workflow?

Qwen Model 'Running Naked' on Its Own Platform
The starting point of this evaluation stemmed from a failed integration.
We originally downloaded the Qwen office desktop version, intending to directly invoke the Qwen 3.8 Flash model within it, but we quickly encountered obstacles. Qwen office explained that the specific model used is determined by the product's session configuration, and users cannot manually switch or specify the underlying model during conversation. The availability of models also depends on the specific entry point, account, and version.
Later, we settled for the web experience version of the Qwen AI platform, found the 3.8Flash entry point, and only then began the formal task execution.
At the time, it seemed like just a change in operational entry point, but upon retrospective review, we realized it was the most significant hidden variable in the entire test.
Let's look at the first round.
The research report task actually contained numerous pitfalls. Among the five material packages and eight task lists we provided, there were data mismatches, unit errors, time frame confusions, and assertions without sources. The game task required the model to develop a match-three puzzle game.
DeepSeek Flash 
U2Flash
U2Flash handled data traps with the most precision, tracing back to the business entity and statistical scope after discovering a discrepancy between "OPPO's 68.72 billion yuan" and the data in Transsion Holdings' annual report. DeepSeek Flash excelled in independent judgment, proposing a bold viewpoint like "chip bifurcation."

The issues with Qwen3.8Flash were concentrated in the "New Phone Landscape" section. The actual new phones mentioned in the materials had not all arrived yet, but Qwen3.8Flash submitted the draft prematurely, using old product lines from its training data to fill the 2026 new phone list: vivo was written as the X200 series, while Huawei had models like Pura 80 and Mate 70. It failed to accurately capture Huawei's 2025 annual report data on revenue, net profit, and terminal business; a set of data for OPPO was also judged incorrectly.
The shortcomings exposed in the game task were even more severe. It only generated a plain text code, which needed to be manually copied into a notepad or editor and saved for execution. After opening in a browser, the deliverable remained at a very primitive web demo stage, playable but with a stuck scoring logic, resembling a running but soulless empty shell.
Qwen3.8flash web version
Of course, these performances were influenced by factors inherent to the model itself, such as using old historical knowledge to fill in current information gaps and lacking a self-checking mechanism after generation. However, in the first round of testing, Qwen3.8Flash did not have access to the same "weapons" as the other two models.
DeepSeek Flash and U2Flash operated within an AI office product like WorkBuddy. When consulting Huawei's annual report, IDC shipment data, or business data, they could conduct internet searches; when encountering data traps in the materials, they could initiate cross-verification; after completing the report writing, they could directly save local files, organize them by chapter, and package them for output.
At this point, the situation became awkward. We thought we were horizontally comparing the technical capabilities of three Flash models, but in reality, we were comparing three vastly different model invocation and delivery methods.
On one side, the model was connected to a workbench with internet access, tool invocation, local reading and writing, and self-checking capabilities; on the other side, the model's capabilities were confined to a single web dialog box, "running naked." The delivery gap obviously could not be simply attributed to the model's "lack of intelligence."
Therefore, we reset the variables, keeping the tasks, materials, and models unchanged, and only moved Qwen3.8 Flash into the same work environment, WorkBuddy, to rerun the entire process.
The results were unexpected. In the research report task, Qwen Flash almost replicated the engineering standards previously set by U2Flash. It immediately output a list of seven error corrections, each with cross-verification.
Qwen Flash retest version
It also established a three-tier source system of "official, institutional, and third-party estimates," with individual annotations in the text. In the initial test, Qwen made numerous factual errors in the "New Phone Landscape" section; after the retest, it not only completed the eight-chapter report but also proactively conducted a round of data error corrections.


The differences were even more stark in the game development task. In the web version, Qwen could only provide a static text code; in the WorkBuddy environment, however, when rendering bugs or scoring logic issues arose, Qwen Flash could re-read the error logs through the workbench's execution environment, completing the closed loop (closed loop) of "locating errors, modifying code, and locally recompiling for verification," ultimately delivering a playable game with decent entertainment value.
The reason models exhibited vastly different delivery ceilings was that they gained the conditions to re-verify, re-read contextual materials, invoke environmental tools, and proactively correct errors. Capabilities previously trapped in a dialog box were maximally released once a complete work chain was established.
DeepSeek Flash 
U2 Flash
This also led us to re-understand the current path divergence in AI office products.
Qwen Office tends to integrate model capabilities, application tools, and interaction interfaces into a highly standardized delivery solution. This "out-of-the-box" experience significantly lowers the usage threshold but also restricts users' ability to customize the underlying model. In pursuit of "absolute uniformity" in the front-end experience, the model's strain potential in complex, dynamic scenarios is inadvertently "sealed."
In contrast, WorkBuddy does not forcibly bind to a single model ecosystem. It allows users to freely mount DeepSeek, U2, or Qwen within the same workspace according to task requirements, treating models as "pluggable computing engines" and returning the choice entirely to the task itself. Here, models are more like computing resources that can be replaced at any time, with the workbench responsible for organizing tasks, context, tools, and models.
Under this path, the core value of AI office products no longer depends on "how strong my model is" but rather on "whether I can precisely and seamlessly embed the most suitable model into complex real-world workflows."
Agent Competition Escalates: From 'Model Entry' to 'Task Entry'
Looking back at this test, the three experiences correspond to three entirely different AI usage methods, also mapping out the three stages of AI application evolution.
The first stage involves conversational AI assistants with single-model chat capabilities, primarily targeting personal scenarios. Users access pre-configured suites provided by vendors, where models, tools, and interaction logic are fully packaged. At its core, it is an auxiliary tool designed to enhance personal efficiency. Its strength lies in simplicity, but the trade-off is a complete surrender of choice.
The second stage introduces a schedulable 'Agent Workbench,' transitioning from 'chatting' to 'task completion' by enabling the on-demand integration of multiple models and fostering team collaboration. At this point, AI evolves into an Agent capable of invoking tools, marking the true beginning of its work execution.
Models shed their 'product gateway' Halo (aura) and are 'dimensionally reduced' to pluggable components within workflows. This fundamentally transforms the logic of 'using models.'
In the past, users would say, 'I open QianWen because I want to use it.' Now, the approach is, 'I have a task, so I schedule the appropriate model into my workbench.' The competition among AI applications has shifted from 'model gateways' to 'task gateways.'
It must be acknowledged, of course, that workbenches cannot conjure intelligence that does not exist within the models themselves. The upper limits of a model's reasoning, factual judgment, and self-inspection capabilities remain hard constraints.
QianWen's 'delivery results falling short of expectations' in certain tasks can be attributed not only to the inherent limitations of the model itself but also to the experience-based platform form (form) of the trial version, which lacks Internet search access and an engineered delivery environment—both significant drawbacks.
Model disparities are both amplified and mitigated by the 'environment.' As mainstream models become increasingly homogeneous in their underlying parameters and foundational reasoning, the decisive factor in user experience has shifted from 'which model to invoke' to 'in what engineered environment the model is invoked.'
Beyond knowledge density, the hierarchical stratification of model capabilities is being redefined by the degree of engineering and the delivery closed loop (closed loop). An 'environment' capable of search, storage, debugging, and delivery has become a prerequisite for unlocking a model's full potential.
A small but revealing detail emerged during testing: we interacted and issued commands through AI hardware co-branded by Ulanzi and WorkBuddy. The integration of hardware introduced a new physical node to workflows previously confined to screens: AI transitioned from a web-based dialog box to a desktop workbench and further extended into real-world hardware devices. 
In the past, discussions about AI often centered on 'where the model runs.' The addition of hardware prompts us to further consider: 'In the future, where should users awaken (activate) AI?'
From web text boxes to desktop-level Agent workbenches and then to integrated physical nodes, the next phase of AI-powered office work no longer belongs to those who preempt (seize) 'model gateways' but to those who control 'task gateways.'
The Next Half of AI-Powered Office Work: From Task Gateways to Enterprise AI Infrastructure
In the past, evaluating an AI product's usability typically involved two metrics: model intelligence and response elegance. However, as Agents genuinely intervene (intervene) in real-world business operations, the evaluation framework has undergone a fundamental transformation.
Can it locate materials? Can it correctly invoke tools? Can it maintain contextual awareness? Can it autonomously correct errors when they occur? Most critically, does the final output represent a 'deliverable' ready for immediate reuse or a 'half-finished product' requiring human intervention to clean up?
If personal efficiency tools address how employees use AI, then as AI office products penetrate enterprise scenarios, they inevitably encounter the next formidable barrier: how enterprises manage AI.
This marks the next stage of AI application evolution: the era of fully governed AI organizations, or enterprise AI infrastructure.
Users ascend from 'using AI' to 'managing AI,' enabling unified scheduling of multiple models, enterprise-level permission isolation, and collaborative governance. Models are further decoupled and integrated into enterprise organizational structures, serving as digital employees. At this stage, the core demands extend beyond mere efficiency gains to include precise control over computational costs and data security risks.
In terms of cost control, enterprises must consider: Who pays the Token bills? How are computational quotas allocated across departments?
Regarding permissions and governance: Who can invoke specific models? Who can access core databases? Where are the security boundaries when AI operates on local files?...
Today, a fierce battle for 'enterprise AI infrastructure' is underway, with major players revealing their strategies.
At the recently concluded Cloud Town Conference, the integrated QianWen Office was propelled to the absolute center of AI implementation demonstrations. Not only did it accelerate the Personal Agent strategy, but it also unveiled private data management capabilities such as 'Enterprise Context,' aiming to embed AI into every cog of enterprise collaboration through deep integration with DingTalk's ecosystem and cloud computing power.
Baidu, meanwhile, is rapidly consolidating its internal office Agent resources. Previously, capabilities from products like Baidu Search, NetDisk, Maps, Miaoda, and Famous were integrated into Baidu Dazi via Skills. Following the July launch of Baidu Dazi Enterprise Edition, Dodo was fully merged into 'Baidu Dazi' the next month, attempting to construct a closed-loop ecosystem for B-end office applications.
Recently, WorkBuddy Enterprise Edition underwent a major upgrade, bundling five office capabilities—Tencent Docs, Tencent Lexiang, Tencent NetDisk, the Security and Efficiency Operations Platform, and Tencent AI Design Agent Ardot—into a unified suite. It also established an 'identity system' for enterprise AI. Currently, WorkBuddy offers enterprise-grade capabilities such as SSO single sign-on, organizational structure management, member usage control, and unified subscriptions.
From accounts and organizational structures to Credit usage, Skills, Agents, and enterprise data, administrators can now manage these AI resources within a unified framework. WorkBuddy is transforming what was once a personal-oriented Agent workbench into an infrastructure capable of supporting organizational-level AI collaboration.
This also imbues the concept of 'task gateways' with new meaning.
For individuals, a task gateway means 'I entrust tasks to AI.' For enterprises, it means 'The enterprise entrusts an entire set of business processes to AI.'
The former addresses how to enable AI to assist individuals in completing work; the latter addresses how to manage, invoke, and constrain a group of AIs within an enterprise while ensuring sustained output.
From models to Agents and then to enterprise AI infrastructure, different vendors follow distinct paths but converge on a single focal point: who can embed models into enterprises' real business processes.
As foundational model capabilities stabilize, the engineering construction centered around 'execution and delivery' is only just beginning.
