08/12 2026
378

Author | Yugu
Disclaimer | The featured image is sourced from the internet. This original article by Jingzhe Research Institute may not be reproduced without permission.
Who could have predicted that the input method—a seemingly niche and static market—would be stirred up again in the age of AI?
Since the latter half of 2025, Doubao Input Method and Qianwen Input Method have been rolled out one after another, with voice input as their core feature. Even WeChat Input Method, launched in 2022, received a significant 3.0 version update in December of the previous year, focusing on enhancing its voice input capabilities. Overseas, Wispr Flow, which began public testing in 2023, has secured a cumulative $81 million in financing, with its valuation surpassing $2 billion.
While everyone is focused on the rapid iteration of large models and intelligent agents, why has the input method suddenly regained popularity? This seemingly odd phenomenon actually signals a “future war” in the AI era.
There is already an industry consensus on why major companies are suddenly investing in voice input methods in unison. As AI dialog boxes replace search boxes as new entry points for traffic, user online time and internet traffic are being redistributed. Beyond large models, input methods—as the most frequently used basic function that spans various applications and scenarios—also have the potential to become a “system-level entry point.” Thus, input methods serving as AI entry points have become a battleground for major companies.
However, this reasoning only explains why companies “want to develop input methods” and not why they “want to develop voice input methods.” The distinction of “voice” is precisely the key.
In the minds of most ordinary users, voice input might simply be “you speak, it writes,” with its core principle being to recognize user speech and transcribe it into text. However, due to early reliance on statistical models for speech recognition, accuracy was significantly affected by environmental factors, accents, and colloquial expressions, making voice input far less accurate than keyboard input. Therefore, although voice input might have had an advantage in speed, its accuracy issues made it difficult to truly gain widespread adoption.

But now, with the rapid advancement of large model capabilities, the accuracy issue has largely been resolved. More importantly, with the enhancement of AI capabilities, input methods are no longer just tools for transcribing speech.
For example, Doubao Input Method is developed based on the same voice large model as Doubao. In addition to boasting an accuracy rate of over 98%, it also features automatic error correction for keyboard input, enabling intelligent association of text, symbols, emojis, and dates through contextual semantic analysis. Qianwen claims to support functions such as colloquial speech purification, intelligent error correction, contextual understanding, creation, Q&A, and translation.
In other words, voice input method products based on AI large models are equivalent to automated text processors with standardized language capabilities. More vividly, voice input methods are like “pen substitutes,” allowing users to not only reply to messages in daily work and life but also directly compose using AI voice input methods. This significant improvement in product functionality has enabled input methods to break through user value barriers, providing an opportunity to reopen the market. But more crucially, the development of AI itself also necessitates voice input methods.
From the rise of generative AI to its current popularity in office settings, a problem has gradually surfaced. Chatbots can provide a travel guide in just a few minutes, but users may need more than ten minutes to input their travel needs into the dialog box. Office agents can complete a work report in just a few minutes, but users must first spend time clarifying the project background and specifying the report format before making their request.
Clearly, as model capabilities become increasingly powerful and AI functions become more diverse, the efficiency of humans using AI does not seem to have achieved a leap forward. The reason for this contradiction is that the “typing” interaction method has become a bottleneck limiting the efficiency of humans using AI.

If we examine the behavioral processes of typing and voice input, we find that typing involves thinking, organizing language, and then typing on the keyboard. Voice input, on the other hand, not only omits the typing step but, with the support of model capabilities such as automatic error correction and contextual semantic analysis, the step of organizing language can also be almost eliminated. Therefore, voice input is more efficient than typing, and voice input methods enhance human productivity when using AI by improving “input efficiency.”
It is worth mentioning that because voice input almost “eliminates” the typing step, making it more efficient, many users who love Vibe Coding have even started to customize voice input keyboards with only four keys, reflecting the unique value of voice input methods for using AI.
In scenarios where AI is used through voice input, a new interaction process is emerging: users speak through the input method, the input method converts the speech into text, and AI understands the text and completes the task. In this process, the input method acts as a “joystick” for humans to use AI, so the input method can be seen as an entry point for calling AI. However, if we look further ahead, must the AI entry point be the input method?
At the end of July, OpenAI provided a highly noteworthy answer. OpenAI announced an update to the ChatGPT desktop application, extending the voice capabilities of ChatGPT Voice further into work scenarios on the desktop, such as ChatGPT Work and Codex.

*Image source: OpenAI social media account
This means that users can directly command intelligent agents to work through voice and can also use voice at any time to check task progress, continue asking questions, interrupt, or reissue instructions to the agent. In this scenario, voice is not just an input method but begins to serve as an operation method for the computer system—this may be the true watershed of human-computer interaction in the AI era.
In the realm of “voice input,” domestic AI giants and OpenAI have actually taken two different paths. What domestic giants are currently doing is integrating AI into input methods and using the tool's attributes to meet specific user scenarios for using AI. OpenAI, on the other hand, is attempting to integrate voice into the AI workbench, allowing users to operate the computer directly using AI.
The fundamental difference between the two approaches is that the former competes for “input rights” in users' WeChat, browsers, chat software, and office software, while the latter competes for “operation rights.” To some extent, OpenAI's approach is clearly more in line with the AI Native style. However, these two approaches are not inherently superior or inferior; they simply differ naturally under different market conditions.
Many people may ask why domestic giants do not directly allow users to operate AI using voice. This is not a technical issue but an ecological one.
In the domestic AI market, almost every major tech company has its own combination of “large model + AI office.” Tencent has the Hunyuan model, WeChat, Enterprise WeChat, and WorkBuddy; Alibaba has Qianwen, DingTalk, and Qianwen Office; ByteDance has Doubao, Feishu, and a vast ecosystem of content and office products.
For these companies, developing a voice input method is hardly challenging. They are fully capable of adding a voice interaction feature to the AI workbench, like OpenAI, allowing users to call their AI anytime in daily scenarios.

However, creating an AI that can control a computer first requires testing user trust in the product. Previously, OpenClaw experienced a sudden loss of control, leading to security incidents involving the leakage of user sensitive data and confidential files, triggering a crisis of trust in agents among users.
Additionally, controlling a computer requires operating system permissions, application permissions, file permissions, and browser permissions, as well as addressing compatibility issues between different software. These issues undoubtedly bring significant communication costs and workloads in terms of user education. Before AI products see a mature commercialization path beyond the “Token economy,” domestic giants have no reason or need to take such a big step.
Input methods, as familiar tool products, are clearly more easily accepted by users. Therefore, the strategy of domestic giants to focus only on input methods is clearly more in line with the overall layout of the product ecosystem, with a clearer commercial logic. However, this does not mean that domestic giants who choose to focus on input methods first do not see the direction of AI operating systems. On the contrary, input methods may just be a stepping stone for them to switch to this direction later.
If we trace the changes in domestic giants' AI products over the past year, we find that the main product direction has shifted from “answer-all” AI assistants to agent platforms. This year, AI office has become a new stage for major companies to deploy AI. On the other hand, voice input method products have also been intensively released and updated since the end of last year.
Simply looking at the timeline, AI office and voice input methods, two seemingly independent product categories, unexpectedly give a sense of gradually converging trends. In reality, AI office also provides a viable entry point for desktop voice input.
In the past, AI office products were essentially just chat boxes where users input their needs and waited for the results. However, now agent-based AI office products no longer just answer questions but can write documents, create tables, search for information, analyze data, and generate code. The “all-in-one” AI office workbench is increasingly resembling a desktop agent.

This feeling is like the AI office workbench integrating all software functions and file materials on the computer, allowing users to simply input their specific needs and let AI do the work. Once users become accustomed to “letting AI do the work,” voice input becomes increasingly important.
Imagine an office scenario in the near future: you turn on your computer, log in to no software, simply press the voice input button on the keyboard, and tell AI, “First, organize yesterday's meeting minutes, list today's to-do items, then update the project report due this week to its best state, and remind me to follow up before I leave work.” AI then automatically executes the tasks. Throughout the process, you do not need to operate any software or even know where the meeting minutes or project report is located.
“Letting users only need to speak and leaving the work to AI” is what a truly mature AI workbench looks like and the ideal way to use AI in the AI era. And voice input methods also have the potential to become a system-level entry point for AI.
In the mobile internet era, when discussing system-level entry points, more people might think of apps, browsers, search engines, and input methods, as users need these tools to access the digital world, obtain information, complete work, and execute tasks online. However, in the AI era, the core capability of AI is not to help users find a tool but to complete tasks directly. Under such an operating system, input methods, as tools directly called by AI, will also “fade into the background” like other tools.
What is even more imaginative is that when voice input becomes a daily and globally callable, seamless interaction method, AI will be able to capture not just task instructions in office scenarios but also users' thoughts, plans, and intentions. What does the user want for lunch? What are their weekend plans? What work problems have they encountered? These needs will be revealed in daily expressions, and AI will not need to “eavesdrop” on users' input methods to obtain “search keywords,” as some apps do today, because AI itself will already serve as the input method function.

Therefore, the purpose of major companies deploying voice input methods is not to create a “new wheel” that runs faster but to secure a ticket before the arrival of a new operation method for human-AI dialogue. Discussions around voice input methods should not be about “whether it can replace the keyboard?” but rather “when voice becomes a system-level operation method, do we still need the act of input?” and “when input becomes invisible, what kind of AI products do we need?”
Because when AI can understand natural language, context, and user intentions, all humans need to do is express their purpose. Just like casting spells in a magical world, perhaps one day AI will enable users to complete all work simply by speaking, truly achieving “speak and it shall be done.”
In the past, technological progress often left behind a generation. Basic operations like downloading apps and searching for information online are almost “factory settings” for the post-2000 generation, but many elderly users still need to relearn and adapt to system operations to use smartphones. In the AI era, where voice becomes a system-level entry point, the barrier to technology use is gradually weakened, and users only need to express their purpose, leaving the rest to AI. This is also the direction where AI truly has the opportunity to change ordinary people's lives.
Voice input technologies do not represent the final destination of AI; rather, they serve as a gateway to the AI era. Whoever can establish themselves as the "natural language" medium for human interaction with the digital realm will unlock this gateway and truly command the "system-level entryway" into the AI age.