AI Smartphone War: The Endgame is a World Without AI Phones

07/27 2026 518

Nearly every industry is discussing 'rebuilding with AI,' and smartphones are no exception.

Apple introduced Apple Intelligence, Samsung continued to double down on Galaxy AI, while Huawei, Xiaomi, OPPO, and vivo have also integrated large models into their phone systems.

From image removal and call summarization to screen understanding and cross-app operations, 'AI smartphones' have quickly become a common theme at new product launches across the board.

In July this year, China's Cyberspace Administration unveiled seven generative AI services for mobile devices that had completed regulatory filing all at once. Nearly all major manufacturers were on the list, signaling that on-device AI is entering a broader phase of compliant implementation.

Source: Cyberspace Administration of China

Research institutions are quite optimistic about this trend.

Counterpoint predicts that by 2026, 45% of global smartphone shipments will feature generative AI capabilities, rising to 52% by 2027; in the premium market with wholesale prices above $400, AI is nearly becoming a standard feature.

Source: Counterpoint

But can these AI capabilities really boost smartphone sales? Not necessarily.

CINNO Research data shows that in the first quarter of 2026, domestic smartphone sales declined by 6% year-on-year. Factors such as rising storage costs and price hikes for new models have extended the user replacement cycle to over 36 months.

Counterpoint also points out that while generative AI has become widespread in premium phones, it has not yet created a sufficiently strong incentive for device upgrades.

This is the most delicate aspect of 'AI smartphones' at present.

Manufacturers are investing more and more, with an increasing number of AI-equipped products, yet a clear new category akin to how smartphones replaced feature phones has not emerged in the market.

The changes brought by AI are more evident at the chip, system, and software service levels rather than in redefining the smartphone itself.

In other words, AI smartphones may not represent the next generation of hardware after smartphones but rather a foundational capability that smartphones are now incorporating.

When AI becomes standard, how meaningful is the concept of an 'AI smartphone'?

More AI smartphones are emerging, yet they increasingly resemble an existing category rather than a new one.

The earliest batch of 'AI smartphones' actually had a relatively clear hardware threshold.

IDC once defined 'next-gen AI smartphones' as those equipped with high-performance NPUs capable of efficiently running generative models on-device, using an NPU computing power of 30 TOPS as the primary criterion.

Source: IDC's 'The Future of Next-Gen AI Smartphones'

According to this definition, AI smartphones are primarily about hardware upgrades.

Stronger NPUs, larger memory, and power and storage configurations adjusted for local models mean these capabilities initially appeared only in a few high-end models.

However, with the compression of smaller models, on-device-cloud collaboration, and improvements in chip computing power, this boundary is rapidly blurring.

Apple's related capabilities now cover multiple iPhone generations; Samsung has extended features like Gemini and Circle to Search to its Galaxy A series. Huawei, Xiaomi, OPPO, and vivo have also brought AI removal, call summarization, and voice assistant capabilities to mid-range products, some even through system updates.

Selling points once exclusive to flagship models are now trickling down the product lineup.

This raises the question: What exactly qualifies as an 'AI smartphone'?

Must it run large models locally, or is cloud access sufficient? Does it count if it can generate images and summarize content, or must it understand user intent and complete tasks across apps?

Different institutions adopt varying standards.

IDC emphasizes on-device computing power and NPU performance, while Counterpoint places more weight on whether a phone can use large pre-trained models to generate content, understand context, and complete tasks, with both on-device and hybrid approaches counted.

By 2026, Counterpoint further proposed standards for 'agent phones,' requiring not only understanding user intent but also the ability to cross two or more services and autonomously complete multi-step tasks.

From content generation to task execution, the expanding definitions are blurring the product boundaries of 'AI smartphones.'

To put it bluntly, a mid-range phone with AI removal, a flagship phone running large models locally, and an agent phone completing tasks across apps may all be called AI smartphones despite being at different capability stages.

A similar process is familiar in the smartphone industry.

In 2000, Sharp and Japanese carrier J-Phone launched the J-SH04, integrating an 11-megapixel camera into a phone, making 'camera phones' a new product label.

In December 2013, China's MIIT issued the first batch of 4G licenses, and for the next few years, '4G phones' were a key selling point at launches.

As cameras and 4G networks became ubiquitous across products, the terms 'camera phone' and '4G phone' gradually faded without forming an independent category beyond smartphones.

'AI smartphones' are clearly repeating this process.

What's truly changing is how phones are operated.

If AI smartphones merely add a few new features, there's no essential difference from past product upgrades.

The earliest AI capabilities were mainly embedded in albums, recording apps, keyboards, and search bars. Image removal, content summarization, call translation, and copywriting generation all had clear entry points.

Users first opened the corresponding feature, then handed a picture, recording, or text to the system for processing.

Apple's writing tools and photo cleanup, OPPO's AI removal, call summarization, and copywriting generation largely fall into this category.

These features indeed improve efficiency but do not alter the phone's original operational path. AI simply enters an existing function and accelerates part of the workflow.

Subsequently, manufacturers began enabling phones to understand more context.

The processing targets were no longer limited to content actively selected by users but also included the current screen, as well as schedules, photos, files, and location information on the device.

vivo's Blue Heart Assistant can recognize text, addresses, and products on the screen, providing search, shopping, and navigation entry points; Honor's YOYO pushes service cards based on flight, delivery, and travel information; OPPO's Xiao Bu Memory can save web pages, images, and meeting minutes, supporting subsequent retrieval.

At this stage, the system began establishing connections between different information, but it still acted as an assistant—organizing information, suggesting services, recommending entry points—with the final choice and operation remaining with the user.

What truly sets AI smartphones apart from ordinary feature upgrades is the agent experiments of the past year.

Huawei's 'Xiao Yi Assistant' can now enter certain shopping, travel, and video apps to complete tasks like shopping, booking flights, and downloading videos according to instructions. Users can also teach Xiao Yi a set of operational steps through demonstration.

Source: Huawei's official website

Honor's YOYO also supports managing app notifications, querying or canceling auto-renewals, and ordering beverages.

Google has begun testing Gemini's multi-step tasks on Pixel 10 and Galaxy S26 series, with initial scenarios including ride-hailing, meal reordering, and grocery shopping.

Source: Google

In the past, users first had to decide which app to open, then proceed with searching, filling in, filtering, and confirming.

At the agent stage, manufacturers hope the system will take over more of these steps. Users only need to state their needs and confirm at critical points like payment and authorization.

However, current phone agents are far from mature.

'Xiao Yi Assistant' is still in public beta and does not yet support many apps like WeChat, finance, email, and office documents; pages involving sensitive or private information also require manual operation.

Google's multi-step tasks are also in beta, initially covering only some markets, models, and apps, with users needing to supervise the process.

In other words, there's still a significant gap between a smooth demo at a launch event and reliably handling accounts, payments, privacy, and exceptions in everyday environments.

Whether the model can understand instructions is just one factor; how much system access is granted, whether apps are willing to integrate, how payments and privacy are handled, and who assumes responsibility for operational errors will all determine how far these features can go.

But this doesn't prevent manufacturers from vying for position early.

In the past, users first chose an app, which then provided the service. In the future, users may first submit their needs to the phone, which then decides which platform to call upon.

This way, the phone system has the opportunity to stand between user demand and app services, participating in the traffic and task distribution once controlled by apps.

This is also the commercial logic behind manufacturers' layout (layout) of AI smartphones.

The person most eager to redefine smartphones still starts by building one.

As agents begin to handle more phone operations, the competition is no longer limited to traditional manufacturers.

Phone makers control the hardware, operating system, and user entry points. Even if large model companies enter the phone space, which data they can access, what permissions they receive, and which apps they can integrate with ultimately depend on coordination with terminal manufacturers.

Unwilling to settle for just providing models, building their own terminals becomes a direct path to retaining complete product definition rights.

A recent notable example is Jueyue Xingchen (StepFun).

In January this year, Yin Qi became chairman of StepFun. A graduate of Tsinghua's Yao Class, Yin co-founded Megvii (Face++) in 2011 and is a leading figure from the previous wave of computer vision entrepreneurship.

From computer vision and smart cars to large model terminals, Yin's layout (layout) over the years has consistently focused on bringing algorithms into specific devices and real-world scenarios.

Thus, StepFun building a phone does not seem like a sudden decision.

In July this year, StepFun unveiled its terminal brand STEPX, agent operating system Step AOS, and personal agent Amoo, along with its first phone, STEPX Neo.

Why would a large model company enter the phone industry, known for its high barriers, thin margins, and intense competition?

Yin explains that under the current phone ecosystem, StepFun's models, operating system, and agents would struggle to enter third-party terminals in their complete form. Without first creating their own product, Step AOS would find it difficult to form a complete closed loop (closed loop).

This is clearly a different approach from phone makers simply adding a few AI features.

According to Yin's vision, users won't need to enter an app to find a function first; instead, they'll submit their needs to Amoo, which will then handle understanding, planning, and service invocation.

StepFun has also partnered with platforms like Ctrip, Alipay, Didi, Meituan, WPS, and CapCut, hoping to invoke services like ticketing, consumption, travel, and office work through interfaces rather than relying solely on screen recognition and simulated taps.

However, STEPX Neo is still in the demonstration phase, with its full specs, pricing, and release date yet to be announced.

More interestingly, to redefine smartphones, StepFun's first offering is still a phone.

STEPX Neo retains a screen, camera, and the basic structure of a conventional smartphone. Step AOS, while emphasizing 'agent-native,' remains compatible with existing systems like Android, Linux, and RTOS, integrating into the mature app and service ecosystem.

The reason is simple.

Accounts, payments, communications, photos, contacts, and schedules are all centralized on phones. Travel, e-commerce, food delivery, and content services have also operated around phones for years. For agents, no consumer electronics device offers such a complete ecosystem of data, permissions, and services.

Yin refers to hardware as the 'container' for agent services. StepFun builds phones not to replicate a traditional phone company but to secure a long-term touchpoint with users for its models and agents.

In fact, attempts to bypass phones and redesign AI terminals are not unheard of.

The most notable example is Humane's AI Pin.

Launched by two former Apple employees, this device lacks a traditional screen, relying instead on voice, camera, and laser projection for interaction, attempting to provide an alternative to phones. However, by February 2025, Humane had stopped selling the consumer version of AI Pin, with its technical assets later acquired by HP, ending in a hasty retreat.

AI glasses, meanwhile, have fared relatively well.

EssilorLuxottica disclosed that its Ray-Ban Meta and Oakley Meta smart glasses, developed in collaboration with Meta, sold over 7 million units in 2025. Functions like shooting, calling, translation, and visual Q&A on AI glasses also handle some phone scenarios.

However, this path comes with a high cost.

In the first quarter of 2026, Meta Reality Labs, which includes smart glasses, Quest, and other VR and AR businesses, reported revenue of $402 million and an operating loss of $4.028 billion, reflecting the long-term and heavy investment required behind new terminals.

It can also be seen that AI Pin attempted to break away from smartphones but ultimately failed to establish a sufficiently complete product and service system. In contrast, AI glasses, which maintain a connection with smartphones and initially support functions like photography, voice, and visual perception, have found a market more quickly.

Of course, Yin Qi also believes that future intelligent terminals may no longer be smartphones in today's sense, and users may spend less time looking at screens or operating devices.

This prediction may come true in the more distant future, but at least for now, those most eager to redefine smartphones still cannot bypass them.

The true value of AI smartphones lies beyond the device itself

At this stage, no consumer electronics device is better suited to carry intelligent agents than smartphones.

However, as large models and intelligent agents gradually become standard features in new devices, hardware upgrades alone will struggle to create long-term differentiation.

Thus, manufacturers must answer another question: beyond selling hardware, what other businesses can AI smartphones pursue?

The visible directions include membership subscriptions, model access fees, value-added services, and transaction commissions from task-based services. The Token revenue envisioned by StepFun also falls under this logic.

However, whether consumers are willing to pay separately for intelligent agents in their smartphones remains uncertain.

Users already pay for smartphones, cloud storage, music, and video services. If intelligent agents merely organize information and generate content, it will be difficult to justify new fees. Only by truly taking over high-frequency tasks and saving significant time can subscription or pay-per-use models become viable.

Compared to directly charging users, service distribution may be more realistic.

When users delegate tasks like "booking a hotel," "buying a plane ticket," or "hailing a ride" to their smartphones, the system must choose among different platforms. Which service is prioritized and where orders ultimately flow both hold commercial value.

Today, apps not only control access points but also manage the entire user journey from search and comparison to checkout.

If intelligent agents become the new primary access point, the relationships between smartphone manufacturers, large model companies, and internet platforms will shift.

Smartphone manufacturers control devices and systems; large model companies provide understanding, planning, and execution capabilities; and super apps possess accounts, products, payment, and fulfillment systems. All three parties hope intelligent agents will access more services, but none are willing to easily surrender their user relationships.

For smartphone manufacturers, the deeper intelligent agents penetrate, the greater the system's influence over apps and transactions.

For large model companies, providing models alone still yields only technical service fees; direct user access is necessary to secure subscriptions, access fees, and transaction commissions.

For super apps, intelligent agents may bring new orders but could also weaken direct platform-user connections. Decisions on how many interfaces to open, how much data to share, and how far to let intelligent agents go involve not just technical considerations but also practical interests.

Thus, the next round of competition may not depend on whose model has more parameters but on who can convince users to entrust their needs to the system first, who can access enough services, and who can earn trust regarding payment, privacy, and liability boundaries.

The greater commercial value of AI smartphones may lie not in the hardware itself but in capturing user demand entry points and participating in subsequent service selection and transaction allocation.

Whether revenue comes from subscriptions, Tokens, or transaction commissions, AI will gradually become a foundational capability of smartphones.

In this way, the endpoint of AI smartphones is a world without "AI smartphones."

Source: Chaoyang Capital Theory

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.