Why Is the 'Last Mile' of Large-Scale AI Models So Challenging?

09/14 2026 345

In the first week of September 2026, the AI sector witnessed an unprecedented "Intensive Release Week." On September 1, Anthropic set the pace by launching Claude Fable 5.1 and Claude Mythos 5.1, models tailored for coding, knowledge work, and extended-duration agent tasks. The following day, Google unveiled Gemini 3.8 Flash and the cybersecurity-focused Gemini 3.8 Flash Cyber—just three weeks after the release of Gemini 3.7 Flash, marking the third iteration of the Flash series within six weeks. Meta soon followed with Muse Spark 1.3. In the early hours of September 4, OpenAI officially launched its new flagship model, GPT-6 Astra. Company president Greg Brockman declared at the event, "Welcome to the AGI era." Astra achieved near-perfect scores across multiple benchmarks: 97.6% on FrontierMath Tier 4, 99.9% on ARC-AGI-3, and a flawless 100% on the ExploitBench cybersecurity test, according to OpenAI's official data.

Model parameters are expanding, context windows are lengthening, and reasoning capabilities are improving. Nearly every metric indicates that the technological prowess of large models is advancing at an unprecedented rate.

Shifting our focus from international to domestic developments, significant progress is evident. In July, Moonshot AI released Kimi K3, boasting 2.8 trillion parameters, a sparse mixture-of-experts architecture, native support for visual understanding, and a context window accommodating up to 1 million tokens—making it the world's largest open-source model at the time. Ranking third in intelligence by independent evaluation agency Artificial Analysis, Kimi K3 quickly garnered global attention from the AI industry and research institutions after its release.

However, when the focus shifts from laboratories and launch events to real-world business and daily life, a different narrative emerges.

01. The Gap Between Hype and Reality

Consider the data. McKinsey's 2025 global survey revealed that approximately 88% of enterprises have integrated AI into at least one business function, yet only about 39% report measurable profit contributions. Around the same time, a report by MIT NANDA noted that roughly 95% of enterprise-level generative AI pilots failed to deliver measurable operational returns in the short term. Gartner's survey was even more direct: over 90% of enterprises worldwide have initiated generative AI pilots, but only about 41% have successfully transitioned to production environments and achieved scalable value. Focusing specifically on agentic AI, only 16% of enterprises have deployed it in production, up just 7 percentage points from 9% in 2025.

These figures converge on a single conclusion: AI experiments are widespread among enterprises, but few succeed or generate profits.

IDC predicts that by 2026, 50% of AI-driven digital application scenarios will fail to meet ROI targets. Another Gartner survey shows that only 11% of CFOs observed tangible financial value from AI in 2025, and just 8% of Chinese enterprises achieved revenue growth through AI.

A substantial gap exists between technological hype and commercial pragmatism.

02. The Eroding Moats

This gap did not arise in isolation. A notable phenomenon is that as large models become more capable, the application ecosystem built around them appears increasingly sparse.

In the early days of ChatGPT, the market saw a surge of AI applications designed to "fill model gaps" as their commercial rationale—legal tech companies reduced model hallucinations using professional databases, medical consultation apps improved reliability through secondary verification, and marketing copy tools generated content using carefully designed prompt chains. These applications capitalized on the early limitations of large models, erecting temporary support structures around the models' weaknesses through engineering solutions.

However, as model iterations accelerated, these moats quickly crumbled. GPT-4 eliminated precision gaps in legal citations, Claude's continuous reasoning improvements significantly enhanced reliability in professional scenarios like medical consultations, and the expansion of context windows from thousands to millions of tokens rendered copy tools' reliance on external memory mechanisms obsolete. The challenge of scaling application-layer value is also reflected in macro data—a report by the University of Oxford revealed that while global annual AI investment has surpassed $400 billion, only about 33% of enterprises have successfully moved AI projects from pilots to scalable applications.

Even more concerning are changes at the user level. While the number of standalone apps labeled "AI-driven" in app stores continues to grow, few head products exceed 1 million monthly active users, and many apps quickly lose users and stall in revenue shortly after launch. For many users, the primary reason for switching AI tools is hearing that "the new model performs better," rather than "the new app solves problems the old one couldn't."

Users lack loyalty because all AI products are starting to look alike—uniform dialog boxes with input fields. When interaction methods converge and underlying capabilities derive from the same batch of large models, where does the independent value of the application layer lie?

03. Structural Dilemmas

If the application layer's predicament is merely superficial, deeper issues lie in the structural obstacles to enterprise AI deployment.

High costs top the list. The computational resources required for large model training, along with the electricity and energy expenses behind data centers, ultimately burden enterprises. While companies like DeepSeek have worked to reduce training costs—in 2025, DeepSeek brought the total post-training cost for R1 series reinforcement learning down to about $294,000—inference costs remain a barrier to scalable applications. Take Google's newly released Gemini 3.8 Flash: while its current promotional pricing stands at just $0.75 per million input tokens and $3.75 per million output tokens (rising to $1.50 and $7.50 respectively in 2027), far below Claude Opus 5's $5 and $25, the expense remains non-trivial for small and medium-sized enterprises making large-scale API calls. Insufficient high-end computational supply and high application costs remain significant barriers.

Next is the data quality dilemma. Gartner surveys show that only 4% of enterprises self-rate as having AI-ready data; most enterprises' data consists merely of traditional business records, lacking structured semantics and business scenario alignment. Data that is "visible but unusable" has become a major bottleneck for enterprise AI deployment.

A deeper issue lies in the cognitive limitations of models themselves. The Blue Book on the Technical System of Causal World Models released by Shuzhi Data points out that most current AI models remain at the associative layer of the "ladder of causation," excelling at answering "what is" but struggling to deduce "what happens after intervention." In zero-tolerance core industrial scenarios like oil and gas, power, and high-end manufacturing, AI outputs lacking causal logic and reasoning capabilities cannot meet rigorous decision-making requirements. This is the core reason why many AI projects are "usable but not trusted, deployed but not deeply integrated."

Gartner Senior Research Director Sun Xin succinctly summarizes the issue: while top AI models complete capability iterations every 1.5 to 3 months, enterprises' ability to adopt AI and translate it into business value has not grown at the same pace. A severe "impact gap" exists between technological supply and enterprise deployment needs.

04. Bubbles and Precipitation

In the first half of 2026, several AI applications once hyped by capital began exiting the market. OpenAI announced the discontinuation of Sora, its video generator, just six months after launch; Yupp.ai, an AI model evaluation platform that raised $33 million in funding, shut down; Google began scaling back its internal AI application lines.

These exits are not without value. Together, they expose a common issue: whether the application layer has formed sufficiently thick independent value as underlying models continue to upgrade. Applications that relied solely on model dividends for support are losing reasons to exist independently.

But the bursting of bubbles does not signify the end of the story. In fact, it resembles a necessary winnowing—clearing out applications that merely "white-label" large model capabilities with interface packaging, allowing truly valuable applications to surface.

The China Academy of Information and Communications Technology noted in its 2025 industry observation that continuous technological iteration has laid a solid foundation for large models' practical application, with intelligent agents emerging as the primary form of large model deployment. GPT-6 Astra has already demonstrated a tangible possibility—according to OpenAI's data, in OSWorld 2.0 offline evaluations, Astra achieved 72.6% performance in computer usage tasks, completing each task in about 40 minutes, a significant improvement over GPT-5.6 Sol's 65.7% and 75 minutes. When models can independently complete tasks like filling out online forms, updating CRM records, organizing schedules, analyzing scientific data, and creating websites, the application layer's form will inevitably be redefined.

The path from "intelligent assistants" to "digital employees," from single-point tools to autonomously collaborative intelligent agent groups, remains long—but the direction is becoming clear.

Models continue to grow stronger; this is a fact. Just four days had passed in September 2026 when four overseas giants took turns making announcements. However, the speed at which AI can truly deliver closed-loop solutions to a broader range of problems lags noticeably behind the growth of model capabilities. Narrowing this gap and translating technological advances into productivity gains may be the most anticipated proposition in AI's next phase.

- End -

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.