Kimi K3 Goes Viral and Suspends Service in 48 Hours: Model Wins Benchmarks, Compute Falls Short Against Big Players

10/10 2026 328

Author|Yinke Editor|Lin Feng


Kimi K3 has dragged China’s large model race into a new phase with its pricing strategy far exceeding its predecessor, a forced halt to new user subscriptions due to compute exhaustion, and a balance sheet showing a sevenfold valuation surge in six months: After shedding the low-cost label, can vendors strike a new balance between quality, service, ecosystem, and capital?

Kimi K3 has dominated headlines in recent days.

It first skyrocketed to the top of Arena’s frontend programming leaderboard with 1,679 points, leaping from 18th to 1st place, surpassing Claude Fable 5 and GPT-5.6 Sol. Then, Elon Musk commented “Impressive” under a review post and officially announced a day later that his Grok 4.6 would “possibly surpass Kimi,” to which Kimi responded, “Welcome to the 2-trillion club,” sparking a high-stakes, viral exchange.

Beyond hype, even pricing—the most criticism-prone aspect—was quickly absorbed by the market. K3’s output price stands at $15 per million tokens, nearly four times higher than K2.6, directly rivaling Anthropic’s Sonnet 5. Historically, price hikes for Chinese large models have been self-sabotaging, but K3 saw overseas paid users and API revenue surge by approximately 400% year-on-year after launch.

However, amid this success, backend servers finally buckled. On July 19, within 48 hours of K3’s launch, Moonshot AI’s GPU clusters reached full capacity under global user demand. The company was forced to suspend new user subscriptions to maintain service quality for existing users.

An AI company launching a flagship product and shutting its doors due to insufficient compute just three days later—this marks the first time in China’s large model race since 2023 that a model proved “too powerful to run.” More critically, this sudden sales halt exposed not just Moonshot AI’s vulnerabilities but systemic flaws across the entire large model industry.

01 The Atmosphere Demanded This Price

Frankly, K3’s API pricing is steep: $3 per million tokens for input and $15 for output, identical to Anthropic’s Sonnet 5 and nearly four times pricier than its predecessor, K2.6. Foreign experts called it “the most expensive model ever released by a Chinese AI lab.”

K3’s pricing breaks from the norm. Other Chinese large models—DeepSeek’s ultra-low-cost approach, Zhipu’s GLM for affordability, Baidu’s ERNIE series competing on price—have all embraced “cheap” as their label. K3 tore it off.

Why take this risk? The answer is money.

From its founding in 2023, Moonshot AI secured rocket-like capital accumulation: $200 million in angel funding, over $1 billion in Series A+, over $300 million in Series B, with Alibaba, Tencent, and Meituan all participating. By late 2025, a Series C round raised another $500 million, valuing the company at $4.3 billion. Founder Yang Zhilin declared, “With over $1 billion in cash on hand, there’s no rush to go public.”

But 2026 brought a dramatic shift. From January to June, Moonshot AI completed four funding rounds, raising over $3.9 billion cumulatively, with its valuation soaring from $4.3 billion to $31.5 billion—a sevenfold increase. Simultaneously, rumors swirled about its Hong Kong IPO, with a potential listing in six months.

Clearly, Moonshot AI aimed to inflate its valuation for a lucrative public debut. However, financing alone wasn’t enough—revenue growth had to keep pace.

And it did. Annual recurring revenue (ARR) surged past $100 million in March, doubled to $200 million by May, and hit $300 million in June—a triple increase in three months. API business accounted for over 70% of revenue. These figures represent elite growth in any SaaS sector, let alone the capital-intensive large model industry.

But behind this growth lies a financial conundrum. While ARR tripled, valuations septupled. Secondary markets focus on expectations; primary markets on narratives. At the IPO threshold, these must align. What ARR level justifies a $30 billion valuation? This is a financial puzzle.

K3’s pricing is the solution. The hike reflects both its perceived value and financial necessity.

02 Servers Overloaded: Kimi’s Bittersweet Dilemma

A humiliating incident occurred 48 hours after K3’s launch. Global user inflows maxed out Moonshot AI’s GPU clusters, forcing the company to halt new subscriptions.

A flagship model needing rapid adoption was sidelined within three days due to compute shortages—a hard landing against computational limits.

China Academy of Information and Communications Technology data shows a 417% surge in AI compute demand in Q1 2026, with supply growing just 128%. K3’s 2.8-trillion-parameter configuration with million-token context windows demands far greater compute per inference than its predecessor, widening the supply-demand gap post-launch. Stronger models require costlier compute—a cycle pressuring all players but fatal for startups.

Contrast this with tech giants. Alibaba planned at least ¥380 billion ($53 billion) in AI infrastructure investments from 2025–2027, with ¥38.6 billion ($5.4 billion) in Q1 2026 capital expenditures for Alibaba Cloud and AI. ByteDance earmarked ¥200 billion ($28 billion) for AI in 2026, including ¥85 billion ($12 billion) on chips. Tencent spent ¥79.2 billion ($11 billion) in 2025 and accelerated to ¥31.9 billion ($4.5 billion) in Q1 2026.

Meanwhile, Moonshot AI’s total financing stands at ¥40 billion ($5.6 billion)—a king’s ransom for a startup but barely a quarter of Alibaba’s quarterly AI spending.

Tech giants subsidize compute with e-commerce profits and gaming revenues; startups rely on financing and revenue. Moonshot AI, having completed six funding rounds, can’t depend on infinite capital. Revenue must follow. K3’s price hike signals a turning point for Chinese large model startups: when free and cheap models can no longer cover costs, only products good enough for markets to pay for will survive.

Yet Moonshot AI’s innovation remains trapped in underpowered servers due to financial constraints.

03 Alibaba Turns Pain Points into Products, But K3 Isn’t on the Menu

Two days after K3’s launch, on July 20, Alibaba Cloud unveiled “Agent-Optimized Instances,” lightweight servers for AI agents.

The logic is straightforward. Developers previously needed separate purchases for cloud servers, API keys, public IPs, and bandwidth—a fragmented, cost-unpredictable process. Alibaba Cloud bundled vCPUs, memory, cloud disks, 200Mbps peak bandwidth, and large model tokens (100 million to 3.2 billion) into prepaid packages. The entry-level tier offers 2 vCPUs, 2GB RAM, and 200 million monthly tokens for ¥262.5 ($37).

Alibaba Cloud estimates 18% monthly savings for AI programming and 68% for content generation. Notably, the instances run on Alibaba’s proprietary ANOLISA OS, claiming 30% less token waste, 30% faster agent execution, and 20% quicker cold starts in mainstream scenarios.

The product’s timing is impeccable. As model competition shifts from benchmarks to commercialization, infrastructure must adapt. Startups build powerful engines, but big tech provides the chassis, tires, and fuel tanks needed for road readiness—their core strength.

However, a critical gap remains: tokens are tied to Alibaba’s ecosystem, likely centered on Tongyi Qianwen, excluding K3.

This creates an awkward disconnect. In the U.S., AWS Bedrock integrates Claude, Google Cloud Vertex AI integrates Gemini, and Microsoft Azure integrates OpenAI—seamless cloud-model alignment. In China, Moonshot AI’s open-source model dominates headlines, while Alibaba Cloud leads in infrastructure, yet the two haven’t aligned on pricing tiers. Developers using K3 must still buy servers from Alibaba and APIs from Moonshot AI, incurring efficiency and cost penalties.

From this perspective , the industrial ecosystem needs refinement. However, K3’s full weights will release on July 27, enabling global developers to freely download and deploy it. Then, real third-party stress testing begins: Can hallucination rates be suppressed? Does inference speed keep pace? How well do multimodal capabilities perform in real-world scenarios? These factors will determine if K3 evolves from a “benchmark champion” to a “commercial leader.” For Moonshot AI, the challenge extends beyond technical iteration to coordinating compute, capital, and ecosystem layers—a task more critical than any benchmark run.

- END -

(Images and public data sources in this article are from online and industry platforms. For infringements, please contact us for prompt corrections.)

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.