2.8 Trillion Parameters Given Away: How Does Moonshot AI Make the Calculation?

07/28 2026 459

Competition Among Open-Source Large Models Is No Longer Just About Strength

Author|Xinjian

Editor|Xiaobai

Illustrations|AI Generated

Produced by|Qiangdiao Next

On the evening of July 27th, Moonshot AI duly uploaded the complete model weights of Kimi K3 to Hugging Face, GitHub, and ModelScope. With 2.8 trillion parameters, it is the world's first open-weight model to approach the 3 trillion threshold, freely available for anyone to download, deploy locally, and further develop.

Ten days earlier, Elon Musk had left a comment under K3's evaluation report saying, "Impressive," only to announce on social media shortly after that xAI's new 2 trillion parameter model would complete initial training the following week, "potentially surpassing Kimi." Kimi's official response was succinct: "Welcome, Musk, to the 2 trillion parameter club."

Then, it directly tore down the club's door: "Still comparing parameter sizes? Sorry, the largest one is now free."

This is clearly not just idealism. A review of Moonshot AI's moves over the past month reveals: In mid-June, ARR (Annual Recurring Revenue) surpassed $300 million, with API revenue accounting for over 70%. On July 19th, new consumer subscriptions were suspended. The previous funding round valued the company at $20 billion had just closed, while the pre-investment valuation for the next round had already reached $31.5 billion.

Eagle-eyed developers noticed that alongside the ~1.56TB model file on Hugging Face, a mere 3.07KB license revealed more commercial details.

Under the Kimi K3 License, if a company operates a "Model-as-a-Service" business and its consolidated revenue with affiliates exceeds $20 million in any 12-month period, it must sign a separate agreement with Moonshot AI before using K3 or its derivatives commercially.

This license shifts the conversation about K3 from technology to business.

01. Moonshot AI Refuses to Work for Free for Cloud Providers

Open models have long faced an awkward reality: model companies bear training costs, while cloud providers may reap subsequent revenues. This is the fundamental reason Baidu's Robin Li earlier dismissed the prospects of open-source models.

Under the MIT license, cloud providers can download models, optimize inference, package them as APIs, and charge using their existing customer base, computing power, and sales channels. Model companies gain downloads and industry buzz, but customer relationships, usage data, and billing remain on cloud platforms. The faster models iterate, the more labs risk becoming upstream R&D departments for cloud providers.

In K3's license, "Model-as-a-Service" refers to allowing third parties substantial control over inputs, parameters, or training data. Embedding K3's capabilities into a specific feature of one's product or simply forwarding requests to another platform does not qualify.

In other words, Moonshot AI does not intend to block developers and most enterprise users; it targets cloud providers and inference platforms that download weights, package them as APIs, and resell inference services under their own brands and pricing.

Overseas analysts specializing in open-source infrastructure have drawn parallels between this threshold and events a decade ago.

MongoDB and Confluent initially scaled through permissive licenses. Only when Amazon and other cloud providers packaged their open-source databases into paid cloud services without sharing revenue did they switch to restrictive SSPL and BSL licenses to prevent resale.

First, scale through openness; then, regain bargaining power through rules. From hyperscale cloud providers to inference platforms like Fireworks and Together focused on model hosting, all theoretically fall under the $20 million revenue threshold. K3 faces commercial challenges similar to MongoDB and Confluent but chooses a gentler legal path: not directly banning MaaS resale but requiring separate agreements for operators above the revenue threshold.

The open-model industry has evolved from "who open-sources first" and "who open-sources most thoroughly" to "who can reclaim pricing power from the open ecosystem."

Similarly releasing weights in July, Zhipu's GLM-5.2 and DeepSeek's V4 Pro used clean MIT licenses without resale restrictions. Of course, this may not stem from greater generosity.

As for Moonshot AI's "certified inference partners" exempt from this threshold, neither the standards nor the list is currently clear, nor is it defined in the license text.

02. From "Parameter Race" to "Scenario Race"

The license is merely K3's commercial choice. Returning to the model itself, K3, DeepSeek, GLM, and Qwen have embarked on four distinct paths.

K3 pursues "comprehensive strength." With 2.8 trillion total parameters, 1 million token context, and native visual understanding, official demos show K3 writing a GPU compiler from scratch, participating in chip design, and generating 3D open-world games. However, at $3 per million input tokens and $15 per million output tokens, with an average task cost of ~$0.94, it is the most expensive among the four.

DeepSeek V4 Pro takes the opposite approach. With 1.6 trillion total parameters but only ~49 billion activated, its input and output prices are $0.435 and $0.87 per million tokens, respectively, with an average task cost of ~$0.04.

DeepSeek also introduces peak/off-peak pricing, doubling charges from 9 AM-12 PM and 2 PM-6 PM, while off-peak usage saves 50%. This aligns with its consistent strategy: parameters can be large, but activation scale and usage costs must be minimized.

Zhipu's GLM-5.2 takes a third path. Using an upgraded DeepSeek Sparse Attention, it reduces inference overhead for 1 million token contexts to a range manageable by commodity hardware. Combined with ~40 billion single-token activated parameters, it achieves the shortest first-character latency among the three.

GLM-5.2's input and output prices are $1.4 and $4.4 per million tokens, respectively, falling between K3 and DeepSeek, with a focus on response speed in coding and agent scenarios.

Alibaba's Qwen does not pursue single-point excellence but breadth. Supporting hundreds of languages and more complete multimodal capabilities, it resembles a general-purpose foundation for enterprise clients rather than a benchmark-chasing flagship product.

However, Kimi has also made Alibaba restless. Three days after K3's release, Alibaba teased Qwen3.8-Max, with 2.4 trillion parameters and explicitly stated plans to open weights. For nine months over the past year, Kimi set the ceiling for open-source model scale. Now, flagship-grade open weights are becoming standard for Chinese large model vendors.

The four paths can no longer be ranked by a single dimension. The "parameter race" has shifted to a "scenario race": long-text processing favors K3, cost-effectiveness favors DeepSeek, response speed favors GLM, and breadth favors Qwen.

Differentiation does not guarantee security. DeepSeek, GLM, and Qwen may continue to strengthen their strengths at any time, and K3's price disadvantage will not disappear automatically.

03. Closed-Source Model Premiums Are Being Squeezed

What truly makes this round of open models striking is the comprehensive comparison of benchmarks and pricing.

On Artificial Analysis's Intelligence Index leaderboard, K3 scored ~57. Ahead of it are Claude Fable 5 and GPT-5.6 Sol, with scores of ~60 and ~59, respectively; Claude Opus 4.8 follows closely at ~56.

This marks the first time an open-weight model has performed within 2-3 points of top-tier closed-source models. On the Arena frontend code leaderboard, K3 surged to first place within 24 hours of release, overtaking Claude Fable 5 and improving 17 ranks from the previous generation.

More critically, these near-world-class results come at significantly lower prices than most closed-source flagships.

K3's average task cost of ~$0.94 is about half that of Claude Opus 4.8. If closed-source models justify their premiums on "only we can achieve this level," then when an open-weight model narrows the gap to 2-3 points at half the price, the rationale for such premiums quickly shrinks.

Microsoft CEO Nadella calls this the "reverse information paradox": buying closed-source models means paying twice—once for the bill, and once by surrendering business know-how to the model.

In January, a New York court ordered OpenAI to hand over 20 million de-identified user conversations, serving as a wake-up call for all enterprises: data on others' servers is no longer under your control. K3 offers the "data stays within your walls" option.

04. Weights Are Just the Appetizer; the Real Battle Is at the Inference Layer

Open weights are no longer novel in this camp. DeepSeek, GLM, and Qwen's flagship models have long been available on Hugging Face, with download, commercial use, and fine-tuning all permitted.

What K3 brings this time that feels truly fresh lies beyond the weights.

Because KDA alters traditional prefix caching mechanisms, Moonshot AI contributed its cache implementation directly to the vLLM project. On the day weights were released, vLLM and SGLang, two major inference engines, synchronization provided deployment solutions. AMD's team completed first-day adaptation for the MI355X chip, cloud platform Modal launched K3 hosting services, and Inferact trained and open-sourced a DSpark speculative decoder compatible with K3.

vLLM's minimal deployment configuration requires 8 NVIDIA B300s or 8 AMD MI355Xs—both top-tier data center chips not yet widely available, far beyond individual developers' reach.

This suggests Moonshot AI seeks more than just headlines for "a larger model." It aims to secure a position within inference frameworks and deployment toolchains.

Anyone can open weights. But if developers use service schemes, cache implementations, and speculative decoders built around your architecture, switching models becomes more than just "replacing a file."

The license's design—"exempting resale terms when used through certified inference partners"—gives Moonshot AI a switch: it can select who enters this ecosystem and who must negotiate first.

Releasing weights is a one-time act; embedding into toolchains offers a chance to truly retain developers. This is why K3's weight release is better understood as a beginning, not an endpoint.

However, custom rules come with costs. Enterprises choosing open models value not just capability and cost but also license simplicity and stability. MIT and Apache 2.0 licenses are popular precisely because legal departments can easily assess usage boundaries.

K3 adds conditions such as revenue thresholds, affiliate definitions, MaaS criteria, and certified partners, requiring more compliance review before deployment. If GLM, DeepSeek, or Qwen offer comparable capabilities with simpler licenses, platforms may bypass K3.

Open models inherently have low switching costs. Technical leads from a single release rarely sustain a custom ruleset long-term.

Rumors suggest K3.1 is already en route, potentially releasing as soon as next month. The "open-source—validate—iterate" playbook may only just be beginning to prove itself.

- END -

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.