07/20 2026
544
© Chaoyong AI Editorial Team
Early this morning, the biggest AI event on the Chinese internet: Kimi released K3.
Moonshot AI openly admitted in its announcement: 'The overall performance of Kimi K3 still lags behind the strongest closed-source models, Claude Fable 5 and GPT-5.6 Sol.'
Pretty straightforward, no pretense.
But the second thing they did: raised prices.
The API output price for K3 has increased to 100 RMB per million tokens, over 3.5 times higher than the previous flagship K2.6's output price of 27 RMB per million tokens. The input price has also risen to 20 RMB per million tokens, more than triple K2.6's input price (cache miss) of 6.5 RMB per million tokens.
So, what makes K3 so impressive? Is it worth using for ordinary users? And is this price hike justified? The Chaoyong AI team uses Kimi daily for work, so today we're sharing our real-world testing experience. 
What Makes K3 So Impressive?
Let's start with the hard data.
K3 has 2.8 trillion parameters, using a Mixture of Experts (MoE) architecture with 896 experts, 16 of which are activated at a time. Its context window reaches 1 million tokens, with native support for visual understanding.
But large parameters don't automatically mean superiority. K3's real strength lies in its global leadership in frontend development.
Frontend Code Arena, a platform where users anonymously vote on which model generates better web pages, gave K3 a score of 1679, surpassing Claude Fable 5 (1631) and GPT-5.6 Sol (1618). This marks the first time an open-source model has outperformed top closed-source models in this dimension. 
Note: This is in the field of frontend development (WebDev), not overall programming capability. Moreover, the current ranking is marked as Preliminary, with K3's vote count (1757) still lower than Fable 5's (2505). As more votes come in, the ranking may change.
In seven frontend subcategories, K3 claimed six first-place finishes.
In other words, if you're involved in frontend development or web design, K3 is currently the best model available.
Moonshot AI also showcased several cases: using K3 to write a GPU programming system, MiniTriton, from scratch; completing chip design autonomously in 48 hours; and editing a video composed of 56 segments.
These aren't just 'answering a question' but 'completing a project'—letting the AI work independently, make revisions, and see the project through to completion. 
Prices Have Risen, But Programming Scenarios Might Be Cheaper Than You Think
K3's listed prices: 20 RMB per million tokens for input (cache miss), 2 RMB per million tokens (cache hit), and 100 RMB per million tokens for output.
For comparison: Claude Fable 5 costs $10 for input and $50 for output; GPT-5.6 Sol costs $5/$30, Terra $2.5/$15, and Luna $1/$6. K3's listed prices are roughly one-third of Fable 5's.
But here's the key detail: the Mooncake architecture.
Moonshot AI disclosed that in programming scenarios, the cache hit rate exceeds 90%. This means that when you repeatedly submit the same codebase, the actual input cost drops significantly.
At a 90% cache hit rate, the actual blended input cost is approximately 3.8 RMB per million tokens (about $0.53), with a minimum of 2 RMB per million tokens (about $0.28) when all inputs hit the cache.
DeepSeek-level pricing with global-leading frontend development capabilities.
This is why K3 dares to raise prices: in the core programming scenario, its actual cost is likely much lower than the listed price. 
Chaoyong AI's Real-World Testing: K3's Impact on Our Workflow
Our Chaoyong AI team uses Kimi for several tasks: writing assistance, long document processing, code support, and multimodal analysis.
In these scenarios, some things have changed with K3, while others haven't. 
Screenshot from Kimi's official website
(1) Long Document Processing: The 1 Million Token Context Window Is a Game-Changer
When writing industry analysis articles, we often handle hundreds of thousands of words in research materials, financial reports, and interview transcripts. Previous Kimi models could handle 200,000 words, but 1 million tokens mean we can now feed entire books, codebases, or video scripts at once.
No more splitting or segmenting—just one feed.
This is a qualitative leap in efficiency.
(2) Multimodal: Not Just an 'Added' Visual Capability
K3 natively supports visual understanding, unlike text models with a separate visual module bolted on. This means deeper comprehension when analyzing images, PDF scans, or design drafts.
We tested a scenario: uploading a complex industry chain diagram and asking K3 to analyze company relationships and competitive dynamics. Previous models missed some key nodes, while K3's recognition accuracy improved significantly.
(3) Programming Assistance: From 'Writing Code' to 'Completing Projects'
Our technical team members use Kimi for scripting and data processing. With K3, the change isn't just 'faster code writing' but 'completing more complex tasks autonomously.'
For example, when asked to write a data collection script, previous models would stop after generating the code. K3 checks if the code runs, identifies potential errors, and automatically fixes them.
This is a leap from 'code generation' to 'project delivery.' 
K3's Weaknesses: Cannot Be Ignored
(1) Still Third in Overall Score
According to the Artificial Analysis Intelligence Index, K3 scored 57 points. The top performer, Claude Fable 5, scored 60, while GPT-5.6 Sol scored 59. This 3-point gap is noticeable in general scenarios. 
While K3 leads globally in frontend development, its overall capability still lags behind top closed-source models.
If your primary needs are chatting, writing, or creative generation, K3's 'global first' label doesn't apply.
(2) Model Weights Not Yet Released, API-Only Access Now
Moonshot AI promised to release the full model weights by July 27. But as of July 17, they're still not available.
This means: currently, K3 can only be accessed via API. Data privacy, service stability, and customization capabilities are limited by Moonshot AI's infrastructure.
Want local deployment, fine-tuning, or model modifications? Wait 10 days.
(3) Low Cache Rates in Non-Programming Scenarios, Actual Prices Aren't Low
The 90% cache rate applies to programming scenarios. For tasks like copywriting, customer service, or general Q&A, the cache hit rate drops significantly. Actual costs may approach the listed prices—around $3 for input and $15 for output.
At these prices, K3 offers lower cost-effectiveness for non-programming users compared to DeepSeek V4 Pro (about $0.435/$0.87).
(4) Local Deployment Barriers
A 2.8 trillion-parameter MoE model, even with only partial experts activated at a time, requires immense computational resources for local deployment. After the weights are released on July 27, the gap between 'wanting to use open-source' and 'being able to run it' remains vast for ordinary users. 
Should Ordinary Users Use K3?
If you're a frontend developer/programmer: Yes.
Global leadership in Frontend Code Arena, programming costs as low as $0.28 per million tokens (full cache hits), and upcoming weight release on July 27. No competitor matches this combination.
If you handle large volumes of long documents (legal, financial, academic): Worth trying.
The 1 million token context window is a hard advantage. But note: K3's overall capability still lags behind top closed-source models, so cross-verify critical information.
If you're a general user (chatting, writing, general Q&A): Maybe wait.
K3 offers lower cost-effectiveness than DeepSeek in this scenario and inferior general capabilities compared to Fable 5. Unless you specifically need a K3 feature, consider waiting.
If you require data privacy/localization: Wait until July 27. After the weight release, K3 will be the largest-parameter, strongest-frontend-capability open-source model among domestic options. However, local deployment demands extremely high GPU resources, making it impractical for most ordinary users.
K3 isn't the 'Chinese version of Fable 5' or an 'all-around champion.' It's the 'frontend development single-event champion + cost-effectiveness king.'
Moonshot AI admits it's inferior to top models but achieves global firsts in specific scenarios at one-third the price of Fable 5, with weights set to open.
For frontend developers, K3 is currently the best choice.
For non-coders, K3 is a 'worth watching' option but not a 'must-switch' one.
After the July 27 weight release, if the open-source community rapidly adopts it, K3's ecosystem advantages could truly explode.