07/21 2026
467
These days, it seems that Chinese large-scale models are gathering for a grand convention. First, K3 made headlines, and then yesterday, Qwen3.8 was unveiled—both boasting large parameter models exceeding 2 trillion. It's rumored that DeepSeek V4 will also be launched today. Initially, I planned to discuss it after V4's release, but I couldn't hold back my curiosity.
Nevertheless, Qwen's recent moves already present numerous fascinating aspects worthy of exploration.
In terms of timing, Kimi K3 has just been launched and has performed exceptionally well. DeepSeek V4 is on the verge of release, and while its capabilities remain unknown, expectations are high, given Liang Sheng's track record.
Caught between these two, Qwen3.8 finds itself in a precarious position. If its model capabilities cannot surpass K3 and its cost-effectiveness cannot outdo V4, external promotion will be challenging.
Another noteworthy aspect is that Alibaba explicitly stated its intention to open-source Qwen3.8 upon release. (The author of the WeChat Official Account post even described the Qwen large model as "poised for open-sourcing.")

One contributing factor to Lin Junyang's departure was the group's adjusted stance on open-sourcing, beginning to weigh the cost-benefit ratio more carefully.
Alibaba had indeed shown signs of this shift earlier, transitioning from "actively embracing open-sourcing" to "conditionally embracing open-sourcing" and "selective open-sourcing." The open-source commitment for Qwen3.8 Max signals Alibaba's return to its main strategy of the past three years.
The Art of Translation and Wording
Interestingly, there's a subtle difference between the introductions published on Qwen's official Chinese and English accounts. The English version directly claims it is "second only to Fable 5," implying it is the second most powerful model after Fable 5.

However, the Chinese version adds the word "possibly," stating, "it may be the most powerful model aside from Fable 5." This raises questions—does the official team lack confidence in Qwen3.8's capabilities, or are they concerned about domestic overhyping leading to a backlash?
Qwen Still Shows Hesitation.
Among domestic large-scale model players, Alibaba tends to overhype its offerings. This is related to its management style; Alibaba is always the most optimistic during earnings calls.
The most typical example is the previous "Happy Horse," which dominated benchmarks upon release but later fell short of expectations after users evaluated it more calmly. The initial overperformance negatively impacted perceptions to some extent.
Of course, the outcome was positive, as Alibaba's leadership recognized it. Zhang Di, the person in charge, now oversees the group's multimodal vision business line, with Wanxiang, Happy Horse, and Happy Oyster all merged into the Future Life Lab.
Perhaps having learned from past mistakes, Alibaba didn't fully hype the base model this time.
Currently, adding the qualifier "possibly" seems artistically and reasonably done.
Most initial community tests conclude that Qwen3.8's current performance lags behind the early-released K3. If the Chinese promotion hadn't included "possibly," the online mockery would likely be overwhelming now.
However, the official team emphasizes that Qwen3.8 is continuously evolving daily, meaning post-training is still ongoing.
As we all know, the potential of post-training remains largely unknown. Referring to Tencent's HY3's improvement from its preview to the official version, Qwen3.8 might similarly turn the tables in its official release.
Moreover, even if Qwen3.8 doesn't immediately outperform K3 in user experience, Alibaba still benefits.
Alibaba holds over 30% of Yuezhi's equity, and the computing power for training K3 mainly comes from Alibaba Cloud. Thus, Alibaba gains regardless: if you use K3, Alibaba Cloud earns rental fees, and its equity appreciates; if you use Qwen, Alibaba secures its open-source and cloud ecosystem positions.
This balanced capital and computing power layout is enviable for other standalone startups.
Advantages in computing power and infrastructure directly reflect confidence in commercial pricing.
Kimi announced yesterday that it had to suspend membership openings due to computing power constraints. This exposes the harshest reality of large-scale model competition: models inherently involve trade-offs between cost and intelligence. Unicorns lacking cloud infrastructure instantly struggle with sudden traffic surges.
Meanwhile, Alibaba's Qwen3.8, though limited in capability, is fast and cheap. As Kimi struggles, Qwen3.8 is offering limited-time 90% discounts and even nighttime double discounts akin to traditional electricity peak-valley pricing.
But why did Alibaba choose to release Qwen3.8 at this node, knowing much post-training remains?
The already-released K3 and the upcoming DeepSeek V4 official version are key factors. If Qwen3.8 delays further, uncertainties become too great.
Alibaba internally likely has a clear assessment of its model's true capabilities. Management isn't foolish—with industry benchmarking now common, they won't be completely misled by team-presented data, as happened with Zuckerberg's LLaMA4 release.
After K3's release, Alibaba likely tested Qwen3.8 and found it didn't hold an absolute advantage.
On the other hand, DeepSeek is known as a "price killer." Liang Sheng's reputation for extreme engineering optimization and low inference costs hangs over everyone.
This created a unique time window.
If Alibaba stuck to its schedule and waited for Qwen3.8's full training before releasing it after DeepSeek V4, the situation would become passive.
Even if V4's model capabilities are average, as long as DeepSeek maintains its cost-cutting approach, slashing prices, Alibaba would lose on both fronts: in terms of absolute model capability, it can't surpass K3, which already dominates public opinion; in terms of cost-effectiveness, it can't beat DeepSeek V4.
Even if Qwen3.8 is a solid model overall, falling into a "neither capable nor cost-effective" position would leave no selling points for promotion, creating an extremely awkward situation.
Qwen's Max Version Returns to Open-Source, Lin Junyang's Departure "In Vain"
Lin Junyang's recent departure from Alibaba sent shockwaves through the large-scale model community. Many rumors circulated about his reasons, one being severe internal disagreements and route struggles over open-source versus closed-source strategies within Alibaba.
Lin Junyang represented Qwen team's early technical idealism (possibly also seeking personal achievements). Under his leadership, Qwen advanced rapidly through unreserved open-sourcing, earning a high reputation in the global developer community.
However, this purely charitable open-source approach inevitably clashed with the group's financial realities as model parameters surged to hundreds of billions or even trillions. Training a top-tier large-scale model requires astronomical computing power and funding, making ROI calculations unavoidable for group leadership.
LatePost reported that Alibaba was reluctant to open-source Qwen Max, but Lin Junyang pushed for it.
Open-source community reputation has benefits, but they're intangible.
In the eyes of some group leaders, open-sourcing the top flagship model without reservation not only fails to directly generate API revenue but also equates to spending real money to build weapons for the entire industry, even competitors, to use against Alibaba Cloud's clients.
After Lin Junyang's departure, Alibaba's large-scale model route noticeably contracted and shifted.
Open-sourcing flagship models halted, adopting a utilitarian approach of "open-sourcing small and medium models, keeping large models closed-source."
Alibaba's reluctance to open-source its largest flagship model was once seen as a signal of its gradual compromise toward closed-source commercialization and the decline of open-source idealism.
For example, Qwen3.6's release page stated, "We will also open-source smaller model versions to reaffirm our firm commitment to technological inclusivity and community-driven innovation."
Why did Alibaba explicitly announce open-sourcing the 2.4T-parameter Max top version with Qwen3.8's release this time? The logical shift from tightening to fully reopening isn't hard to understand—it's a correction back to common sense.
Frankly, in the current large-scale model circle, the situation is: if you don't open-source, plenty of others will.
At this stage, open-sourcing essentially uses ecosystem and reputation to compensate for gaps in absolute model capability.
Imagine if your model truly competes with OpenAI or Anthropic, alternating in leadership. Even if closed-source, people would line up to use your API for the best results.
But currently, open-source models still lag behind closed-source ones, never surpassing the top contemporary closed-source models.
In the future, if independent large-scale model vendors like Zhipu or Yuezhi's Moonshot truly catch up to OpenAI or Anthropic, they'll likely shift to closed-source.
Otherwise, sustainable independence becomes difficult, unless we ignore AGI—you can't rely solely on investors as customers.
Alibaba's situation differs. I've never fully understood why Alibaba internally thought "shifting to closed-source is more cost-effective."
As a cloud vendor with resource advantages and a self-trained base model, open-sourcing costs Alibaba little substantive interest.
Even if Alibaba open-sources full weights, others deploying them independently or small cloud vendors offering services are unlikely to match Alibaba Cloud's native deployment in inference efficiency and per-token cost.
If they do, it only proves Alibaba Cloud is incompetent.
Thus, Lin Junyang's departure over route disagreements was regrettable, but fortunately, after some wavering, Alibaba finally calculated the numbers correctly.
Qwen3.8 Max's renewed open-source commitment and return to the past three years' main route indicate that Alibaba, as a cloud giant, has finally grasped its core narrative:
Rather than scrimping on closed-source API profits like independent model vendors, it's better to openly make friends with open-source models and earn big money from computing efficiency as the "utilities" of the AI era.
Last month in Paris, Joe Tsai said AI's enormous value is certain, but which layer it ultimately concentrates in is uncertain—bottom-layer chips, cloud computing infrastructure, model layer, or application layer. All four layers are possible. Alibaba's approach is full-stack, meaning strong positioning at every layer.
This analysis holds industry-wide, but probabilistically and in terms of revenue trends, Alibaba's greatest opportunity lies in its MaaS-based cloud computing business.