08/13 2026
529
On Wednesday evening, DeepSeek suddenly rolled out the official version of its highly anticipated V4 Pro large language model.
The first to be updated was the DeepSeek API documentation, which revealed that the new model is named DeepSeek-V4-Pro-0813. Moreover, the official API pricing remains unchanged for now.

On Terminal Bench, DeepSeek V4 Pro scored 87.9, just 0.1 points below Claude Fable 5’s 88.0.

Nearly all discussions have centered on this narrow 0.1-point margin. However, two other evaluation tables have gone largely unnoticed.
According to leaked comparative data, in the AI security agent evaluation CyberGym and the high-difficulty agent evaluation AutomationBench, the official version of V4 Pro outperformed Fable 5. Terminal Bench measures terminal operation capabilities, CyberGym assesses autonomous decision-making in security attack and defense scenarios, and AutomationBench targets more complex automated task orchestration. The difficulty of these three evaluations increases progressively, with the latter two representing the most challenging aspects for large models in real-world deployment.
The preview version of V4 Pro still trailed Fable 5 in these two areas. However, after the official release, the tables have turned.
This is where the real focus should lie. While general benchmarks have been narrowing the gap over the past two years, with differences shrinking to 0.1 points, the conversation has begun to lose sight of what truly matters. What genuinely distinguishes models is their Agent capability—the ability to consistently make correct decisions in multi-step, long-chain tasks. Such capabilities cannot be achieved through mere training exercises.
The official version of V4 Pro supports both a 1M context window and a maximum output of 384K. While a 1M context is not uncommon, a 384K output is currently the most ambitious among publicly available APIs. A large output window enables the model to complete longer code generation, more complete reasoning chains, and more complex tool invocation orchestration in a single inference. The core capabilities tested by AutomationBench align precisely with this.
By extending the output window to 384K, DeepSeek is addressing engineering bottlenecks in Agent scenarios. This move is made at the deployment architecture level.
Behind this advancement lies the support of China’s domestic computing power. Currently, domestic super-nodes are entering a phase of large-scale deployment. A super-node refers to a logically unified computing unit formed by connecting multiple AI accelerator cards via high-speed networks, breaking through the limitations of single-card memory and bandwidth to support the training and inference of trillion-parameter-scale models. The memory and bandwidth requirements for a 1M context plus 384K output far exceed the capacity of a single card, making the super-node architecture essential for running this specification.
The leap in capabilities from the preview to the official version of V4 Pro occurred within a timeframe of just a few weeks. Model algorithm iterations typically take months, making it difficult to explain the jump in Agent evaluation scores—from lagging to surpassing—within a few weeks solely through algorithmic optimization. A more plausible inference is that the underlying inference infrastructure has been expanded. The pace of super-node deployment is directly determining the pace of model capability release.

According to data from China Insights Consultancy, the market size of China’s intelligent computing chips is expected to grow from USD 30.1 billion in 2024 to USD 201.2 billion in 2029, with a compound annual growth rate of 46.3%. The path of demand transmission is becoming clearer.
Guojin Securities points out that as the supply of domestic chips increases, demand will sequentially transmit to switching networks, servers and entire cabinets, computing power leasing, AIDC deployment, and cloud services. Chip supply breakthroughs, system architecture upgrades, and demand expansion are occurring simultaneously.
Huatai Securities notes that domestic GPUs have entered the era of ten thousand cards, transitioning from "usable" to "user-friendly." The gap between the two lies precisely in system-level engineering capabilities like super-nodes.
The competitive focus of domestic computing power has shifted from single-card performance to system-level efficiency. Whoever first successfully implements super-nodes will be able to push model specifications beyond the reach of competitors.
The API pricing for the official version of V4 Pro remains unchanged. The cost is 0.025 yuan per million tokens for cache hits, 3 yuan for cache misses (input), and 6 yuan for output—the same as the preview version. Despite the significant leap in model capabilities, the price stays the same.
This approach does not resemble conventional commercial pricing strategies. DeepSeek is compatible with both OpenAI and Anthropic access formats and has opened up Beta-stage dialogue prefix continuation and FIM completion functions. Each move lowers the migration threshold for developers, widening the ecological interface.
With increased capabilities, unchanged prices, and lowered access thresholds, all these actions point toward a single goal: capturing Token call volume.
Guosheng Securities has shifted the narrative of the hard technology sector from expected storytelling to performance delivery in its latest research report. For DeepSeek, the indicator of this shift lies in the slope of the API call volume curve. Sustained growth in Token call volume is the anchor point for the return on investment in super-nodes.
As of press time, DeepSeek has not yet released the complete technical report for the official version of V4 Pro. The specific scores for CyberGym and AutomationBench have also not been officially confirmed. However, one trend is already clear: Agent evaluation is replacing general benchmarks as the true measure of generational gaps between models.
Super-nodes support larger context and output windows, which in turn support more complex Agent tasks. These complex Agent tasks are precisely the aspects that general benchmarks fail to measure but are most needed in real-world deployments. This chain of transmission may be more worthy of attention than the 0.1-point difference.