08/25 2026
331
Today, NVIDIA has, for the first time, unveiled the preliminary on-chip test results for its next-generation flagship cabinet, the Vera Rubin NVL72.

Notably, NVIDIA's testing eschewed the conventional use of AI large language models (LLMs) for chat-based Q&A, opting instead for AI agents. They directly harnessed the power of the most robust open-source large model, DeepSeek V4 Pro, to conduct benchmarks on SemiAnalysis's AgentX, which comprises entirely of conversational trajectories generated by agents writing code in real-world scenarios.
Why opt for agent benchmarking over LLM inference?
According to the latest data from OpenRouter, in real-world applications, an AI agent task consumes 15 times more tokens than an ordinary chat conversation!
Upon task completion, an agent's context can swell from 60,000 tokens to nearly 400,000. Such expansive contexts are commonplace when agents undertake lengthy tasks, interspersed with tool invocations and requests to sub-agents.
The more interactions an agent engages in, the higher the throughput. Hence, NVIDIA has made it abundantly clear that performance metrics must evolve! It's no longer adequate to measure single inference requests; capturing the entire agent workflow is imperative. For this reason, they employed the AgentX benchmark.

The results are nothing short of remarkable: Aiming for 160 tokens per second interaction, the Vera Rubin NVL72 achieved up to a 30-fold increase in throughput per megawatt and up to a 35-fold reduction in cost per million tokens compared to its predecessor, the GB300 NVL72!
NVIDIA has proclaimed the end of the LLM era and the dawn of the Agent era, rendering all previous AI benchmarking standards obsolete!

Compared to the older H200, the GB300 generation already boasted a 15-fold increase in throughput per megawatt and slashed the cost per million tokens to roughly one-tenth on DeepSeek V4 Pro. Building upon this foundation, the Vera Rubin achieved an additional 30-fold improvement. When compounding the gains across two generations, from the H200 to the Vera Rubin, the volume of agent work achievable with the same unit of electricity has surged nearly 450-fold!
Industry pricing trends indicate that the sticker price for a full cabinet roughly doubles with each generation, making it about four times pricier over two generations. However, with the same unit of electricity now producing over 400 times more tokens, the corresponding cost per million tokens has plummeted by dozens of times.
Despite the higher machine costs, the unit cost of utilizing them has drastically decreased.
NVIDIA's most audacious move in the past two years has been to overhaul the entire cabinet design, sell it at a significantly higher price point, and yet propel the output per unit of electricity to unprecedented levels, far surpassing those of older machines.