NVIDIA Unveils Vera Rubin NVL72 With 10x More Tokens per Megawatt for AI Factories
Updated
Updated · InfoWorld · Aug 14
NVIDIA Unveils Vera Rubin NVL72 With 10x More Tokens per Megawatt for AI Factories
3 articles · Updated · InfoWorld · Aug 14
Summary
NVIDIA framed data centers as “AI factories” where compute drives revenue, arguing operators should optimize tokens per watt, cost per token, uptime and time to first token rather than raw hardware specs.
Vera Rubin NVL72 is the centerpiece: NVIDIA said it delivers 10x more tokens per megawatt than GB200 NVL72, while its cableless rack design cuts tray assembly from 2 hours to 5 minutes.
That pitch extends across the stack, with Vera CPU offering 2x higher single-threaded performance and 40% lower memory latency, and sixth-generation NVLink claiming 3x lower latency than off-the-shelf Ethernet.
NVIDIA also highlighted Spectrum-X at up to 1.6x higher performance across deployments above 100,000 GPUs, BlueField-4 at 800Gb/s, and storage processors that lift agentic-inference tokens per second by up to 5x.
The broader message is that extreme co-design—compute, networking, storage, software and security—can lower token costs; NVIDIA said Blackwell optimizations cut DeepSeek V4 token costs by up to 5x in one month.
As NVIDIA redefines data centers into AI factories, will their push for extreme co-design secretly lock enterprises into a proprietary hardware monopoly?
If raw hardware speed no longer dictates AI profitability, how will older data centers survive the shift toward power-constrained, token-optimized ecosystems?
With agentic AI demanding sequential CPU-GPU loops, are traditional cloud infrastructures doomed to waste millions on idle accelerators and hidden latency bottlenecks?