LIVE
NVIDIA's Vera Rubin NVL72 Debuts in MLPerf Inference v6.116/09/26 · NVIDIA|OpenAI expands ChatGPT advertising with Sponsored Agents16/09/26 · OpenAI|OpenAI moves into advertising with 'Sponsored Agents'16/09/26 · OpenAI|Google DeepMind Introduces Gemini 3.8 Live and Its Extended Thinking Variant15/09/26 · Google DeepMind|What's at stake in AI's trillion-dollar infrastructure bet15/09/26|A Flaw in Chain-of-Thought Safety Monitoring14/09/26|Stellar Colosseum: A Multi-Agent System for Long-Horizon Mathematical Research14/09/26|Apple Code Hints Siri Could Be Swapped for ChatGPT or Claude14/09/26 · Apple|Anthropic says Houthi-linked actors used Claude Code for missile guidance software13/09/26 · Anthropic|Yoshua Bengio examines why AI agents lie, cheat and coordinate13/09/26|Terence Tao warns of a 'severe misalignment' in OpenAI's mathematical claims11/09/26 · OpenAI|Terence Tao warns of a 'severe misalignment' of AI in mathematics11/09/26|NVIDIA's Vera Rubin NVL72 Debuts in MLPerf Inference v6.116/09/26 · NVIDIA|OpenAI expands ChatGPT advertising with Sponsored Agents16/09/26 · OpenAI|OpenAI moves into advertising with 'Sponsored Agents'16/09/26 · OpenAI|Google DeepMind Introduces Gemini 3.8 Live and Its Extended Thinking Variant15/09/26 · Google DeepMind|What's at stake in AI's trillion-dollar infrastructure bet15/09/26|A Flaw in Chain-of-Thought Safety Monitoring14/09/26|Stellar Colosseum: A Multi-Agent System for Long-Horizon Mathematical Research14/09/26|Apple Code Hints Siri Could Be Swapped for ChatGPT or Claude14/09/26 · Apple|Anthropic says Houthi-linked actors used Claude Code for missile guidance software13/09/26 · Anthropic|Yoshua Bengio examines why AI agents lie, cheat and coordinate13/09/26|Terence Tao warns of a 'severe misalignment' in OpenAI's mathematical claims11/09/26 · OpenAI|Terence Tao warns of a 'severe misalignment' of AI in mathematics11/09/26|
ReleaseNVIDIA

NVIDIA's Vera Rubin NVL72 Debuts in MLPerf Inference v6.1

NVIDIA has published the first MLPerf Inference v6.1 results for its upcoming Vera Rubin NVL72 platform, claiming notable throughput and efficiency gains over the current Blackwell generation.

September 16, 20263 min readPublished byNVIDIA

NVIDIA used the release of MLPerf Inference v6.1, the industry benchmark suite overseen by MLCommons, to unveil the first performance figures for its upcoming Vera Rubin NVL72 platform, positioned as the successor to the current Blackwell generation. The company frames its pitch around three factors it considers central to inference economics at scale: raw system throughput, the ability to scale performance proportionally as hardware is added, and ongoing software optimization to extract more value from existing infrastructure.

NVIDIA argues these factors combine to produce more tokens generated per unit of time, which in turn can translate into higher revenue for operators billing customers for inference workloads. The NVL72 architecture, which links 72 GPUs within a shared memory domain via NVLink, is specifically designed to reduce the bottlenecks that arise when very large models must be distributed across multiple chips.

The MLPerf debut comes ahead of Vera Rubin's actual commercial availability, underscoring NVIDIA's pattern of publicizing performance claims for future chip generations well in advance. This is happening against a backdrop of intensifying competition from AMD, Google's TPU line, and custom silicon developed by major cloud providers. While the submitted figures come from NVIDIA itself, MLPerf results are generally produced under a shared, audited methodology across the industry, giving them some comparative weight.

For infrastructure buyers, the practical value of this announcement lies mainly in forward planning: cloud providers and large AI labs use such early benchmark disclosures to shape medium-term hardware procurement decisions, at a time when inference compute demand is growing faster than that dedicated to model training.

Tags
nvidiamlperfinferencevera-rubinbenchmarkdata-center-gpu

Read also