The Vera Rubin platform sets new standards for data center energy efficiency by delivering 30x higher inference performance and lowering the cost of generating intelligent tokens by roughly 35x.
Article index:
- An exceptional leap in inference efficiency and cost reduction
- Architectural details of the Vera CPU
- Launching the Scale-In networking architecture as a fifth pillar
- SpaceX xAI partnership and the orbital satellite project
- Frequently asked questions
An exceptional leap in inference efficiency and cost reduction
Nvidia published detailed performance data on Monday showing that its next-generation platform, the Vera Rubin NVL72, achieves superior inference throughput up to 30 times higher per megawatt compared to its current GB300 NVL72 system when running agentic AI workloads, while cutting the financial cost per million tokens by about 35 times. Concurrently, the company showcased the technical architecture of its new “Vera” central processor at the Hot Chips 2026 conference and introduced the “Scale-In” architecture as a fundamental fifth pillar for AI networks.
Productivity rates were measured using the SemiAnalysis AgentX benchmark, which simulates real agentic programming sessions including conversational context growth, external tool calling, and sub-agent generation. When running the DeepSeek-V4 Pro model, the Vera Rubin system outperformed the GB300 system by 30 times, which in turn delivers performance 15 times higher per megawatt than the previous Hopper architecture. This leap is exceptionally important because intelligent agent tasks consume about 15 times more tokens than simple chat requests, according to OpenRouter data, helping data center operators maximize workload output within available power budgets via DSX Max LPS power management technology.
Architectural details of the Vera CPU
During the Hot Chips conference, Nvidia detailed the design of the Vera processor, equipped with 88 cores built on the custom Olympus architecture. The processor prioritizes single-threaded performance and low latency rather than increasing the number of weak cores. It features a 10-lane decode front-end, simultaneous multithreading, eight LPDDR5X memory controllers providing up to 1.2 terabytes per second of bandwidth, and a six-die design connected by ultra-fast NVLink interconnects. Ian Buck, vice president of hyperscale and HPC at Nvidia, stated: “Agentic AI requires a new kind of computing system; a system built not just to generate answers, but to take practical actions and execute them.”
Launching the Scale-In networking architecture as a fifth pillar
Nvidia also introduced the Scale-In networking architecture powered by BlueField-4 data processing units to secure and manage agentic AI factories at scale, joining scale-up, scale-out, data-center-interconnect, and storage networks as a fifth foundational pillar.
SpaceX xAI partnership and the orbital satellite project
In a separate announcement, SpaceX xAI revealed that it will deploy Vera processors across its infrastructure dedicated to the Grok model, with plans to send a custom Vera Rubin system into space aboard the first generation of its Starmind smart satellites.
Frequently asked questions
Question: What is the performance improvement offered by the Vera Rubin platform?
Answer: It delivers 30x higher inference throughput per megawatt and a 35x reduction in token generation cost compared to the GB300 system.
Question: What are the technical specifications of the Vera CPU?
Answer: It includes 88 custom Olympus high-performance cores, delivers up to 1.2 TB/s of memory bandwidth, and features advanced NVLink interconnects.
Question: How will SpaceX xAI benefit from these processors?
Answer: It will deploy them to run and train the Grok model on the ground, and plans to send custom systems into space via Starmind orbital satellites.