On January 5, 2026, at CES in Las Vegas, NVIDIA officially launched the Rubin platform, a system composed of six chips designed to form a single artificial intelligence supercomputer. The announcement marks a turning point in the race for AI infrastructure, as the demand for computation for training and inference reaches unprecedented levels.
Six chips designed as a single system
What distinguishes Rubin from all previous generations is the design principle NVIDIA calls "extreme codesign." Rather than optimizing each component separately, the company designed the entire architecture simultaneously and integrated it. The six chips that make up the platform are the Vera CPU, the Rubin GPU, the NVLink 6 Switch, the ConnectX-9 SuperNIC, the BlueField-4 DPU, and the Spectrum-6 Ethernet Switch.
This approach treats the data center as a computing unit, rather than just a GPU server. It ensures that performance and efficiency are maintained in real-world deployments, not just in isolated benchmarks. It is a profound shift in philosophy in how to design AI infrastructure at scale.
It should be noted that in March 2026, the platform integrated a seventh chip: the Groq 3 LPX, a very low-latency inference accelerator, thus bringing the architecture to seven distinct components.
Massive performance gains over Blackwell
The figures announced by NVIDIA are considerable. The Rubin GPU achieves 50 petaflops in NVFP4 inference and 35 petaflops in NVFP4 training, which is 5x and 3.5x the performance of Blackwell, respectively, for only 1.6x more transistors.
Economically, the Rubin platform allows for up to a 10x reduction in the cost per inference token and requires 4x fewer GPUs to train MoE (mixture-of-experts) models compared to the Blackwell platform. Jensen Huang, founder and CEO of NVIDIA, clearly stated this goal: to reduce the cost of generating tokens to about one-tenth of that of the previous generation, in order to make the deployment of AI at scale much more economically accessible.
These gains are not just about raw speed. The Vera Rubin platform is designed to process hundreds of thousands of input tokens in long context, an essential capability for agentic workflows, complex reasoning, and multimodal pipelines.
The NVL72: a rack that reinvents installation
The Vera Rubin NVL72 groups 72 GPUs in a single rack. Several assembled NVL72s form the DGX SuperPOD, a very large-scale AI supercomputer that hyperscalers like Microsoft, Google, Amazon, and Meta are spending billions to acquire.
Physically, the NVL72 introduces radical changes. The system is entirely liquid-cooled, with no fans, no exposed tubes, and no visible wiring. This design reduces installation time to five minutes, compared to two hours for equivalent Blackwell systems.
NVIDIA also claims that the Vera Rubin NVL72 rack offers more bandwidth than the entire internet, a spectacular marketing statement that illustrates the connectivity density aimed for in agentic uses and the most demanding reasoning models.
A global ecosystem ready for deployment
The platform is in full production, and the first products will be available from partners in the second half of 2026. Among the first cloud providers to deploy instances based on Vera Rubin will be AWS, Google Cloud, Microsoft, and OCI, as well as NVIDIA’s cloud partners: CoreWeave, Lambda, Nebius, and Nscale.
Microsoft will deploy Vera Rubin NVL72 systems in its future Fairwater AI superfactory sites, intended to house hundreds of thousands of superchips. On the server manufacturer side, Cisco, Dell, HPE, Lenovo, and Supermicro plan to offer a wide range of Rubin-based systems.
Major AI labs are also on board. Anthropic, OpenAI, Meta, xAI, Mistral AI, Cohere, Perplexity, and Runway are among the organizations planning to use the Rubin platform to train larger models and improve their large-scale inference capabilities.
Industry CEOs speak out
The announcement was accompanied by statements from several top executives. “Intelligence scales with compute. When we add more compute, models get more capable, solve harder problems, and have a greater impact for people,” said Sam Altman, CEO of OpenAI, in an official statement. (“Intelligence scales with compute. When we add more compute, models get more capable, solve harder problems and make a bigger impact for people.”)
“The efficiency gains of the NVIDIA Rubin platform represent the kind of infrastructure progress that enables longer memory, better reasoning, and more reliable results,” said Dario Amodei, co-founder and CEO of Anthropic, in an official statement. (“The efficiency gains in the NVIDIA Rubin platform represent the kind of infrastructure progress that enables longer memory, better reasoning and more reliable outputs.”)
“We are building the world’s most powerful AI factories to serve any workload, anywhere, with maximum performance and efficiency,” said Satya Nadella, Executive Chairman and CEO of Microsoft, in an official statement. (“We are building the world’s most powerful AI factories to serve any workload, anywhere, with maximum performance and efficiency.”)
Jensen Huang, for his part, summarized NVIDIA’s vision in these terms: “Rubin arrives at exactly the right moment, as AI computing demand for both training and inference is skyrocketing. With our annual cadence of delivering a new generation of AI supercomputers, Rubin takes a giant leap towards the next frontier of AI.” (“Rubin arrives at exactly the right moment, as AI computing demand for both training and inference is going through the roof.”)
The roadmap continues with Rubin Ultra in 2027
The Rubin platform is not an endpoint; it will be followed by Rubin Ultra, expected in the second half of 2027, continuing the annual pace that NVIDIA has set for renewing its generations of AI supercomputers.
However, this sustained cadence exposes NVIDIA to very high market expectations. The company must constantly exceed Wall Street’s forecasts and reassure that AI infrastructure spending does not outpace actual demand. Competition is also intensifying, with AMD pushing its own rack-scale solutions, and clients like Google and Amazon developing their own custom chips.
Nevertheless, the Rubin platform represents a structuring step for the entire AI industry. By treating the data center as a unified computer rather than a collection of components, NVIDIA is redefining what it means to build artificial intelligence infrastructure. Players who deploy these systems starting in the second half of 2026 will have a significant cost and capacity advantage for training and deploying the next generation of reasoning models.



No comments yet — start the discussion!