What Is NVIDIA Vera Rubin NVL72? Hardware, Architecture, Specs, and AI Factory Explained

Picture of Swastik Lamsal

Swastik Lamsal

Director of Digital Strategy & Marketing

NVIDIA Vera Rubin NVL72 is more than a new generation of GPU hardware. It is a rack-scale AI platform that combines 72 Rubin GPUs, 36 Vera CPUs, NVLink 6, ConnectX-9 SuperNICs, and BlueField-4 DPUs to support demanding AI training, reasoning, inference, and agentic AI workloads.

For buyers, the important question is not simply how fast Rubin GPUs are. It is how the complete NVL72 architecture moves data between CPUs and GPUs, connects multiple racks, feeds accelerators from storage, and handles the networking, power, and liquid cooling required to keep the system productive.

This guide breaks down the NVIDIA Vera Rubin NVL72 hardware, architecture, preliminary specifications, and AI factory design in practical terms. It also explains how Rubin, Vera, NVLink 6, ConnectX-9, BlueField-4, Spectrum-X or Quantum-X800, storage, and cooling work together and what organizations should evaluate before planning a deployment.

What Is NVIDIA Vera Rubin NVL72?

NVIDIA Vera Rubin NVL72 is NVIDIA’s third-generation rack-scale AI platform built around 72 Rubin GPUs and 36 Vera CPUs. NVIDIA designed the rack to function as one tightly connected compute system rather than 72 independent GPU servers.

The platform sits above an individual GPU or conventional server in the infrastructure stack. Buyers familiar with DGX and HGX platforms can view NVL72 as another step toward treating compute, interconnect, networking, power, and cooling as one coordinated architecture.

The current Vera Rubin NVL72 design combines compute, scale-up communication, external networking, infrastructure processing, cooling, and management. NVIDIA also supports larger POD-scale architectures that connect Vera Rubin compute racks with specialized CPU, storage, networking, and inference systems.

What Does NVL72 Mean?

For buyers, NVL72 means a rack-scale architecture with a 72-GPU NVLink domain. All 72 Rubin GPUs can communicate through the NVLink fabric instead of relying on standard Ethernet for communication inside the rack.

NVIDIA does not define “NVL” as a separate spelled-out product name on its current product pages. The useful distinction is the number: NVL72 identifies the 72-GPU NVLink domain.

What Hardware Is Inside NVIDIA Vera Rubin NVL72?

Infographic showing NVIDIA Vera Rubin NVL72 hardware, including Rubin GPUs, Vera CPUs, NVLink 6 fabric, ConnectX-9 SuperNICs, and BlueField-4 DPUs.

Vera Rubin NVL72 combines several purpose-built processors rather than depending on one type of chip. NVIDIA’s rack design uses 18 compute trays and nine NVLink switch trays, with compute, networking, infrastructure control, power, and cooling designed together.

Hardware layerVera Rubin NVL72 implementationPrimary role
AI accelerators72 NVIDIA Rubin GPUsTraining, reasoning, inference, AI compute
Host compute36 NVIDIA Vera CPUsData movement, orchestration, agentic workloads
Scale-up fabricNVIDIA NVLink 6 + NVLink switchesGPU-to-GPU communication inside the rack
Scale-out I/ONVIDIA ConnectX-9 SuperNICsConnects the rack to external AI fabrics
InfrastructureNVIDIA BlueField-4 DPUsNetworking, storage, security, isolation

NVIDIA’s current rack design places two Vera Rubin Superchips in each compute tray. The tray also integrates ConnectX-9 networking and BlueField-4 infrastructure processing, reducing the number of separate components that operators must assemble manually.

That system-level dependency is also why GPU server build planning must account for CPUs, memory, storage, networking, power, cooling, and service access rather than focusing only on accelerator count.

NVIDIA Rubin GPU

Rubin is the accelerator at the center of the platform, but it should not be viewed as a standalone graphics card. Each Rubin GPU includes 288 GB of HBM4 memory with up to 22 TB/s of memory bandwidth and delivers up to 50 PFLOPS of NVFP4 inference compute.

Rubin advances the same system-level principle seen in earlier dense SXM platforms. Catalyst’s H100 SXM architecture coverage shows why the accelerator, GPU baseboard, server design, networking, power, and cooling must be evaluated together.

Across 72 GPUs, Vera Rubin NVL72 provides 20.7 TB of HBM4 and up to 1,580 TB/s of aggregate GPU memory bandwidth. That capacity helps large models, mixture-of-experts workloads, and long-context inference keep more active data close to the accelerators.

NVIDIA Vera CPU

Each Vera CPU contains 88 custom NVIDIA Olympus cores. Across 36 processors, the NVL72 rack provides 3,168 CPU cores and 54 TB of LPDDR5X CPU memory.

The CPU has a larger role in agentic AI than simply launching GPU jobs. Agents retrieve data, execute code, call tools, manage sandboxes, preprocess information, and orchestrate repeated inference requests, which increases pressure on CPU compute and memory bandwidth.

NVIDIA NVLink 6

NVLink 6 forms the scale-up network inside Vera Rubin NVL72. It lets the 72 GPUs exchange data through a rack-scale fabric with 260 TB/s of aggregate NVLink switch bandwidth, or 3.6 TB/s per GPU.

This matters for workloads such as model parallelism, mixture-of-experts routing, collective operations, and synchronized training. When GPUs spend too much time waiting for other GPUs, accelerator utilization falls even when the individual chips remain fast.

NVIDIA ConnectX-9 SuperNIC

ConnectX-9 connects Vera Rubin compute to the scale-out network outside the NVLink domain. NVIDIA lists up to 1.6 Tb/s of total ConnectX-9 throughput, with support for both high-speed Ethernet and InfiniBand connectivity.

This layer becomes critical when a model, dataset, or workload spans multiple NVL72 racks. Internal NVLink handles communication inside the rack, while ConnectX-9 feeds external switches that link many racks into a larger AI factory.

NVIDIA BlueField-4 DPU

BlueField-4 handles infrastructure work that would otherwise consume CPU and GPU resources. NVIDIA positions the DPU for networking, storage, security, telemetry, isolation, and infrastructure control.

BlueField-4 integrates up to 800 Gb/s of network connectivity, a 64-core Grace CPU, PCIe Gen6, and memory dedicated to infrastructure processing. It is a DPU, not another GPU.

How Does NVIDIA Vera Rubin NVL72 Architecture Work?

Infographic explaining NVIDIA Vera Rubin NVL72 architecture, including CPU-to-GPU links, NVLink 6 scale-up, ConnectX-9 networking, scale-out fabric, and BlueField-4 infrastructure services.

Vera Rubin NVL72 uses different interconnects for different distances. NVLink 6 scales up inside one rack, while ConnectX-9 and Ethernet or InfiniBand scale the system out across racks. BlueField-4 handles supporting infrastructure services.

Architecture layerTechnologyWhat it connectsWhy it matters
CPU-to-GPUNVLink-C2CVera CPUs to Rubin GPUsHigh-bandwidth host-to-accelerator data movement
GPU scale-upNVLink 672 Rubin GPUsCreates one tightly coupled GPU domain
Server/rack I/OConnectX-9NVL72 to external fabricMoves traffic beyond the rack
Scale-outSpectrum-X / Spectrum-6 or Quantum-X800Many racks and clustersBuilds large AI factories
Infrastructure servicesBlueField-4Network, storage, securityOffloads control-plane and data-path work

A simplified workload path looks like this:

  • Vera CPUs prepare, move, and orchestrate data and tasks.
  • Rubin GPUs execute the accelerated AI workload.
  • NVLink 6 moves data among GPUs inside the rack.
  • ConnectX-9 sends external traffic through Ethernet or InfiniBand.

The result is a hierarchy rather than one flat network. That separation helps each communication layer use technology designed for its specific latency, bandwidth, and scale requirements.

NVIDIA Vera Rubin NVL72 Specifications

NVIDIA currently labels the published Vera Rubin specifications as preliminary and subject to change. Buyers should confirm final specifications against the current NVIDIA datasheet and the selected OEM configuration before ordering infrastructure.

SpecificationNVIDIA Vera Rubin NVL72
GPU configuration72 NVIDIA Rubin GPUs
CPU configuration36 NVIDIA Vera CPUs
GPU memory20.7 TB HBM4
GPU memory bandwidthUp to 1,580 TB/s
CPU memory54 TB LPDDR5X
CPU cores3,168 NVIDIA Olympus cores
NVFP4 inferenceUp to 3,600 PFLOPS
NVFP4 trainingUp to 2,520 PFLOPS
FP8/FP6 trainingUp to 1,260 PFLOPS
NVLink 6 bandwidth260 TB/s rack scale
NVLink-C2C bandwidth65 TB/s rack scale
Scale-out bandwidth28.8 TB/s

The training specifications above use NVIDIA’s published precision definitions and footnotes. Peak FLOPS should not be treated as application performance because model architecture, software, communication, storage, and utilization can change real-world results substantially.

Why Does NVIDIA Call Vera Rubin NVL72 an “AI Factory”?

NVIDIA uses AI factory to describe infrastructure that continuously converts data and electrical power into useful AI output. The concept shifts attention from one GPU benchmark to the productivity of the entire data center.

Vera Rubin targets several stages of that process:

  • Large-scale model pretraining.
  • Post-training and reinforcement learning.
  • Test-time scaling and AI reasoning.
  • Real-time agentic AI inference.

These workloads place pressure on more than GPU compute. CPU orchestration, memory movement, network congestion, storage throughput, cooling capacity, and available electrical power can all limit the amount of useful AI work a rack produces.

That is why Vera Rubin represents Rubin AI infrastructure, not simply a new accelerator generation.

What Networking Does Vera Rubin NVL72 Need?

Vera Rubin uses ConnectX-9 SuperNICs at the compute edge and supports two major scale-out paths: NVIDIA Spectrum-X Ethernet and NVIDIA Quantum-X800 InfiniBand. The correct choice depends on cluster size, workload behavior, operating model, latency goals, and existing network standards.

NVIDIA Spectrum-6 now anchors the newer Spectrum-X generation. NVIDIA describes Spectrum-6 as a 102.4 Tb/s Ethernet switch platform, giving large AI factories a path toward higher-density and larger-scale Ethernet fabrics.

Quantum-X800 provides an InfiniBand path with up to 800 Gb/s end-to-end connectivity and hardware-assisted features such as adaptive routing and in-network computing. NVIDIA positions it for very large AI and HPC clusters.

Buyers must also account for broader AI networking challenges because synchronized collective traffic, congestion, microbursts, and inconsistent paths can leave accelerators waiting even when headline network bandwidth looks sufficient.

What Optics and Cabling Are Required?

The answer depends on whether the design uses Ethernet or InfiniBand, rack distance, port speed, switch model, and topology. High-speed AI fabrics may use DACs for short copper runs and active optical cables or transceivers where distance requires fiber.

Port speed alone does not prove compatibility. Teams should validate switch ports, OSFP form factors, breakout behavior, fiber type, cable length, firmware, and the complete approved bill of materials before deployment.

For Rubin-generation infrastructure, the correct optic or cable should come from the validated switch, NIC, topology, and OEM design rather than assumptions based only on connector type or advertised speed.

What Storage Does Vera Rubin NVL72 Need?

Vera Rubin NVL72 does not remove the need for high-performance external storage. Training datasets, model checkpoints, inference context, RAG data, and shared model files must reach compute fast enough to avoid leaving expensive GPUs waiting for data.

A practical architecture may need four storage functions:

  • High-throughput storage for training datasets.
  • Fast checkpoint storage for long training jobs.
  • Context or KV-cache storage for agentic inference.
  • Durable storage for models, results, and recovery.

NVIDIA has added BlueField-4 STX and its context-memory architecture to the broader Vera Rubin platform. STX targets long-context and agentic AI workloads where systems need to share and reuse KV-cache data efficiently.

That does not make STX the only storage answer. Buyers still need to size capacity, throughput, latency, redundancy, protocols, and network connections around the actual training or inference workload.

What Cooling Does Vera Rubin NVL72 Require?

Vera Rubin NVL72 uses direct liquid cooling as part of its third-generation MGX rack architecture. NVIDIA describes a warm-water, single-phase design that can operate with a 45°C supply temperature, while still maintaining the same physical rack footprint as earlier MGX generations.

Facility teams should treat cooling as a core system requirement, not a post-purchase detail. Broader AI cooling strategies highlight why multiple factors must be evaluated together:

  • Rack density and heat load
  • Sustained GPU utilization
  • Water availability and facility integration
  • Heat rejection capacity
  • Redundancy and failover design
  • Retrofit and space constraints

The rack design integrates several cooling and infrastructure elements:

  • Redesigned liquid manifolds for high-density compute
  • Dedicated cooling paths for networking and compute components
  • Modular service and maintenance architecture

As a result, data centers need more than physical rack space. They must also provide:

  • Compatible facility water systems
  • CDUs (coolant distribution units)
  • Liquid manifolds and plumbing design
  • Heat rejection and chiller capacity
  • Redundant cooling loops
  • Service and maintenance access planning

NVIDIA also incorporates rack-level energy storage and Intelligent Power Smoothing. This helps reduce sudden power spikes caused by synchronized AI workloads, making power delivery part of the compute architecture itself rather than just a facility utility.

NVIDIA does not publish a single fixed rack power number for NVL72. Instead, power varies by configuration. Buyers should validate:

  • Final OEM system configuration
  • Rack design and density
  • Cooling infrastructure capability
  • Deployment scale and cluster size

Assuming a universal power figure can lead to under-provisioning or infrastructure mismatch.

Vera Rubin NVL72 vs GB200 NVL72 vs GB300 NVL72

Infographic comparing GB200 NVL72, GB300 NVL72, and Vera Rubin NVL72 across GPUs, CPUs, memory, NVLink, bandwidth, and networking.

Vera Rubin keeps the 72-GPU rack-scale concept introduced with Blackwell but changes the GPU, CPU, memory, networking, and scale-up fabric. The largest architectural jump is the move from Grace and fifth-generation NVLink to Vera and NVLink 6.

FeatureGB200 NVL72GB300 NVL72Vera Rubin NVL72
GPU generationBlackwellBlackwell UltraRubin
GPUs727272
CPU36 Grace36 Grace36 Vera
GPU memory13.4 TB HBM3E20 TB HBM3E20.7 TB HBM4
CPU cores2,5922,5923,168
NVLink5th generation5th generationNVLink 6
Rack NVLink bandwidth130 TB/s130 TB/s260 TB/s
NVIDIA NIC generationConnectX-7 classConnectX-8ConnectX-9
Primary evolutionFirst 72-GPU Blackwell rackMore memory and reasoning capacityAgentic AI and next-gen scale

Organizations tracing the move from Hopper to newer rack-scale architectures can use the H100 SXM5 accelerator as a practical reference for how earlier SXM systems depended on NVLink, qualified server platforms, power, cooling, and high-speed networking.

GB300 NVL72 remains an available Blackwell Ultra platform, while Vera Rubin is the newer architecture now moving through its production and partner deployment ramp. A buyer should compare deployment timing, software requirements, facility readiness, cost, and workload,not simply pick the newest generation.

What Workloads Fit NVIDIA Vera Rubin NVL72?

Vera Rubin NVL72 primarily fits environments where large-scale compute utilization matters enough to justify rack-scale infrastructure. Workload requirements should come before accelerator selection because training, inference, and HPC place different pressure on compute, storage, networking, and scheduling.

Catalyst’s GPU deployment strategy provides useful context for those differences, especially when organizations need to balance AI training, inference, traditional HPC, or converged AI and scientific workloads.

Primary Vera Rubin NVL72 workloads include:

  • Frontier and mixture-of-experts model training.
  • Large reasoning and test-time scaling workloads.
  • High-throughput or long-context inference.
  • Large agentic AI and neocloud services.

Scientific computing can also use Rubin technology, although NVIDIA positions Vera Rubin NVL4 and specialized Rubin HPC systems more directly toward some AI-for-science and high-precision workloads.

Who Is NVIDIA Vera Rubin NVL72 For?

Strong candidates include hyperscalers, neocloud providers, sovereign AI operators, major research organizations, and enterprises building very large centralized AI factories.

It becomes more compelling when the organization can keep 72 GPUs busy and already plans for high-speed storage, large scale-out networks, direct liquid cooling, and high-density electrical infrastructure.

Who Might Not Need Vera Rubin NVL72?

NVL72 will not be the right starting point for every AI project.

  • Small AI teams running development or moderate inference.
  • Enterprises that need only one or a few GPU servers.
  • Facilities without rack-scale liquid-cooling capability.
  • Buyers prioritizing lower entry cost over maximum scale.

A smaller platform can offer better utilization and a simpler facility requirement when the workload cannot consume an NVL72 rack efficiently.

What Are the Main Vera Rubin NVL72 Alternatives?

HGX Rubin NVL8 gives organizations an eight-GPU Rubin platform and supports Vera or x86 host architectures. It can fit enterprises that want Rubin without adopting a 72-GPU rack-scale system.

Existing Hopper infrastructure can also remain practical for workloads that do not require Rubin-scale deployment. A dense H100 server can support eight-GPU AI training and HPC where platform availability, budget, or current software qualification favors an established generation.

Vera Rubin NVL4 targets scientific computing and AI-for-science with four Rubin GPUs and two Vera CPUs. It provides a different balance for simulation, HPC, and high-precision workloads.

GB300 NVL72 remains a current Blackwell Ultra rack-scale option. It may make more sense when availability, existing Blackwell infrastructure, deployment timing, or current qualification matters more than moving immediately to Rubin.

AMD Helios provides a non-NVIDIA rack-scale path using 72 Instinct MI455X GPUs, EPYC Venice CPUs, Pensando networking, and ROCm software. AMD says Helios shipments began moving toward customer deployments in the second half of 2026.

What Should Be Bought With NVIDIA Vera Rubin NVL72?

NVIDIA Vera Rubin NVL72 Deployment Checklist

A Vera Rubin deployment should be planned as a complete bill of materials, not as an accelerator order.

Deployment layerWhat buyers need to verify
Rack-scale computeDGX/OEM implementation, configuration, support and deployment model
Scale-out networkSpectrum-X/Spectrum-6 or Quantum-X800 topology and capacity
ConnectivityCorrect OSFP optics, DACs, AOCs, fiber and breakout design
StorageDataset, checkpoint, object, RAG and context-memory requirements
Facility infrastructureRack power, liquid cooling, CDU, manifolds and redundancy

Management networking, software, security, spare parts, service contracts, deployment labor, and cluster orchestration also need to fit the design.

The central question should be: Can every layer keep the Rubin GPUs productive? Oversized compute with undersized networking, storage, power, or cooling creates an expensive bottleneck.

Is NVIDIA Vera Rubin NVL72 Available?

Vera Rubin is in production and moving into partner deployments, but availability should still be checked for the exact system and region. NVIDIA announced the production ramp in May 2026 and reported in July that Vera Rubin racks were already running at several cloud and AI infrastructure partners.

This does not mean every Vera Rubin NVL72 configuration has the same lead time or order status. OEM platform, partner, geography, allocation, qualification, and deployment requirements can affect when a buyer can receive a system.

For that reason, announced, in production, orderable, shipping, deployed, and generally available should not be treated as identical terms.

How Can Catalyst Support a Vera Rubin NVL72 Deployment?

Catalyst Data Solutions Inc can help organizations evaluate complete AI and HPC configurations across compute, servers, networking, storage, optics, cabling, power, and supporting infrastructure. The goal should be to match the architecture to workload, facility limits, availability, budget, and deployment timeline.

Buyers can request current availability for Vera Rubin-related infrastructure, OEM alternatives, lead-time questions, or a complete configuration quote.

Catalyst can also compare NVIDIA, AMD, Dell, HPE, Supermicro, networking, and supporting hardware without assuming one platform fits every deployment. Buyers planning several infrastructure layers together can use a complete AI infrastructure bundle to frame the sourcing request around the complete system rather than one component.

A complete quote request should include the intended workload, rack quantity, deployment location, network fabric, storage requirements, optics and cabling, available power, cooling design, desired support, and project timeline.

FAQs

Is NVIDIA Vera Rubin NVL72 a GPU?

No. Vera Rubin NVL72 is a rack-scale system architecture containing 72 Rubin GPUs, 36 Vera CPUs, NVLink 6, ConnectX-9 networking, and BlueField-4 DPUs.

How many GPUs are in Vera Rubin NVL72?

Vera Rubin NVL72 contains 72 NVIDIA Rubin GPUs connected through a sixth-generation NVLink fabric.

How much GPU memory does Vera Rubin NVL72 have?

The rack provides 20.7 TB of HBM4 GPU memory with up to 1,580 TB/s of aggregate memory bandwidth, based on NVIDIA’s current preliminary specifications.

What does NVL72 mean?

NVL72 identifies a rack architecture with a 72-GPU NVLink domain. The GPUs use NVLink for high-bandwidth communication within the rack.

Does Vera Rubin NVL72 require liquid cooling?

Yes. NVIDIA designed the third-generation MGX NVL72 architecture around direct liquid cooling, including a 45°C warm-water cooling design.

What network does Vera Rubin NVL72 use?

Inside the rack, it uses NVLink 6. For scale-out connectivity, ConnectX-9 links the system to NVIDIA Spectrum-X Ethernet, including Spectrum-6 platforms, or Quantum-X800 InfiniBand.

Is Vera Rubin NVL72 shipping now?

Production and partner deployment are ramping during 2026. NVIDIA reported Vera Rubin racks running at partners in July, but buyers should check the exact OEM, region, configuration, and lead time before treating a system as immediately orderable.

Is Vera Rubin NVL72 better than GB300 NVL72?

It is newer, but that does not make it the automatic choice for every buyer. Vera Rubin adds HBM4, Vera CPUs, NVLink 6, and ConnectX-9, while GB300 NVL72 remains an available Blackwell Ultra platform that may better fit current infrastructure or deployment schedules.

Research and Fact-Check Record

Research Date: August 14, 2026
Official Datasheet Version: 
https://nvdam.widen.net/s/7hztspzswk/gpu-architecture-datasheet-vera-rubin-nvidia-us-5198950-web

Official Vera Rubin NVL72 product/specification:
https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72

Research Key Point 1: NVIDIA Vera Rubin NVL72 combines 72 Rubin GPUs, 36 Vera CPUs, NVLink 6, ConnectX-9 SuperNICs, and BlueField-4 DPUs, with NVIDIA Quantum-X800 InfiniBand or Spectrum-X Ethernet for scale-out connectivity.

Research Key Point 2: NVIDIA’s current preliminary specifications list 20.7 TB of HBM4 GPU memory, up to 1,580 TB/s of aggregate GPU memory bandwidth, 260 TB/s of NVLink 6 switch bandwidth, and 3,168 custom NVIDIA Olympus CPU cores per Vera Rubin NVL72 rack.

Research Key Point 3: NVIDIA says Vera Rubin production is ramping worldwide and reported in July 2026 that Vera Rubin NVL72 racks were running at CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure, and Nebius. Buyers should separately verify current orderability, OEM configuration, regional availability, and lead time before procurement.