NVIDIA Vera Rubin NVL4: Why This Platform Matters for HPC and AI-for-Science

Picture of Swastik Lamsal

Swastik Lamsal

Director of Digital Strategy & Marketing

NVIDIA Vera Rubin NVL4 is built for a growing challenge in scientific computing: researchers increasingly need one platform that can handle traditional simulation, AI training, inference, and data-heavy analysis without treating each workload as a separate system.

The platform combines four NVIDIA Rubin GPUs with two NVIDIA Vera CPUs, linked through NVLink 6 and NVLink-C2C. That design targets demanding HPC and AI-for-Science workloads where FP64 performance, memory bandwidth, and fast CPU-to-GPU communication can directly affect how quickly simulations and scientific models run.

For organizations evaluating next-generation HPC infrastructure, the important question is not simply how fast Rubin is. It is how NVL4 fits into a complete environment that includes the server, networking, storage, liquid cooling, power, and cluster architecture needed to keep those accelerators productive.

This guide explains the NVIDIA Vera Rubin NVL4 architecture, specifications, HPC advantages, AI-for-Science use cases, interconnects, deployment requirements, and differences from NVL72 and Grace Hopper, using official NVIDIA specifications and clearly distinguishing announced performance claims from independently verified results.

What Is NVIDIA Vera Rubin NVL4?

NVIDIA Vera Rubin NVL4 is a four-GPU accelerated computing platform built specifically around the convergence of HPC and AI. NVIDIA connects four Rubin GPUs to two Vera CPUs and supports the design in liquid-cooled NVIDIA MGX modular servers.

The reason NVL4 exists is different from the reason NVIDIA offers massive NVL72 systems. NVL4 gives scientific computing organizations a dense node for simulation, modeling, AI training, AI inference, and hybrid workflows without requiring a 72-GPU NVLink rack.

That positioning also makes NVL4 relevant to universities, national laboratories, industrial R&D groups, energy companies, and HPC centers. NVIDIA specifically identifies climate modeling, computational fluid dynamics, quantum chemistry, and energy exploration among the workloads targeted by Vera Rubin scientific computing systems.

How Does Vera Rubin NVL4 Work?

Infographic explaining how the NVIDIA Vera Rubin NVL4 platform works, showing its architecture layers: four NVIDIA Rubin GPUs, two NVIDIA Vera CPUs, sixth-generation NVLink GPU-to-GPU scaling, NVLink-C2C CPU-to-GPU connection, HBM4 memory, liquid-cooled NVIDIA MGX-compatible systems, high-speed InfiniBand or Ethernet cluster networking, and direct liquid cooling. The diagram presents the complete hardware stack and communication path of Vera Rubin NVL4.

The Vera Rubin NVL4 architecture combines CPU compute, GPU acceleration, high-bandwidth memory, and two types of NVIDIA interconnect technology.

Four Rubin GPUs communicate through a second-generation NVLink bridge running NVLink 6. Two Vera CPUs connect to the GPUs through NVLink-C2C, giving CPU and GPU memory a high-speed coherent communication path.

Architecture LayerNVIDIA Vera Rubin NVL4 Role
Accelerator4 NVIDIA Rubin GPUs
Host compute2 NVIDIA Vera CPUs
GPU-to-GPU scale-upSixth-generation NVLink
CPU-to-GPU connectionNVLink-C2C
GPU memoryHBM4
Server architectureLiquid-cooled NVIDIA MGX-compatible systems
Cluster networkingHigh-speed InfiniBand or Ethernet, depending on implementation
CoolingDirect liquid cooling

This relationship matters more than any one component. A scientific application may use the CPU for orchestration and data preparation, send intensive numerical or AI operations to the GPUs, exchange intermediate data repeatedly, and then write large outputs to external storage.

That is why an NVL4 deployment should be treated as a complete GPU server design problem rather than simply a four-GPU purchase.

What Hardware Is Inside Vera Rubin NVL4?

The core NVL4 design includes four Rubin GPUs and two 88-core Vera CPUs. NVIDIA says the GPUs use HBM4, while Vera combines custom Olympus CPU cores with high-bandwidth LPDDR5X memory.

The main hardware relationships are:

  • Four Rubin GPUs: Handle FP64 scientific computing as well as lower-precision AI training and inference.
  • Two Vera CPUs: Provide host compute, data movement, orchestration, preprocessing, and CPU-heavy scientific work.
  • NVLink 6 and NVLink-C2C: Connect GPUs to each other and CPUs to GPU compute with high bandwidth.
  • MGX and liquid cooling: Place the components into deployable high-density server designs.

NVIDIA says Vera provides up to 1.8 TB/s of coherent NVLink-C2C bandwidth between Vera CPUs and NVIDIA GPUs. The company also lists up to 1.5 TB of LPDDR5X capacity and 1.2 TB/s of memory bandwidth for Vera CPU configurations.

Actual server memory, networking, local storage, and power configuration can vary by OEM implementation. This distinction matters when comparing an NVIDIA platform architecture with a finished Dell, Supermicro, or other manufacturer’s system.

For broader platform context, Catalyst’s DGX and HGX guide explains how NVIDIA accelerator architectures become complete data center systems.

What Are NVIDIA Vera Rubin NVL4 Specs?

NVIDIA’s published Rubin figures remain preliminary and subject to change. NVIDIA currently publishes detailed Rubin GPU and Vera Rubin Superchip specifications, while some NVL4 totals below are simple calculations from four published Rubin GPU figures rather than a separate final NVL4 datasheet value.

SpecificationVera Rubin NVL4
GPU configuration4 NVIDIA Rubin GPUs
CPU configuration2 NVIDIA Vera CPUs
CPU cores88 Olympus cores per Vera CPU; 176 across two CPUs
GPU memory288 GB HBM4 per Rubin GPU; about 1.152 TB across 4 GPUs*
GPU memory bandwidth22 TB/s per GPU; about 88 TB/s across 4 GPUs*
Native FP6433 TFLOPS per GPU; about 132 TFLOPS across 4 GPUs*
NVFP4 inference50 PFLOPS per GPU; about 200 PFLOPS across 4 GPUs*
NVFP4 training35 PFLOPS per GPU; about 140 PFLOPS across 4 GPUs*
GPU interconnectSixth-generation NVIDIA NVLink
NVLink bandwidthUp to 3.6 TB/s per Rubin GPU
CPU-GPU interconnectNVIDIA NVLink-C2C
NVLink-C2C bandwidthUp to 1.8 TB/s coherent bandwidth
Server supportLiquid-cooled NVIDIA MGX modular servers

Calculated from NVIDIA’s preliminary per-GPU figures. Because NVIDIA rounds some published performance values, these calculated totals should not be treated as final OEM system specifications.

One important distinction involves FP64 numbers. NVIDIA lists 33 TFLOPS of native FP64 per Rubin GPU and separately lists higher DGEMM performance that uses Tensor Core-based emulation algorithms. Those values should not be presented as though they describe the same type of FP64 execution.

Why Does Vera Rubin NVL4 Matter for HPC?

Traditional HPC depends heavily on double-precision FP64 compute because scientific simulations often need more numerical accuracy than many AI workloads. Climate models, fluid simulations, chemistry calculations, and physics codes can accumulate errors when numerical precision is insufficient.

Rubin provides native FP64 capability while also supporting the AI performance required for emerging scientific machine learning. NVIDIA describes Vera Rubin as a single accelerated platform for numerical solvers, AI models, instrument data, and real-time analytics.

Three characteristics are especially important for HPC:

  • FP64 performance supports high-precision numerical simulation.
  • HBM4 bandwidth helps feed large scientific datasets to GPU compute.
  • NVLink 6 reduces communication limits when multiple GPUs cooperate.
  • NVLink-C2C improves data movement between host CPUs and GPUs.

The result is not simply a faster scientific computing GPU. It is a tightly connected CPU-GPU node intended to keep simulation, data movement, and accelerated analysis working together.

That systems approach matches how modern HPC GPU deployments must be planned: compute performance matters only when networking, storage, power, cooling, and application architecture can keep the accelerators productive.

Why Is NVL4 Important for AI-for-Science?

AI-for-Science uses machine learning alongside traditional scientific computing to speed up discovery, analyze complex data, approximate expensive calculations, and improve how researchers explore large design spaces.

Vera Rubin NVL4 matters because the same node provides native scientific compute and high-throughput AI acceleration. NVIDIA specifically positions Vera Rubin for combining simulation, AI, and data-intensive research.

Key AI-for-Science workflows include:

  • Scientific foundation models: Train models on scientific data across chemistry, biology, physics, climate, or materials.
  • Surrogate models: Use AI to approximate expensive simulations and evaluate possibilities faster.
  • Simulation plus AI: Combine numerical solvers with learned models inside one research workflow.
  • AI-assisted analysis: Apply training and inference to simulation output, instrument data, and scientific datasets.

A researcher might first run a physics-based simulation, use a neural model to approximate part of the solution space, evaluate the AI prediction, and then return selected cases to a higher-fidelity simulation. That hybrid workflow places demands on CPU compute, GPU compute, memory, storage, and communication at the same time.

What Does NVLink-C2C Do in Vera Rubin NVL4?

NVLink-C2C is NVIDIA’s high-speed coherent connection between Vera CPUs and NVIDIA GPUs. NVIDIA states that the current Vera implementation provides up to 1.8 TB/s of coherent CPU-GPU bandwidth.

Coherency matters because tightly coupled HPC applications frequently move data between CPU-managed operations and GPU kernels. A faster coherent path can reduce the communication bottleneck that appears when CPUs repeatedly prepare, inspect, transform, or coordinate data used by GPUs.

NVLink-C2C serves a different purpose from NVLink 6. NVLink-C2C connects CPU and GPU resources, while sixth-generation NVLink provides the high-bandwidth GPU-to-GPU communication needed inside the NVL4 scale-up domain.

What Scientific Workloads Can Vera Rubin NVL4 Run?

Infographic showing scientific workloads supported by the NVIDIA Vera Rubin NVL4 platform, including climate modeling, computational fluid dynamics, quantum chemistry, materials science, energy exploration, physics, scientific AI training, and scientific inference. Each section explains why NVL4 is relevant and highlights infrastructure requirements such as FP64 computing, HBM bandwidth, GPU communication, storage, networking, and AI acceleration.

NVIDIA is targeting Vera Rubin scientific systems at workloads that combine large numerical calculations, accelerated libraries, AI, and data processing. The exact application performance will still depend on software optimization, dataset size, precision requirements, and cluster design.

Scientific WorkloadWhy NVL4 Is RelevantMain Infrastructure Pressure
Climate modelingFP64 simulation plus AI-assisted forecastingCompute, memory, networking
Computational fluid dynamicsLarge numerical solvers and dense data movementFP64, HBM bandwidth, fabric
Quantum chemistryHigh-precision calculations plus AI modelsFP64, memory, GPU communication
Materials scienceSimulation, surrogate modeling, AI analysisCompute, storage, networking
Energy explorationSimulation and large-scale data analyticsGPU compute, storage throughput
PhysicsNumerical modeling and data-intensive analysisFP64, fabric, parallel storage
Scientific AI trainingFoundation and surrogate model trainingAI compute, memory, networking
Scientific inferenceRapid AI evaluation of scientific dataGPU throughput, storage, latency

NVIDIA has also announced Vera Rubin systems for future supercomputers supporting molecular dynamics, high-energy physics, fusion energy, materials science, drug discovery, astronomy, earth science, and related fields.

NVIDIA Vera Rubin NVL4 vs NVL72: What’s the Difference?

The main difference is deployment scale and workload emphasis.

NVL4 centers on four Rubin GPUs and two Vera CPUs in a dense scientific computing configuration. Vera Rubin NVL72 integrates 72 Rubin GPUs and 36 Vera CPUs into a rack-scale system designed for massive AI training, reasoning, and inference.

FeatureVera Rubin NVL4Vera Rubin NVL72
Rubin GPUs472
Vera CPUs236
Primary emphasisHPC and AI-for-ScienceRack-scale frontier AI
GPU scale-upNVLink 6 within four-GPU design72-GPU NVLink 6 domain
Server modelMGX-compatible modular serverFull rack-scale MGX architecture
Typical buyerHPC center, research lab, scientific computing teamHyperscaler, neocloud, frontier AI operator
Scaling approachBuild clusters from dense HPC nodesBegin with a large rack-scale GPU domain

NVL4 should therefore not be described simply as a “smaller NVL72.” Its four-GPU topology specifically gives HPC and AI-for-Science buyers a modular building block that can scale through conventional supercomputing cluster architecture.

The complete NVL72 hardware architecture provides a useful contrast because NVL72 places much more of the scale-up communication inside one rack.

Vera Rubin NVL4 vs Grace Hopper: What Improved?

Grace Hopper combined NVIDIA Grace CPU technology with Hopper GPUs and established a tightly coupled CPU-GPU architecture for AI and HPC. Vera Rubin advances that idea with Rubin GPUs, Vera CPUs, HBM4, newer NVLink technology, and substantially more AI and scientific computing capability.

AreaGrace Hopper GenerationVera Rubin NVL4 Direction
GPU generationHopperRubin
Host CPUGraceVera
GPU memory generationHopper-era HBMHBM4
CPU-GPU connectionEarlier NVLink-C2C generationSecond-generation NVLink-C2C
GPU scale-upEarlier NVLink generationNVLink 6
Scientific simulationBaselineNVIDIA claims up to 4× performance
AI-for-Science trainingBaselineNVIDIA claims up to 6× performance
AI-for-Science inferenceBaselineNVIDIA claims up to 8× performance

NVIDIA makes those 4×, 6×, and 8× claims against Grace Hopper on its Vera Rubin product page. They are vendor performance claims, not independent benchmark results, and actual gains will depend on the application, software, precision, system design, and test methodology.

Teams maintaining Hopper-generation infrastructure can use Catalyst’s H100 SXM review as a reference point for how the previous GPU generation fits into complete servers and data center deployments.

Where Does NVL4 Fit in HPC Infrastructure?

A Vera Rubin NVL4 node is only the compute layer. A production HPC environment also needs scale-out networking, high-throughput storage, rack-level power distribution, liquid cooling, management, optics, cabling, and software that can use the hardware efficiently.

Servers and OEM Systems

NVIDIA says NVL4 supports liquid-cooled MGX modular servers. OEM implementations convert the reference architecture into systems that organizations can actually rack, network, power, cool, manage, and support.

Dell has announced the PowerEdge XE8812 around Vera Rubin NVL4 for HPC and AI, with configurations scaling to as many as 144 GPUs per rack. Dell describes the system as fanless and direct-liquid-cooled for demanding scientific workloads.

Supermicro has also introduced a 1U NVL4 implementation with four Rubin GPUs, two Vera CPUs, ConnectX-9 networking, BlueField-4 connectivity, and direct liquid cooling. That is an OEM configuration example, not a mandatory specification for every NVL4 system.

Networking and Cluster Fabric

Once multiple NVL4 servers form a cluster, traffic must move outside the four-GPU NVLink domain. This scale-out layer can become critical for distributed simulation, MPI workloads, AI training, shared storage, and multi-node scientific applications.

NVIDIA’s Vera Rubin ecosystem includes ConnectX-9 SuperNICs and high-speed scale-out networking. Supermicro’s published NVL4 HPC blueprint, for example, pairs NVL4 clusters with NVIDIA Quantum-X800 InfiniBand.

The correct network depends on node count, communication pattern, oversubscription, latency targets, storage traffic, and software. Catalyst’s discussion of AI networking challenges provides additional context for why accelerator performance and fabric design must be evaluated together.

Storage

Scientific computing often generates large simulation results, checkpoints, training datasets, and instrument data. Storage therefore needs enough throughput and concurrency to keep compute nodes working rather than waiting for files or checkpoint operations.

There is no single storage array that every NVL4 deployment must use. Buyers should verify protocol, network connectivity, throughput, latency, availability, checkpoint behavior, application requirements, and GPU validation before defining the storage layer.

Liquid Cooling and Power

NVIDIA specifically describes NVL4 as compatible with liquid-cooled MGX servers. OEM implementations such as Dell XE8812 and Supermicro’s NVL4 systems also emphasize direct liquid cooling.

Infrastructure teams therefore need to evaluate CDU capacity, facility-water conditions, manifolds, redundancy, rack power, heat rejection, serviceability, and retrofit requirements. The broader liquid-cooling strategies become part of the server decision, not a separate facilities question.

What Should Be Bought With Vera Rubin NVL4?

A buyer should plan around the full HPC node and cluster rather than requesting four Rubin GPUs without the surrounding architecture.

  • Server platform: A validated liquid-cooled MGX or OEM NVL4 system with the required Vera CPU, local NVMe, management, and power design.
  • Cluster fabric: Appropriate SuperNICs, switches, optics, and cables for InfiniBand or Ethernet scale-out.
  • Storage: High-throughput storage sized for datasets, checkpoints, simulation output, and AI training pipelines.
  • Facility infrastructure: Racks, power distribution, CDUs, manifolds, monitoring, and redundant cooling capacity.

Compatibility should be verified from OEM documentation before ordering. Similar port types or physical form factors do not prove that servers, adapters, optics, cables, firmware, and network platforms form a validated configuration.

Who Is NVIDIA Vera Rubin NVL4 For?

NVL4 is most relevant to organizations that need both scientific computing and modern AI acceleration.

  • National laboratories and major university research centers.
  • HPC centers running simulation, modeling, and scientific AI.
  • Industrial R&D organizations in engineering, energy, chemistry, or materials.
  • Research teams building scientific foundation models or simulation-plus-AI workflows.

NVIDIA’s announced Vera Rubin supercomputing deployments reinforce this positioning. The company has identified future systems supporting national laboratories, energy research, earth science, materials science, fusion, molecular dynamics, and other research workloads.

Who Might Not Need Vera Rubin NVL4?

Not every AI or scientific workload requires an NVL4 system. Teams running modest inference, smaller machine-learning models, lightly accelerated analytics, or applications without strong FP64 and multi-GPU demands may obtain better economics from less dense hardware.

Existing Hopper or Blackwell systems may also remain appropriate when they already meet performance, software, power, cooling, budget, and deployment requirements. The correct platform depends on the workload rather than the age of the GPU generation.

What Should Buyers Verify Before Planning an NVL4 Deployment?

Before treating NVL4 as a procurement project, buyers should confirm the complete configuration with the server manufacturer and infrastructure team.

  • Verify the exact OEM server, Rubin GPU configuration, Vera CPU design, firmware, and supported software stack.
  • Define ConnectX/SuperNIC, switch, optics, cabling, and cluster topology requirements.
  • Validate storage performance against simulation I/O, checkpoints, datasets, and AI pipelines.
  • Confirm rack power, CDU capacity, facility water, redundancy, and deployment schedule.

Availability also needs verification. NVIDIA announced in June 2026 that NVL4-based systems were expected to become available from global system manufacturers in Q4 2026, so announced systems should not automatically be described as generally available before the OEM confirms orderability and shipping status.

How Can Catalyst Support a Vera Rubin NVL4 HPC Deployment?

Photorealistic image of HPC infrastructure experts supporting a data center deployment, showing engineers reviewing an open server rack with networking equipment, cables, and compute hardware. One specialist explains system components while another observes with a laptop nearby, surrounded by rows of enterprise server cabinets in a modern high-performance computing facility.

Catalyst Data Solutions works across OEM, channel, and distribution ecosystems to help organizations source AI, HPC, and data center infrastructure. For an NVL4 project, that can include the server platform, networking, storage, optics, cabling, rack infrastructure, power, and supporting cooling components.

Because frontier hardware availability can change quickly, the configuration should reflect workload, software requirements, facility limits, budget, existing infrastructure, and deployment timeline rather than assuming one vendor or architecture fits every project.

Buyers can request current availability or use the same request to ask about lead time, request a configuration quote, or build a complete AI infrastructure bundle.

A useful request should include the target scientific workloads, node and GPU count, storage capacity and throughput, network fabric, rack power, cooling architecture, software environment, support requirements, and expected deployment date.

Frequently Asked Questions

What is NVIDIA Vera Rubin NVL4?

NVIDIA Vera Rubin NVL4 is a liquid-cooled HPC and AI-for-Science platform that combines four Rubin GPUs with two Vera CPUs. The GPUs communicate through NVLink 6, while NVLink-C2C provides high-speed coherent communication between Vera CPUs and Rubin GPUs.

What does NVL4 mean?

NVIDIA uses NVL4 to identify the four-Rubin-GPU NVLink-connected configuration. It distinguishes this dense server-level platform from larger architectures such as Vera Rubin NVL72, which uses 72 Rubin GPUs.

How many GPUs are in Vera Rubin NVL4?

Vera Rubin NVL4 contains four NVIDIA Rubin GPUs. NVIDIA connects those GPUs through a second-generation NVLink bridge running sixth-generation NVLink.

Does Vera Rubin NVL4 use Vera CPUs?

Yes. NVIDIA Vera Rubin NVL4 pairs its four Rubin GPUs with two NVIDIA Vera CPUs connected through NVLink-C2C. Vera handles host compute, data movement, orchestration, and other CPU-intensive work around accelerated applications.

Does Vera Rubin NVL4 support FP64?

Yes. NVIDIA lists 33 TFLOPS of native FP64 performance per Rubin GPU in its preliminary specifications. That capability makes Rubin relevant to scientific simulations that require high numerical precision.

What is NVLink-C2C?

NVLink-C2C is NVIDIA’s coherent CPU-to-GPU interconnect. With Vera and NVIDIA GPUs, NVIDIA states that it provides up to 1.8 TB/s of coherent bandwidth, helping CPUs and GPUs share data with fewer transfer bottlenecks.

Is Vera Rubin NVL4 designed for HPC?

Yes. NVIDIA specifically positions NVL4 for scientific computing and AI-for-Science, including workloads such as climate modeling, computational fluid dynamics, quantum chemistry, and energy exploration.

NVL4 vs NVL72: which is better for HPC?

NVL4 is more specifically positioned as a dense building block for HPC and AI-for-Science, while NVL72 is a much larger 72-GPU rack-scale system aimed primarily at frontier AI. The better choice depends on application scale, communication patterns, software, facility capacity, and budget.

Research and Fact-Check Record

Research Date: August 15, 2026

Official Datasheet: 

https://dam-cdn.nvd.orangelogic.com/AssetLink/v5rf2icnf86o26e464tf6djn23r8ibhe.pdf

https://shop.catalystdatasolutionsinc.com/wp-content/uploads/2026/08/gpu-architecture-datasheet-vera-rubin-nvidia-us-5198950-web.pdf

Research Key Point 1: NVIDIA Vera Rubin NVL4 combines four Rubin GPUs and two Vera CPUs, connected through NVLink 6 and NVLink-C2C in liquid-cooled MGX-compatible systems.

Research Key Point 2: NVIDIA claims up to 4× scientific simulation, 6× AI-for-Science training, and 8× inference performance versus a four-Grace-Hopper-Superchip GH200 configuration. These remain NVIDIA projected performance claims.

Research Key Point 3: NVIDIA stated that Vera Rubin NVL4 systems from global manufacturers were expected in Q4 2026, with Bull, Dell, GIGABYTE, HPE, and Supermicro among the announced system partners.

OEM Availability Note: OEM timelines vary. Dell lists the PowerEdge XE8812 for global availability in early 2027, while Supermicro has announced its NVL4-based ARS-123GL-NRB-ALC without a specific general-availability date in the reviewed official materials.