NVIDIA Vera Rubin NVL4 is built for a growing challenge in scientific computing: researchers increasingly need one platform that can handle traditional simulation, AI training, inference, and data-heavy analysis without treating each workload as a separate system.
The platform combines four NVIDIA Rubin GPUs with two NVIDIA Vera CPUs, linked through NVLink 6 and NVLink-C2C. That design targets demanding HPC and AI-for-Science workloads where FP64 performance, memory bandwidth, and fast CPU-to-GPU communication can directly affect how quickly simulations and scientific models run.
For organizations evaluating next-generation HPC infrastructure, the important question is not simply how fast Rubin is. It is how NVL4 fits into a complete environment that includes the server, networking, storage, liquid cooling, power, and cluster architecture needed to keep those accelerators productive.
This guide explains the NVIDIA Vera Rubin NVL4 architecture, specifications, HPC advantages, AI-for-Science use cases, interconnects, deployment requirements, and differences from NVL72 and Grace Hopper, using official NVIDIA specifications and clearly distinguishing announced performance claims from independently verified results.
What Is NVIDIA Vera Rubin NVL4?
NVIDIA Vera Rubin NVL4 is a four-GPU accelerated computing platform built specifically around the convergence of HPC and AI. NVIDIA connects four Rubin GPUs to two Vera CPUs and supports the design in liquid-cooled NVIDIA MGX modular servers.
The reason NVL4 exists is different from the reason NVIDIA offers massive NVL72 systems. NVL4 gives scientific computing organizations a dense node for simulation, modeling, AI training, AI inference, and hybrid workflows without requiring a 72-GPU NVLink rack.
That positioning also makes NVL4 relevant to universities, national laboratories, industrial R&D groups, energy companies, and HPC centers. NVIDIA specifically identifies climate modeling, computational fluid dynamics, quantum chemistry, and energy exploration among the workloads targeted by Vera Rubin scientific computing systems.
How Does Vera Rubin NVL4 Work?

The Vera Rubin NVL4 architecture combines CPU compute, GPU acceleration, high-bandwidth memory, and two types of NVIDIA interconnect technology.
Four Rubin GPUs communicate through a second-generation NVLink bridge running NVLink 6. Two Vera CPUs connect to the GPUs through NVLink-C2C, giving CPU and GPU memory a high-speed coherent communication path.
| Architecture Layer | NVIDIA Vera Rubin NVL4 Role |
| Accelerator | 4 NVIDIA Rubin GPUs |
| Host compute | 2 NVIDIA Vera CPUs |
| GPU-to-GPU scale-up | Sixth-generation NVLink |
| CPU-to-GPU connection | NVLink-C2C |
| GPU memory | HBM4 |
| Server architecture | Liquid-cooled NVIDIA MGX-compatible systems |
| Cluster networking | High-speed InfiniBand or Ethernet, depending on implementation |
| Cooling | Direct liquid cooling |
This relationship matters more than any one component. A scientific application may use the CPU for orchestration and data preparation, send intensive numerical or AI operations to the GPUs, exchange intermediate data repeatedly, and then write large outputs to external storage.
That is why an NVL4 deployment should be treated as a complete GPU server design problem rather than simply a four-GPU purchase.
What Hardware Is Inside Vera Rubin NVL4?
The core NVL4 design includes four Rubin GPUs and two 88-core Vera CPUs. NVIDIA says the GPUs use HBM4, while Vera combines custom Olympus CPU cores with high-bandwidth LPDDR5X memory.
The main hardware relationships are:
- Four Rubin GPUs: Handle FP64 scientific computing as well as lower-precision AI training and inference.
- Two Vera CPUs: Provide host compute, data movement, orchestration, preprocessing, and CPU-heavy scientific work.
- NVLink 6 and NVLink-C2C: Connect GPUs to each other and CPUs to GPU compute with high bandwidth.
- MGX and liquid cooling: Place the components into deployable high-density server designs.
NVIDIA says Vera provides up to 1.8 TB/s of coherent NVLink-C2C bandwidth between Vera CPUs and NVIDIA GPUs. The company also lists up to 1.5 TB of LPDDR5X capacity and 1.2 TB/s of memory bandwidth for Vera CPU configurations.
Actual server memory, networking, local storage, and power configuration can vary by OEM implementation. This distinction matters when comparing an NVIDIA platform architecture with a finished Dell, Supermicro, or other manufacturer’s system.
For broader platform context, Catalyst’s DGX and HGX guide explains how NVIDIA accelerator architectures become complete data center systems.
What Are NVIDIA Vera Rubin NVL4 Specs?
NVIDIA’s published Rubin figures remain preliminary and subject to change. NVIDIA currently publishes detailed Rubin GPU and Vera Rubin Superchip specifications, while some NVL4 totals below are simple calculations from four published Rubin GPU figures rather than a separate final NVL4 datasheet value.
| Specification | Vera Rubin NVL4 |
| GPU configuration | 4 NVIDIA Rubin GPUs |
| CPU configuration | 2 NVIDIA Vera CPUs |
| CPU cores | 88 Olympus cores per Vera CPU; 176 across two CPUs |
| GPU memory | 288 GB HBM4 per Rubin GPU; about 1.152 TB across 4 GPUs* |
| GPU memory bandwidth | 22 TB/s per GPU; about 88 TB/s across 4 GPUs* |
| Native FP64 | 33 TFLOPS per GPU; about 132 TFLOPS across 4 GPUs* |
| NVFP4 inference | 50 PFLOPS per GPU; about 200 PFLOPS across 4 GPUs* |
| NVFP4 training | 35 PFLOPS per GPU; about 140 PFLOPS across 4 GPUs* |
| GPU interconnect | Sixth-generation NVIDIA NVLink |
| NVLink bandwidth | Up to 3.6 TB/s per Rubin GPU |
| CPU-GPU interconnect | NVIDIA NVLink-C2C |
| NVLink-C2C bandwidth | Up to 1.8 TB/s coherent bandwidth |
| Server support | Liquid-cooled NVIDIA MGX modular servers |
Calculated from NVIDIA’s preliminary per-GPU figures. Because NVIDIA rounds some published performance values, these calculated totals should not be treated as final OEM system specifications.
One important distinction involves FP64 numbers. NVIDIA lists 33 TFLOPS of native FP64 per Rubin GPU and separately lists higher DGEMM performance that uses Tensor Core-based emulation algorithms. Those values should not be presented as though they describe the same type of FP64 execution.
Why Does Vera Rubin NVL4 Matter for HPC?
Traditional HPC depends heavily on double-precision FP64 compute because scientific simulations often need more numerical accuracy than many AI workloads. Climate models, fluid simulations, chemistry calculations, and physics codes can accumulate errors when numerical precision is insufficient.
Rubin provides native FP64 capability while also supporting the AI performance required for emerging scientific machine learning. NVIDIA describes Vera Rubin as a single accelerated platform for numerical solvers, AI models, instrument data, and real-time analytics.
Three characteristics are especially important for HPC:
- FP64 performance supports high-precision numerical simulation.
- HBM4 bandwidth helps feed large scientific datasets to GPU compute.
- NVLink 6 reduces communication limits when multiple GPUs cooperate.
- NVLink-C2C improves data movement between host CPUs and GPUs.
The result is not simply a faster scientific computing GPU. It is a tightly connected CPU-GPU node intended to keep simulation, data movement, and accelerated analysis working together.
That systems approach matches how modern HPC GPU deployments must be planned: compute performance matters only when networking, storage, power, cooling, and application architecture can keep the accelerators productive.
Why Is NVL4 Important for AI-for-Science?
AI-for-Science uses machine learning alongside traditional scientific computing to speed up discovery, analyze complex data, approximate expensive calculations, and improve how researchers explore large design spaces.
Vera Rubin NVL4 matters because the same node provides native scientific compute and high-throughput AI acceleration. NVIDIA specifically positions Vera Rubin for combining simulation, AI, and data-intensive research.
Key AI-for-Science workflows include:
- Scientific foundation models: Train models on scientific data across chemistry, biology, physics, climate, or materials.
- Surrogate models: Use AI to approximate expensive simulations and evaluate possibilities faster.
- Simulation plus AI: Combine numerical solvers with learned models inside one research workflow.
- AI-assisted analysis: Apply training and inference to simulation output, instrument data, and scientific datasets.
A researcher might first run a physics-based simulation, use a neural model to approximate part of the solution space, evaluate the AI prediction, and then return selected cases to a higher-fidelity simulation. That hybrid workflow places demands on CPU compute, GPU compute, memory, storage, and communication at the same time.
What Does NVLink-C2C Do in Vera Rubin NVL4?
NVLink-C2C is NVIDIA’s high-speed coherent connection between Vera CPUs and NVIDIA GPUs. NVIDIA states that the current Vera implementation provides up to 1.8 TB/s of coherent CPU-GPU bandwidth.
Coherency matters because tightly coupled HPC applications frequently move data between CPU-managed operations and GPU kernels. A faster coherent path can reduce the communication bottleneck that appears when CPUs repeatedly prepare, inspect, transform, or coordinate data used by GPUs.
NVLink-C2C serves a different purpose from NVLink 6. NVLink-C2C connects CPU and GPU resources, while sixth-generation NVLink provides the high-bandwidth GPU-to-GPU communication needed inside the NVL4 scale-up domain.
What Scientific Workloads Can Vera Rubin NVL4 Run?

NVIDIA is targeting Vera Rubin scientific systems at workloads that combine large numerical calculations, accelerated libraries, AI, and data processing. The exact application performance will still depend on software optimization, dataset size, precision requirements, and cluster design.
| Scientific Workload | Why NVL4 Is Relevant | Main Infrastructure Pressure |
| Climate modeling | FP64 simulation plus AI-assisted forecasting | Compute, memory, networking |
| Computational fluid dynamics | Large numerical solvers and dense data movement | FP64, HBM bandwidth, fabric |
| Quantum chemistry | High-precision calculations plus AI models | FP64, memory, GPU communication |
| Materials science | Simulation, surrogate modeling, AI analysis | Compute, storage, networking |
| Energy exploration | Simulation and large-scale data analytics | GPU compute, storage throughput |
| Physics | Numerical modeling and data-intensive analysis | FP64, fabric, parallel storage |
| Scientific AI training | Foundation and surrogate model training | AI compute, memory, networking |
| Scientific inference | Rapid AI evaluation of scientific data | GPU throughput, storage, latency |
NVIDIA has also announced Vera Rubin systems for future supercomputers supporting molecular dynamics, high-energy physics, fusion energy, materials science, drug discovery, astronomy, earth science, and related fields.
NVIDIA Vera Rubin NVL4 vs NVL72: What’s the Difference?
The main difference is deployment scale and workload emphasis.
NVL4 centers on four Rubin GPUs and two Vera CPUs in a dense scientific computing configuration. Vera Rubin NVL72 integrates 72 Rubin GPUs and 36 Vera CPUs into a rack-scale system designed for massive AI training, reasoning, and inference.
| Feature | Vera Rubin NVL4 | Vera Rubin NVL72 |
| Rubin GPUs | 4 | 72 |
| Vera CPUs | 2 | 36 |
| Primary emphasis | HPC and AI-for-Science | Rack-scale frontier AI |
| GPU scale-up | NVLink 6 within four-GPU design | 72-GPU NVLink 6 domain |
| Server model | MGX-compatible modular server | Full rack-scale MGX architecture |
| Typical buyer | HPC center, research lab, scientific computing team | Hyperscaler, neocloud, frontier AI operator |
| Scaling approach | Build clusters from dense HPC nodes | Begin with a large rack-scale GPU domain |
NVL4 should therefore not be described simply as a “smaller NVL72.” Its four-GPU topology specifically gives HPC and AI-for-Science buyers a modular building block that can scale through conventional supercomputing cluster architecture.
The complete NVL72 hardware architecture provides a useful contrast because NVL72 places much more of the scale-up communication inside one rack.
Vera Rubin NVL4 vs Grace Hopper: What Improved?
Grace Hopper combined NVIDIA Grace CPU technology with Hopper GPUs and established a tightly coupled CPU-GPU architecture for AI and HPC. Vera Rubin advances that idea with Rubin GPUs, Vera CPUs, HBM4, newer NVLink technology, and substantially more AI and scientific computing capability.
| Area | Grace Hopper Generation | Vera Rubin NVL4 Direction |
| GPU generation | Hopper | Rubin |
| Host CPU | Grace | Vera |
| GPU memory generation | Hopper-era HBM | HBM4 |
| CPU-GPU connection | Earlier NVLink-C2C generation | Second-generation NVLink-C2C |
| GPU scale-up | Earlier NVLink generation | NVLink 6 |
| Scientific simulation | Baseline | NVIDIA claims up to 4× performance |
| AI-for-Science training | Baseline | NVIDIA claims up to 6× performance |
| AI-for-Science inference | Baseline | NVIDIA claims up to 8× performance |
NVIDIA makes those 4×, 6×, and 8× claims against Grace Hopper on its Vera Rubin product page. They are vendor performance claims, not independent benchmark results, and actual gains will depend on the application, software, precision, system design, and test methodology.
Teams maintaining Hopper-generation infrastructure can use Catalyst’s H100 SXM review as a reference point for how the previous GPU generation fits into complete servers and data center deployments.
Where Does NVL4 Fit in HPC Infrastructure?

A Vera Rubin NVL4 node is only the compute layer. A production HPC environment also needs scale-out networking, high-throughput storage, rack-level power distribution, liquid cooling, management, optics, cabling, and software that can use the hardware efficiently.
Servers and OEM Systems
NVIDIA says NVL4 supports liquid-cooled MGX modular servers. OEM implementations convert the reference architecture into systems that organizations can actually rack, network, power, cool, manage, and support.
Dell has announced the PowerEdge XE8812 around Vera Rubin NVL4 for HPC and AI, with configurations scaling to as many as 144 GPUs per rack. Dell describes the system as fanless and direct-liquid-cooled for demanding scientific workloads.
Supermicro has also introduced a 1U NVL4 implementation with four Rubin GPUs, two Vera CPUs, ConnectX-9 networking, BlueField-4 connectivity, and direct liquid cooling. That is an OEM configuration example, not a mandatory specification for every NVL4 system.
Networking and Cluster Fabric
Once multiple NVL4 servers form a cluster, traffic must move outside the four-GPU NVLink domain. This scale-out layer can become critical for distributed simulation, MPI workloads, AI training, shared storage, and multi-node scientific applications.
NVIDIA’s Vera Rubin ecosystem includes ConnectX-9 SuperNICs and high-speed scale-out networking. Supermicro’s published NVL4 HPC blueprint, for example, pairs NVL4 clusters with NVIDIA Quantum-X800 InfiniBand.
The correct network depends on node count, communication pattern, oversubscription, latency targets, storage traffic, and software. Catalyst’s discussion of AI networking challenges provides additional context for why accelerator performance and fabric design must be evaluated together.
Storage
Scientific computing often generates large simulation results, checkpoints, training datasets, and instrument data. Storage therefore needs enough throughput and concurrency to keep compute nodes working rather than waiting for files or checkpoint operations.
There is no single storage array that every NVL4 deployment must use. Buyers should verify protocol, network connectivity, throughput, latency, availability, checkpoint behavior, application requirements, and GPU validation before defining the storage layer.
Liquid Cooling and Power
NVIDIA specifically describes NVL4 as compatible with liquid-cooled MGX servers. OEM implementations such as Dell XE8812 and Supermicro’s NVL4 systems also emphasize direct liquid cooling.
Infrastructure teams therefore need to evaluate CDU capacity, facility-water conditions, manifolds, redundancy, rack power, heat rejection, serviceability, and retrofit requirements. The broader liquid-cooling strategies become part of the server decision, not a separate facilities question.
What Should Be Bought With Vera Rubin NVL4?
A buyer should plan around the full HPC node and cluster rather than requesting four Rubin GPUs without the surrounding architecture.
- Server platform: A validated liquid-cooled MGX or OEM NVL4 system with the required Vera CPU, local NVMe, management, and power design.
- Cluster fabric: Appropriate SuperNICs, switches, optics, and cables for InfiniBand or Ethernet scale-out.
- Storage: High-throughput storage sized for datasets, checkpoints, simulation output, and AI training pipelines.
- Facility infrastructure: Racks, power distribution, CDUs, manifolds, monitoring, and redundant cooling capacity.
Compatibility should be verified from OEM documentation before ordering. Similar port types or physical form factors do not prove that servers, adapters, optics, cables, firmware, and network platforms form a validated configuration.
Who Is NVIDIA Vera Rubin NVL4 For?
NVL4 is most relevant to organizations that need both scientific computing and modern AI acceleration.
- National laboratories and major university research centers.
- HPC centers running simulation, modeling, and scientific AI.
- Industrial R&D organizations in engineering, energy, chemistry, or materials.
- Research teams building scientific foundation models or simulation-plus-AI workflows.
NVIDIA’s announced Vera Rubin supercomputing deployments reinforce this positioning. The company has identified future systems supporting national laboratories, energy research, earth science, materials science, fusion, molecular dynamics, and other research workloads.
Who Might Not Need Vera Rubin NVL4?
Not every AI or scientific workload requires an NVL4 system. Teams running modest inference, smaller machine-learning models, lightly accelerated analytics, or applications without strong FP64 and multi-GPU demands may obtain better economics from less dense hardware.
Existing Hopper or Blackwell systems may also remain appropriate when they already meet performance, software, power, cooling, budget, and deployment requirements. The correct platform depends on the workload rather than the age of the GPU generation.
What Should Buyers Verify Before Planning an NVL4 Deployment?
Before treating NVL4 as a procurement project, buyers should confirm the complete configuration with the server manufacturer and infrastructure team.
- Verify the exact OEM server, Rubin GPU configuration, Vera CPU design, firmware, and supported software stack.
- Define ConnectX/SuperNIC, switch, optics, cabling, and cluster topology requirements.
- Validate storage performance against simulation I/O, checkpoints, datasets, and AI pipelines.
- Confirm rack power, CDU capacity, facility water, redundancy, and deployment schedule.
Availability also needs verification. NVIDIA announced in June 2026 that NVL4-based systems were expected to become available from global system manufacturers in Q4 2026, so announced systems should not automatically be described as generally available before the OEM confirms orderability and shipping status.
How Can Catalyst Support a Vera Rubin NVL4 HPC Deployment?

Catalyst Data Solutions works across OEM, channel, and distribution ecosystems to help organizations source AI, HPC, and data center infrastructure. For an NVL4 project, that can include the server platform, networking, storage, optics, cabling, rack infrastructure, power, and supporting cooling components.
Because frontier hardware availability can change quickly, the configuration should reflect workload, software requirements, facility limits, budget, existing infrastructure, and deployment timeline rather than assuming one vendor or architecture fits every project.
Buyers can request current availability or use the same request to ask about lead time, request a configuration quote, or build a complete AI infrastructure bundle.
A useful request should include the target scientific workloads, node and GPU count, storage capacity and throughput, network fabric, rack power, cooling architecture, software environment, support requirements, and expected deployment date.
Frequently Asked Questions
What is NVIDIA Vera Rubin NVL4?
NVIDIA Vera Rubin NVL4 is a liquid-cooled HPC and AI-for-Science platform that combines four Rubin GPUs with two Vera CPUs. The GPUs communicate through NVLink 6, while NVLink-C2C provides high-speed coherent communication between Vera CPUs and Rubin GPUs.
What does NVL4 mean?
NVIDIA uses NVL4 to identify the four-Rubin-GPU NVLink-connected configuration. It distinguishes this dense server-level platform from larger architectures such as Vera Rubin NVL72, which uses 72 Rubin GPUs.
How many GPUs are in Vera Rubin NVL4?
Vera Rubin NVL4 contains four NVIDIA Rubin GPUs. NVIDIA connects those GPUs through a second-generation NVLink bridge running sixth-generation NVLink.
Does Vera Rubin NVL4 use Vera CPUs?
Yes. NVIDIA Vera Rubin NVL4 pairs its four Rubin GPUs with two NVIDIA Vera CPUs connected through NVLink-C2C. Vera handles host compute, data movement, orchestration, and other CPU-intensive work around accelerated applications.
Does Vera Rubin NVL4 support FP64?
Yes. NVIDIA lists 33 TFLOPS of native FP64 performance per Rubin GPU in its preliminary specifications. That capability makes Rubin relevant to scientific simulations that require high numerical precision.
What is NVLink-C2C?
NVLink-C2C is NVIDIA’s coherent CPU-to-GPU interconnect. With Vera and NVIDIA GPUs, NVIDIA states that it provides up to 1.8 TB/s of coherent bandwidth, helping CPUs and GPUs share data with fewer transfer bottlenecks.
Is Vera Rubin NVL4 designed for HPC?
Yes. NVIDIA specifically positions NVL4 for scientific computing and AI-for-Science, including workloads such as climate modeling, computational fluid dynamics, quantum chemistry, and energy exploration.
NVL4 vs NVL72: which is better for HPC?
NVL4 is more specifically positioned as a dense building block for HPC and AI-for-Science, while NVL72 is a much larger 72-GPU rack-scale system aimed primarily at frontier AI. The better choice depends on application scale, communication patterns, software, facility capacity, and budget.
Research and Fact-Check Record
Research Date: August 15, 2026
Official Datasheet:Â
https://dam-cdn.nvd.orangelogic.com/AssetLink/v5rf2icnf86o26e464tf6djn23r8ibhe.pdf
Research Key Point 1: NVIDIA Vera Rubin NVL4 combines four Rubin GPUs and two Vera CPUs, connected through NVLink 6 and NVLink-C2C in liquid-cooled MGX-compatible systems.
Research Key Point 2: NVIDIA claims up to 4× scientific simulation, 6× AI-for-Science training, and 8× inference performance versus a four-Grace-Hopper-Superchip GH200 configuration. These remain NVIDIA projected performance claims.
Research Key Point 3: NVIDIA stated that Vera Rubin NVL4 systems from global manufacturers were expected in Q4 2026, with Bull, Dell, GIGABYTE, HPE, and Supermicro among the announced system partners.
OEM Availability Note: OEM timelines vary. Dell lists the PowerEdge XE8812 for global availability in early 2027, while Supermicro has announced its NVL4-based ARS-123GL-NRB-ALC without a specific general-availability date in the reviewed official materials.