Supermicro Vera Rubin Systems Explained: NVL72 vs HGX NVL8 vs NVL4

Picture of Sophan Pheng

Sophan Pheng

VP of Sales & Product Management

Supermicro Vera Rubin systems take NVIDIA’s Rubin architecture from individual accelerators to deployable AI and HPC infrastructure. The portfolio spans rack-scale NVL72, modular HGX Rubin NVL8 servers, and Vera Rubin NVL4 systems designed around scientific computing and AI. The underlying Rubin GPU architecture provides the accelerator foundation, while Supermicro adds servers, racks, networking, cooling, power, storage, and deployment engineering.

The quick difference is scale and workload fit. Vera Rubin NVL72 combines 72 Rubin GPUs and 36 Vera CPUs as one rack-scale AI system. HGX Rubin NVL8 packages eight Rubin GPUs into a flexible 2U server with several CPU choices. NVL4 combines four Rubin GPUs with Vera CPUs for HPC, scientific simulation, and AI-for-science workloads.

What Are Supermicro Vera Rubin Systems?

Infographic titled “What Are Supermicro Vera Rubin Systems?” showing a Supermicro server and three system configurations: NVL72 with 72 GPUs for rack-scale frontier AI, HGX NVL8 with 8 GPUs in a 2U system for enterprise training and inference, and NVL4 with 4 GPUs per node for HPC and AI research.

Supermicro Vera Rubin systems are server, rack, and data-center implementations of NVIDIA’s Vera Rubin architecture. Supermicro does not make a competing “Rubin GPU.” NVIDIA supplies the Rubin accelerator and Vera platform technologies; Supermicro integrates them into usable infrastructure. That distinction is a core requirement of Catalyst’s content strategy.

The portfolio covers several deployment scales instead of forcing every buyer into the same architecture. NVIDIA itself positions Vera Rubin as a platform that treats the data center as the unit of compute, with products ranging from NVL4 and eight-GPU HGX systems to the rack-scale NVL72.

Supermicro then surrounds those compute platforms with DCBBS infrastructure. That can include storage, networking, racks, power distribution, cooling equipment, cabling, site planning, and SuperCloud management rather than just a GPU chassis.

For buyers comparing today’s architecture with the complete rack design, the NVL72 hardware architecture provides useful context on how Vera, Rubin, NVLink, and the network layers fit together.

PlatformBasic scaleMain roleTypical buyer
Vera Rubin NVL7272 GPUs, rack-scaleFrontier AI factoriesHyperscalers, neoclouds, sovereign AI
HGX Rubin NVL88 GPUs, 2UFlexible training and inferenceEnterprises, service providers, AI clusters
Vera Rubin NVL44 GPUs per nodeHPC + AI convergenceResearch centers, universities, HPC sites

Vera Rubin NVL72: Built for Rack-Scale AI

Vera Rubin NVL72 rack-scale AI system in a modern data center, showing a central server rack filled with networking and compute hardware, surrounded by rows of connected racks and blue cabling.

Vera Rubin NVL72 is Supermicro’s highest-scale Rubin configuration for buyers that want one tightly integrated rack-scale compute system. NVIDIA specifies 72 Rubin GPUs and 36 Vera CPUs connected through sixth-generation NVLink and NVLink-C2C.

Supermicro describes the NVL72 rack as a single rack-scale accelerator. Its design brings Rubin GPUs, Vera CPUs, NVLink 6, ConnectX-9 SuperNICs, BlueField-4 DPUs, and the scale-out network into one coordinated architecture rather than treating the GPUs as independent servers.

That distinction matters for very large models. NVLink 6 supplies the high-bandwidth scale-up fabric inside the system, while ConnectX-9 connects GPU resources to the wider cluster. Quantum-X800 InfiniBand or Spectrum-X Ethernet can then connect racks into larger AI factories.

Supermicro currently describes NVL72 as delivering 75 TB of fast memory and approximately 1.6 PB/s of HBM4 bandwidth. NVIDIA’s current preliminary specification table lists 20.7 TB of HBM4 GPU memory, 54 TB of CPU memory, and 1,580 TB/s of aggregate HBM4 bandwidth, which explains Supermicro’s rounded figures.

NVL72 primarily targets workloads that benefit from very large scale:

  • Frontier-model pretraining and post-training
  • Large reasoning and mixture-of-experts models
  • High-volume agentic and real-time inference
  • Large neocloud, hyperscale, or sovereign AI factories

The network surrounding that compute becomes as important as the GPUs themselves. Large NVL72 deployments must control east-west congestion, oversubscription, optics, and storage traffic, all of which shape AI data center networking at rack and cluster scale. 

NVL72 therefore is not simply “72 GPUs in a rack.” Buyers need compatible networking, checkpoint and dataset storage, high-density power distribution, direct liquid cooling, and facility capacity that can sustain the rack continuously.

HGX Rubin NVL8: The Flexible 8-GPU Option

HGX Rubin NVL8 8-GPU module shown against a black background, with eight large accelerator components arranged in two rows and connected by high-speed interconnect hardware.

HGX Rubin NVL8 is the modular choice in Supermicro’s Rubin portfolio. It places eight Rubin GPUs in a 2U liquid-cooled system while giving customers more freedom over host CPU architecture and deployment scale.

Supermicro says its NVL8 platform can support NVIDIA Vera CPUs as well as next-generation AMD and Intel x86 processors. That flexibility lets organizations match the CPU to software requirements, data-processing needs, existing standards, and workload characteristics rather than adopting a fixed host architecture.

NVIDIA also defines HGX Rubin NVL8 as an eight-Rubin-GPU platform connected through sixth-generation NVLink. HGX supports either Vera or x86 CPU baseboards, reinforcing its role as an adaptable system platform rather than a fixed turnkey appliance.

The architectural distinction between integrated NVIDIA systems and OEM-built HGX servers also matters. HGX gives server manufacturers more flexibility in CPU, storage, networking, and system design, which is an important part of the DGX and HGX comparison 

Supermicro says up to nine 2U HGX Rubin NVL8 systems can fit within a rack, creating a maximum of 72 Rubin GPUs while preserving server-level modularity. This does not make nine NVL8 nodes architecturally identical to one NVL72; their scale-up topology, CPU design, service model, and system boundaries remain different.

HGX NVL8 fits organizations that need:

  • Enterprise AI training and inference
  • Incremental cluster growth
  • CPU and platform flexibility
  • Eight-GPU server serviceability

That makes NVL8 especially useful when a full rack-scale NVL72 deployment would exceed workload needs, facility capacity, or purchasing timelines.

Vera Rubin NVL4: Where HPC and AI Converge

Vera Rubin NVL4 server with its chassis open, showing four large GPU modules and liquid-cooling tubes inside, positioned in a modern data center with server racks in the background.

Vera Rubin NVL4 targets environments where traditional high-performance computing and modern AI increasingly share the same infrastructure. NVIDIA connects four Rubin GPUs with two Vera CPUs through NVLink-C2C, creating a platform aimed at scientific simulation, AI-for-science, and accelerated HPC.

Supermicro’s current ARS-123GL-NRB-ALC implementation is a 1U liquid-cooled NVL4 system. The published specification lists four Rubin GPUs, two 88-core Vera CPUs, sixth-generation NVLink, 1,536 GB of onboard LPDDR5X memory, and high-speed ConnectX-9 networking.

The current system supports four 800Gb/s ConnectX-9 InfiniBand/Ethernet interfaces, with BlueField-4 connectivity options. That gives an HPC architect a clearer path from compute node to high-speed fabric instead of selecting the accelerator before considering the cluster network.

Supermicro also built an NVL4 DCBBS blueprint around native FP64 performance, which matters for double-precision numerical workloads. Its published use cases include scientific research and simulation alongside AI, while NVIDIA specifically highlights climate modeling, computational fluid dynamics, energy exploration, and other scientific workloads for Vera Rubin.

NVL4 therefore fits workloads such as:

  • Computational fluid dynamics and physics
  • Molecular and materials simulation
  • AI-assisted scientific discovery
  • Mixed FP64 simulation and AI workflows

At scale, Supermicro aligns the NVL4 blueprint with Quantum-X800 InfiniBand and its DLC-2 cooling architecture. This creates an HPC cluster design rather than merely a smaller version of an NVL72 AI factory.

NVL72 vs HGX NVL8 vs NVL4: Key Differences

The most important difference between the three Supermicro Vera Rubin systems is not simply GPU count. Each platform changes the system boundary, CPU relationship, networking model, cooling density, deployment scale, and workload profile.

FeatureVera Rubin NVL72HGX Rubin NVL8Vera Rubin NVL4
Compute scaleFull rack2U server1U Supermicro node
Rubin GPUs7284
Host CPU36 Vera CPUsVera, AMD or Intel options2 Vera CPUs
Main focusRack-scale AI factoryTraining and inferenceHPC + scientific AI
Scale-up fabricNVLink 6 rack architectureNVLink 6 within HGXNVLink 6 / NVLink-C2C
Scale-out directionSpectrum-X or Quantum-X800High-speed Ethernet/InfiniBandHPC fabric, including Quantum-X800
CoolingDirect liquid coolingDirect liquid cooling; L2A optionDirect liquid cooling
Best fitFrontier AIFlexible enterprise/service-provider AIResearch and HPC

CPU flexibility

NVL72 uses Vera as part of a tightly co-designed rack. Supermicro’s HGX NVL8 gives buyers the widest CPU choice, including Vera and next-generation x86 processors. NVL4 returns to a tightly connected Vera-plus-Rubin architecture because its goal includes scientific computing.

Networking

NVL72 needs both scale-up and scale-out networking. NVLink 6 joins accelerator resources inside the rack, while ConnectX-9, BlueField-4, Spectrum-X, and Quantum-X800 handle external communication and infrastructure functions.

NVL8 keeps the eight GPUs tightly connected but scales by adding separate servers. NVL4 emphasizes HPC-grade node-to-node communication, with Supermicro’s current system exposing 800G ConnectX-9 connectivity and its larger blueprint aligning with Quantum-X800.

Memory and workload behavior

NVL72 concentrates enormous aggregate GPU and CPU memory capacity around large AI workloads. NVL8 provides a smaller eight-GPU working set with host-memory characteristics influenced by CPU choice.

NVL4 puts fewer GPUs in each node but targets workloads where native FP64 computation and tight CPU-GPU communication can matter more than maximizing accelerator count.

Cooling and scalability

All three platforms move cooling from a server-level consideration to infrastructure planning. The difference is how far that planning extends: one NVL4 node, modular NVL8 racks, or complete NVL72 AI factories with dedicated cooling, power, storage, and network layers.

Which Supermicro Vera Rubin System Should You Choose?

The right system depends on workload and deployment constraints, not on which option has the largest GPU count.

ChooseWhen it makes senseMain reason
NVL72Frontier training, massive inference, sovereign AI, hyperscale AI factoriesMaximum rack-scale integration and accelerator density
HGX NVL8Enterprise AI, modular clusters, mixed deployment requirementsCPU flexibility and incremental scaling
NVL4HPC, FP64 simulation, research, AI-for-scienceCombines scientific computing with Rubin-generation AI

Choose NVL72 when the workload already justifies rack-scale infrastructure. An organization should have a plan for storage throughput, high-bandwidth scale-out networking, redundant cooling, and high-density power before treating NVL72 as a compute purchase.

Choose HGX NVL8 when flexibility matters more than building one large rack-scale accelerator. It lets infrastructure teams select a host architecture and expand cluster capacity in server-sized increments.

Choose NVL4 when scientific computing remains central. It gives research and HPC environments a Rubin platform without requiring them to design primarily around large LLM training.

Some enterprises may not need any of these frontier configurations immediately. Current Blackwell systems or smaller GPU servers can remain more practical when power, budget, software qualification, or workload scale does not justify Rubin-generation infrastructure. The broader GPU server deployment guide provides useful context for those sizing decisions.

Need Help Matching a Vera Rubin System to Your Infrastructure?

Catalyst Data Solutions works across OEM, channel, and distribution ecosystems to help organizations compare and source AI, HPC, and data-center infrastructure based on workload, power, cooling, networking, storage, budget, and deployment requirements.

For teams evaluating NVL72, HGX Rubin NVL8, or NVL4, Catalyst can help check current availability, compare platform options, and identify the supporting infrastructure needed to build a complete deployment.

What Should Be Bought With a Supermicro Vera Rubin System?

Supermicro Vera Rubin server in a blue-lit data center, with five infrastructure categories shown below: Compute (NVL72/NVL8/NVL4), Network (ConnectX-9/BlueField-4), Fabric (Spectrum-X/Quantum-X800), Storage (NVMe and context storage), and Facility (power, cooling, rack, and CDU).

A Rubin deployment needs more than the compute chassis. Catalyst’s strategy specifically requires accelerator and server articles to connect the CPU, NIC, fabric, storage, optics, cooling, power, and validated system architecture rather than assuming compatibility.

Infrastructure layerWhat buyers should validate
ComputeExact NVL72, NVL8, or NVL4 architecture and CPU configuration
NetworkConnectX-9, BlueField-4, Spectrum-X or Quantum-X800 topology
StorageDataset throughput, checkpointing, context storage, redundancy
Optics/cablingPort type, speed, reach, transceiver and cable compatibility
FacilityRack design, busbar/power distribution, CDU, heat rejection and redundancy

A practical NVL72 training deployment, for instance, could combine compute racks with Spectrum-X or Quantum-X800 fabric, high-performance NVMe storage for checkpoints, context-memory storage for inference workflows, and redundant DCBBS cooling. The exact configuration must come from validated OEM designs, not inferred port compatibility.

Cooling, Networking and Power at Rubin Scale

High-density Rubin infrastructure makes facility engineering part of the server decision. Supermicro uses its DLC-2 direct liquid-cooling stack across the Vera Rubin portfolio, including cold plates, manifolds, coolant distribution units, rear-door heat exchangers, and facility-side cooling components.

The data-center cooling strategies become especially relevant because the facility must remove heat continuously without allowing thermal limits to reduce accelerator utilization.

LayerRubin-scale consideration
Scale-upNVLink 6 handles high-bandwidth accelerator communication
Scale-outSpectrum-X Ethernet or Quantum-X800 InfiniBand links systems and racks
CoolingDirect-to-chip liquid cooling, CDUs and facility heat rejection
PowerHigh-density rack distribution, redundancy and available site capacity
StorageEnough throughput for datasets, checkpoints and context data

Supermicro’s current NVL72 blueprint sizes its DLC-2 design around approximately 227 kW per rack and includes four 110 kW power shelves per NVL72 rack for power delivery and redundancy. These figures describe infrastructure design capacity, not a promise that every deployment continuously consumes those exact values.

Supermicro also offers liquid-to-air options for facilities without an existing liquid-cooling loop. That may reduce retrofit barriers, but buyers still need to validate electrical capacity, heat rejection, rack loading, network cabling, floor layout, and redundancy before installation.

Supermicro’s Role in the Vera Rubin AI Factory

Supermicro’s role is system and data-center integration. NVIDIA provides the underlying Vera Rubin compute and networking technologies; Supermicro turns those components into servers, racks, cooling architectures, storage designs, networking configurations, and deployable facilities.

Its DCBBS framework can extend from a single system to multi-megawatt infrastructure. Supermicro’s June 2026 blueprints describe compute, high-performance and context-memory storage, Spectrum-X or Quantum-X800 networking, power distribution, liquid cooling, and site infrastructure within one deployment model.

SuperCloud software adds another layer. Supermicro says its management suite provides infrastructure control, deployment automation, developer tools, and multi-tenant GPU cloud management, which matters for neoclouds and shared AI environments.

DCBBS should therefore be understood as a deployment architecture, not as an accelerator technology. Catalyst’s source strategy makes the same distinction: Supermicro builds infrastructure around NVIDIA and AMD accelerator platforms rather than competing with their GPUs.

What These Three Systems Mean for AI Infrastructure

Vera Rubin shows how AI infrastructure is becoming more workload-specific. One GPU generation can now appear in architectures designed for very different operating models.

NVL72 treats rack-scale integration as the starting point for frontier AI. HGX NVL8 keeps the familiar multi-GPU server model while adding Rubin performance and greater host-CPU flexibility.

NVL4 takes another route. It uses Vera and Rubin to connect conventional HPC requirements, including native FP64 computing, with AI training, inference, and AI-for-science.

The result is a more useful buying framework than simply comparing GPU counts:

NVL72 maximizes rack-scale integration, NVL8 maximizes configuration flexibility, and NVL4 bridges AI with HPC.

That also means storage, networking, cooling, power, and software must follow the workload. The accelerator alone cannot determine the right architecture.

Conclusion

Supermicro Vera Rubin systems address three different infrastructure requirements. NVL72 targets frontier-scale AI factories with 72 Rubin GPUs and tightly integrated rack-scale networking. HGX Rubin NVL8 provides an eight-GPU, 2U platform for organizations that value CPU choice and modular growth. NVL4 combines four Rubin GPUs with Vera CPUs for HPC and scientific AI.

The best choice depends on workload scale, software requirements, network architecture, cooling, available power, storage throughput, and deployment timeline,not simply peak GPU density.

For infrastructure teams, the practical selection is straightforward: NVL72 for AI factories, NVL8 for flexible AI clusters, and NVL4 for HPC and scientific computing.

Frequently Asked Questions

1. When will Supermicro Vera Rubin systems be available?

Supermicro said in June 2026 that its NVL72 and HGX NVL8 DCBBS blueprints were available for customer engagements, with deployments planned for the second half of 2026. NVIDIA now describes Vera Rubin as in full production, but individual Supermicro SKU availability and lead times still require confirmation.

2. How much does a Supermicro Vera Rubin server cost?

The official sources reviewed do not provide one standard public price. Cost depends on the system, GPU count, CPU choice, networking, storage, rack architecture, liquid cooling, services, and deployment scale, so buyers should request a configuration-specific quote.

3. What CPU options are available for Supermicro HGX Rubin NVL8?

Supermicro says HGX Rubin NVL8 supports NVIDIA Vera CPUs as well as next-generation AMD and Intel x86 processors. The final supported processor models should be verified against the specific orderable Supermicro configuration before purchase.

4. What software stack supports Supermicro Vera Rubin systems?

Vera Rubin sits within NVIDIA’s accelerated computing software ecosystem, while exact OS, orchestration, drivers, and enterprise software support depend on the system configuration. Supermicro also provides SuperCloud management and automation tools for its infrastructure.

5. Can NVL72 be deployed in an existing data center?

Potentially, but the facility must meet its power, cooling, space, networking, and structural requirements. Supermicro has described liquid-to-air sidecar options for sites without an existing facility liquid loop, but a site assessment remains necessary before deployment.

6. What power requirements should organizations consider for Rubin systems?

Plan at the rack and facility level, not just from a server power-supply rating. Supermicro’s NVL72 DCBBS designs involve high-density rack power distribution and liquid cooling, so teams must validate electrical feeds, redundancy, heat rejection, CDU capacity, and expansion headroom.

7. How does Supermicro’s Rubin platform compare with NVIDIA DGX?

NVIDIA DGX provides a more standardized, NVIDIA-integrated system and software experience. Supermicro uses NVIDIA platforms such as HGX and Vera Rubin inside a broader configurable infrastructure model, giving buyers more choices around server architecture, racks, storage, cooling, and deployment design.

8. Can Supermicro Vera Rubin systems support multi-tenant AI clouds?

Yes, the architecture can support multi-tenant deployments when configured appropriately. Supermicro specifically describes DCBBS blueprints for single- or multi-tenant environments and says SuperCloud includes multi-tenant GPU cloud management capabilities.

Research & Fact-Check 

Research Date: September 3, 2026

Official Datasheet Version: https://www.supermicro.com/en/products/system/datasheet/ars-123gl-nrb-alc

Research Key Points:

  1. Supermicro currently positions NVL72, HGX Rubin NVL8, and NVL4 as different infrastructure designs rather than interchangeable GPU-count variants. NVL72 emphasizes rack-scale AI, NVL8 CPU flexibility, and NVL4 HPC/AI convergence.
  2. Supermicro’s June 2026 DCBBS architecture extends beyond compute into storage, Spectrum-X or Quantum-X800 networking, power distribution, liquid cooling, management, and site infrastructure.
  3. Supermicro should be described as the system/platform integrator around NVIDIA Vera Rubin, not as the supplier of a separate competing Rubin accelerator.