HPE Compute XD700: High-Density HGX Rubin NVL8 for Training and Inference

Picture of Swastik Lamsal

Swastik Lamsal

Director of Digital Strategy & Marketing
HPE Compute XD700: HGX Rubin NVL8 for AI Training

AI infrastructure is changing as large language models, reasoning systems, multimodal models, and AI agents demand more compute and memory. Enterprises are moving beyond isolated GPU servers toward dense systems where GPUs, CPUs, networking, storage, power, and cooling operate as one platform.

The HPE Compute XD700 brings that approach to HPE’s next-generation AI infrastructure. Announced as an Open Compute Project-inspired server built on NVIDIA HGX Rubin NVL8, it targets high-density AI training and inference deployments and is expected to become available in early 2027.

What is HPE Compute XD700? HPE Compute XD700 is an upcoming high-density AI server based on NVIDIA HGX Rubin NVL8. Each HGX platform connects eight Rubin GPUs through sixth-generation NVIDIA NVLink, creating a tightly coupled accelerator domain for demanding AI training, inference, and HPC workloads.

What Is HPE Compute XD700?

HPE Compute XD700 system highlighting high-performance compute, scalability, high-speed connectivity, AI analytics, data processing, cloud-scale infrastructure, and rack integration.

HPE Compute XD700 is a purpose-built HPE AI accelerator server designed around GPU-intensive workloads rather than conventional enterprise applications. HPE positions the platform for organizations building AI factories, large training clusters, and high-throughput inference environments.

Its core is NVIDIA HGX Rubin NVL8, which places eight Rubin GPUs into one high-bandwidth computing domain. NVIDIA designed HGX Rubin for generative AI, agentic AI, scientific computing, training, and inference workloads that depend heavily on fast GPU communication.

Understanding the accelerator itself is only one part of deployment planning. The underlying Rubin GPU architecture also changes memory bandwidth, GPU interconnect requirements, networking, power density, and cooling requirements.

XD700 characteristicCurrent announced information
Server classHigh-density AI/HPC server
GPU platformNVIDIA HGX Rubin NVL8
GPUs per HGX system8 NVIDIA Rubin GPUs
Primary workloadsTraining, inference, generative AI, HPC
AvailabilityTargeted for early 2027

HPE has not yet published a final production data sheet covering every XD700 CPU, storage, networking, power, and rack configuration. Buyers should therefore treat detailed pre-release specifications as planning information until HPE finalizes orderable configurations.

NVIDIA HGX Rubin NVL8 Architecture Explained

HGX Rubin NVL8 means eight NVIDIA Rubin GPUs operating as a tightly interconnected accelerator platform. NVIDIA combines the GPUs with NVLink, NVLink Switch technology, high-bandwidth memory, networking interfaces, and its AI software ecosystem.

Current preliminary NVIDIA specifications list 2.3 TB of aggregate HBM4 memory, 176 TB/s of aggregate GPU-memory bandwidth, and 28.8 TB/s of total NVLink Switch bandwidth for an eight-GPU Rubin NVL8 system. NVIDIA notes that these values remain subject to change.

HGX Rubin NVL8 componentPreliminary platform specification
Rubin GPUs8
Total GPU memory2.3 TB HBM4
Aggregate memory bandwidth176 TB/s
NVLinkSixth generation
NVLink Switch bandwidth28.8 TB/s

NVLink matters because an AI workload often cannot remain inside one GPU. Large models divide parameters, tensors, expert layers, or training batches across multiple accelerators, making communication performance almost as important as raw GPU compute.

The CPU remains responsible for host-side functions, orchestration, data preparation, and other system tasks. HGX Rubin NVL8 supports x86-based platforms, while NVIDIA also offers a Vera CPU configuration as HGX Vera Rubin NVL8.

For organizations evaluating CPU architecture alongside accelerator design, Vera CPU requirements provide useful context on how the host processor fits into next-generation agentic AI infrastructure.

How HPE XD700 Supports AI Training Workloads

HPE XD700 AI training architecture highlighting massive parallel compute, high-speed GPU connectivity, large model capacity, cluster-ready infrastructure, and optimized training performance.

Training a large AI model creates three major infrastructure demands: enormous parallel compute, enough accelerator memory to hold working model data, and extremely fast communication between GPUs and between servers.

HPE Compute XD700 addresses the first two requirements through eight-GPU HGX Rubin NVL8 nodes. Sixth-generation NVLink supports the intensive GPU-to-GPU data transfers involved in tensor parallelism, pipeline parallelism, mixture-of-experts models, and other distributed training techniques.

Training environments may use XD700 for workloads including:

  • Large language model pretraining and fine-tuning
  • Foundation and multimodal model development
  • Mixture-of-experts and reasoning models
  • Scientific AI and accelerated HPC

NVIDIA states that HGX Rubin NVL8 can match the training performance of HGX B200 using four times fewer GPUs in a specified projected mixture-of-experts workload. That figure represents NVIDIA’s projected workload comparison rather than a universal fourfold improvement across every model.

Scaling beyond one server also changes the purchasing decision. A production AI cluster needs a low-latency compute fabric, high-performance storage, optics, cabling, management networking, power distribution, and cooling alongside the GPU nodes.

The AI cluster networking stack becomes particularly important once training spans several servers because slow inter-node communication can leave expensive accelerators waiting for data.

HPE XD700 for AI Inference and Enterprise Deployment

Training focuses on building or modifying models. Inference focuses on running those models efficiently for users, applications, APIs, copilots, and AI agents.

That distinction changes how infrastructure gets optimized.

RequirementAI trainingAI inference
Main objectiveComplete training fasterServe requests efficiently
Key metricTraining throughputTokens, requests, and latency
GPU communicationExtremely importantDepends on model size
Memory demandModel + optimizer statesModel + KV/context data
Scaling priorityMaximum parallel computeThroughput and cost per request

Rubin’s large HBM4 capacity and memory bandwidth can benefit inference workloads where model weights and expanding context windows place heavy pressure on accelerator memory.

Agentic AI adds another challenge. One user request may trigger several models, tools, retrieval operations, and reasoning steps, increasing token generation and infrastructure utilization even when the number of users remains unchanged.

NVIDIA positions Rubin around these reasoning and agentic workloads and claims up to 10 times the token-factory throughput of HGX B200 in a specified projected workload. Actual gains will vary by model, precision, batching strategy, latency target, and software stack.

Why High-Density AI Servers Matter

GPU performance is rising faster than many data centers can expand floor space, power delivery, or cooling capacity. As a result, performance per rack has become an important AI infrastructure metric alongside performance per accelerator.

HPE designed the XD700 specifically around higher GPU density. Its March 2026 announcement stated that an XD700 rack could support up to 128 Rubin GPUs and deliver twice the GPU density of the previous generation.

A later HPE ISC 2026 page lists up to 72 Rubin GPUs per XD700 rack, creating a discrepancy in HPE’s pre-release materials. Final rack density may depend on configuration or evolving production specifications, so buyers should confirm the orderable rack design once HPE publishes the final documentation.

Higher density can reduce:

  • Rack space required for a target GPU count
  • Distance between compute and networking components
  • Infrastructure overhead per accelerator
  • Expansion pressure on data center floor space

Density does not automatically make deployment easier. Concentrating more accelerators into a rack also concentrates electrical demand and heat, making facility readiness part of the server selection process.

Inside the HPE XD700 AI System Architecture

HPE XD700 AI system architecture showing accelerators, host CPU and memory, NVLink fabric, scale-out networking, and storage for AI and HPC workloads.

An HPE GPU server for AI should not be evaluated by accelerator benchmarks alone. The usable performance of an AI system depends on how quickly the entire architecture can feed, connect, cool, and manage its GPUs.

A practical XD700 deployment can be viewed as five connected layers.

LayerRole in the AI system
AcceleratorsRubin GPUs execute AI/HPC computation
Host CPU and memoryRun orchestration and CPU-side workloads
NVLink fabricConnect GPUs inside the HGX domain
Scale-out networkConnect multiple AI servers
StorageFeed training data, models and checkpoints

HPE’s final XD700 CPU configuration should be verified against its production data sheet. Launch reporting quotes HPE executives describing an Intel Xeon 6-based system, while NVIDIA’s HGX Rubin platform itself supports x86 OEM configurations.

Storage becomes equally important as clusters grow. NVIDIA’s Rubin NVL8 SuperPOD reference architecture uses high-performance shared storage connected through InfiniBand or high-speed Ethernet and recommends RDMA-capable architectures to reduce host CPU overhead.

This interconnected approach is also important when planning GPU deployment infrastructure because the server is only one component of a production AI environment.

Cooling, Networking and Memory: The Challenges of AI Infrastructure

Cooling

Eight Rubin accelerators running continuously create substantial heat density. At this performance level, facilities must evaluate liquid-cooling capability, coolant distribution, rack manifolds, heat rejection, and redundancy rather than assuming traditional air cooling will be sufficient.

NVIDIA’s DGX Rubin NVL8 reference architecture uses a liquid-cooled, DC-busbar design, illustrating the facility requirements associated with this class of accelerator. HPE also emphasizes liquid cooling throughout its high-density AI and HPC portfolio.

The choice of AI cooling strategy should therefore begin before hardware arrives, not after racks have been installed.

Networking

NVLink provides high-speed GPU communication inside an NVL8 server, but distributed training needs another fabric between nodes. That scale-out layer may use high-performance InfiniBand or Ethernet depending on architecture, software requirements, cluster size, and cost.

NVIDIA’s Rubin SuperPOD reference design includes 800 Gb/s-class networking and separates compute, storage, management, and other traffic according to workload requirements.

Poor fabric design can turn a GPU problem into a network problem. The practical data center networking challenges include switch capacity, oversubscription, NIC selection, optics, cabling, topology, and storage traffic.

Memory and Storage

Rubin NVL8’s 2.3 TB of HBM4 gives the GPUs fast working memory, but HBM does not replace enterprise storage. Training datasets, model checkpoints, vector data, user content, and model artifacts still need a storage architecture capable of keeping the compute layer supplied.

NVIDIA’s Rubin reference architecture specifically treats high-performance storage as part of the AI cluster rather than an external afterthought.

HPE XD700 vs NVIDIA DGX Rubin and Dell AI Systems

HPE XD700 and NVIDIA DGX Rubin NVL8 use the same underlying eight-GPU Rubin accelerator class, but they serve buyers through different system and operational ecosystems.

Dell has also announced PowerEdge systems using NVIDIA HGX Rubin NVL8. Its Rubin portfolio includes Intel, AMD, and NVIDIA Vera CPU approaches across different system configurations, giving buyers another OEM path for deploying the same accelerator generation.

CategoryHPE Compute XD700NVIDIA DGX Rubin NVL8Dell Rubin systems
AcceleratorHGX Rubin NVL88 Rubin GPUsHGX Rubin NVL8 options
Primary approachHPE AI/HPC ecosystemNVIDIA turnkey platformDell AI Factory ecosystem
CPUFinal SKU details pending2× Intel Xeon 6776P listedIntel, AMD and Vera options announced
ManagementHPE ecosystemNVIDIA Mission ControlDell + NVIDIA ecosystem
Best fitHPE-standardized environmentsNVIDIA-standardized AI stacksDell-standardized environments

NVIDIA currently lists its DGX Rubin NVL8 with two Intel Xeon 6776P processors, 2.3 TB GPU memory, and 28.8 TB/s total NVLink bandwidth. NVIDIA still labels the specifications preliminary.

The best choice is therefore not simply the server with the newest GPU. Existing support contracts, deployment standards, networking, software, power, cooling, availability, and operational expertise may matter just as much.

Organizations considering different Rubin configurations may also compare a Dell Rubin server option when an eight-GPU HGX node is not the only architecture under consideration.

Why HGX Rubin NVL8 Matters for Future AI Factories

AI factories treat computer infrastructure as a system for producing training results, inference responses, tokens, embeddings, and other AI outputs. That approach changes purchasing from choosing individual servers to engineering complete compute environments.

HGX Rubin NVL8 reflects that shift. Its eight GPUs, HBM4 memory, NVLink fabric, scale-out networking requirements, storage dependencies, and cooling requirements all influence real-world performance.

The same principle applies when comparing NVIDIA with alternative accelerator ecosystems. AMD’s emerging rack-scale AI platforms may suit organizations whose workload, software, cost, availability, or infrastructure strategy points toward a different architecture.

Future AI infrastructure will increasingly be measured by complete system capability rather than individual GPU performance. A fast accelerator provides limited value when networking, storage, memory movement, cooling, or power prevents the cluster from keeping it busy.

What Buyers Actually Need to Deploy HPE Compute XD700

Premium HPE Compute XD700 data-center infographic showing a central HPE server tower surrounded by six connected callouts for node configuration, compute and storage networking, high-performance storage, rack power and cooling, software and management, and workload-specific design.

Buying XD700 hardware will represent only one part of a production deployment. Organizations should define the complete bill of infrastructure before committing rack space or data center capacity.

A typical design review should cover:

  • XD700 node count, CPUs, system memory, and Rubin GPUs
  • Compute and storage networking, switches, NICs, optics, and cables
  • High-performance storage capacity, bandwidth, and checkpoint requirements
  • Rack power, busbar/PDU architecture, liquid cooling, and facility capacity

Software also matters. Teams need supported drivers, NVIDIA CUDA and AI software, cluster orchestration, monitoring, workload scheduling, security, and model-serving tools that match their production requirements.

The best configuration depends on workload rather than a single benchmark. Training clusters, enterprise inference farms, AI-agent environments, scientific computing systems, and service-provider AI factories can require very different storage, networking, and redundancy designs.

Conclusion

HPE Compute XD700 represents HPE’s next generation of high-density enterprise AI infrastructure, combining NVIDIA HGX Rubin NVL8 acceleration with an architecture intended for large-scale AI training and inference.

Its real value will depend on more than eight Rubin GPUs. Buyers need to evaluate CPU configuration, NVLink, scale-out networking, storage bandwidth, liquid cooling, power density, software, and facility readiness as one integrated design.

Because XD700 remains a pre-release platform with availability targeted for early 2027, organizations planning deployments today should validate final specifications, rack density, lead time, and configuration options before designing production infrastructure.

How Can Catalyst Support an HPE XD700 AI Infrastructure Deployment?

Technician servicing an HPE XD700 AI system in a data center, illustrating enterprise deployment, infrastructure integration, and ongoing support.

Catalyst Data Solutions works across leading OEM, channel, and distribution ecosystems to help organizations source AI, HPC, and data center infrastructure without forcing every project toward one vendor.

A complete HPE XD700 project may require servers, memory, storage, switching, NICs, DPUs, optics, cabling, racks, power components, cooling infrastructure, and supporting management hardware. Availability for frontier AI equipment can also change as vendors move from announcement to production.

Catalyst can help compare HPE, NVIDIA, Dell, AMD-based, and other infrastructure options based on workload, software ecosystem, data center limitations, deployment timeline, budget, and current supply.

Buyers can request a configuration quote or evaluate the broader AI hardware catalog when planning a complete compute, networking, storage, power, and cooling bundle.

Frequently Asked Questions

1. When will HPE Compute XD700 be available?

HPE currently targets early 2027 availability for HPE Compute XD700. As of August 2026, the server has been announced but is not yet a generally available production platform.

2. How many NVIDIA Rubin GPUs are inside HPE XD700?

Each NVIDIA HGX Rubin NVL8 platform contains eight Rubin GPUs connected through sixth-generation NVLink. Final HPE rack configurations may combine multiple XD700 systems.

3. Is HPE XD700 better for AI training or inference?

It targets both training and inference. Training benefits from GPU density, HBM4 capacity, and high-speed GPU communication, while inference can use the same resources for high-throughput generative and agentic AI serving.

4. Can HPE XD700 run large language models?

Yes. HPE specifically positions XD700 for AI training and inference, while NVIDIA designed HGX Rubin NVL8 for generative AI, reasoning models, mixture-of-experts workloads, and high-performance computing.

5. Does HPE XD700 support AI agent workloads?

The Rubin architecture targets agentic AI and reasoning workloads that require intensive token generation, model communication, and memory movement. Actual agent performance will depend on the model, software stack, networking, storage, and serving architecture.

6. What cooling technology does HPE XD700 use?

HPE has not yet published a complete final XD700 production data sheet covering every cooling configuration. Rubin NVL8-class deployments are highly power-dense, and NVIDIA’s DGX Rubin NVL8 reference implementation requires liquid cooling, so facility liquid-cooling capability should be part of XD700 planning.

7. How does HGX Rubin NVL8 differ from previous HGX platforms?

Rubin NVL8 adds Rubin GPUs, HBM4, sixth-generation NVLink, greater accelerator-memory bandwidth, and higher AI performance compared with previous Blackwell-generation HGX systems. NVIDIA currently lists 2.3 TB HBM4 and 176 TB/s aggregate memory bandwidth for Rubin NVL8 as preliminary specifications.

8. How does HPE XD700 compare with Dell and NVIDIA AI systems?

XD700 uses NVIDIA’s HGX Rubin NVL8 accelerator platform but packages it within HPE’s enterprise AI and data center ecosystem. NVIDIA DGX emphasizes an integrated NVIDIA platform, while Dell provides its own Rubin-based PowerEdge and AI Factory configurations; the best option depends on infrastructure standards, software, support, cooling, availability, and budget.

Research & Fact-Check

Last fact-checked: September 2, 2026

Official Product Source: https://www.hpe.com/us/en/newsroom/press-release/2026/03/hpe-unveils-next-generation-ai-factory-and-supercomputing-advancements-with-nvidia.html

Key Research Findings 

1. HGX Rubin NVL8 platform: HPE Compute XD700 is an AI server built on NVIDIA HGX Rubin NVL8, connecting eight Rubin GPUs through sixth-generation NVLink.

2. Preliminary specifications: NVIDIA lists 2.3 TB HBM4, up to 176 TB/s memory bandwidth, and 28.8 TB/s NVLink Switch bandwidth. These figures may change.

3. Rack-density discrepancy: HPE cites both 128 and 72 Rubin GPUs per rack in different materials. Final density should be confirmed when HPE publishes official QuickSpecs.