NVIDIA Groq 3 LPX inference rack with 256 Groq 3 LPU accelerators, high-speed SRAM, and low-latency token generation designed to extend Vera Rubin NVL72 for agentic AI.

Ground, Overnight

Fast Shipping

In Stock

100k+ SKUs

Sales and Support

8/5 Support

New / Refurbished

Certified

NVIDIA Groq 3 LPX Interactive AI Inference Accelerator Rack

SKU: GROQ3-LPX Category: Brand:

⭐ Top Rated Verified Buyers
We are trusted for quality, reliability and fast shipping 💯
REQUEST A QUOTE
0 People watching this product now!
Description

NVIDIA Groq 3 LPX Interactive AI Inference Accelerator (GROQ3-LPX)

NVIDIA Groq 3 LPX is an interactive AI inference accelerator rack designed for low-latency, high-throughput token generation in agentic AI systems. Each LPX rack integrates 256 Groq 3 LPU accelerators and is designed to work alongside NVIDIA Vera Rubin NVL72, complementing Rubin GPU compute with deterministic, SRAM-centric decode acceleration for long-context and latency-sensitive inference.

Availability / RFQ: Deployment remains a high-end infrastructure project that should be planned with the surrounding Vera Rubin, networking, rack, power, and cooling environment. Contact Catalyst Data Solutions to discuss current status, compatibility, configuration requirements, and sourcing options.
End-user / export compliance notice: Advanced AI accelerators and systems may be subject to applicable U.S. export controls, end-user/end-use restrictions, destination requirements, and OEM/channel policies. Catalyst may request purchaser, final end-user, intended-use, and deployment-location information before a quote or fulfillment can be approved. An RFQ does not guarantee allocation or shipment.

Technical Specifications

SKU GROQ3-LPX
Product Type Interactive AI Inference Accelerator Rack
Accelerators 256 NVIDIA Groq 3 LPU accelerators per LPX rack
Rack SRAM 128 GB SRAM
Rack DDR5 Memory 12 TB DDR5
SRAM Bandwidth 40 PB/s per rack
Scale-Up Bandwidth 640 TB/s across the LPX rack
Primary Platform Pairing NVIDIA Vera Rubin NVL72
Availability Status Mass Production / RFQ

Deployment and Compatibility Considerations

  • Plan LPX as an extension to a Vera Rubin inference architecture rather than as a generic PCIe accelerator.
  • Validate high-speed connectivity between LPX and Vera Rubin NVL72 racks for the intended inference topology.
  • Confirm rack power, cooling, network, management, and facility requirements before deployment.
  • Size the architecture around context length, decode latency, concurrency, and token-throughput objectives.

Common Use Cases

  • Low-latency agentic AI token generation
  • Interactive coding and tool-using agents
  • Long-context reasoning and inference
  • Premium latency-sensitive AI services
  • Vera Rubin token-factory deployments

Frequently Paired Infrastructure

  • NVIDIA Vera Rubin NVL72
  • NVIDIA Spectrum-X Ethernet or Quantum-X800 InfiniBand
  • NVIDIA BlueField-4 infrastructure processing
  • AI-native context-memory storage
  • High-density power and cooling infrastructure

Official NVIDIA Documentation

NVIDIA Groq 3 LPX Overview

Frequently Asked Questions

What role does Groq 3 LPX play next to Vera Rubin NVL72?

LPX is designed to accelerate low-latency decode and token generation while Rubin GPUs handle high-bandwidth model computation, allowing a co-designed inference architecture for agentic workloads.

How many LPU accelerators are in an LPX rack?

NVIDIA documents 256 Groq 3 LPU accelerators per LPX rack.

Additional information
product typeAI Inference Accelerator Rack
Architecture / SeriesNVIDIA Groq 3 LPX
Accelerator Configuration256 Groq 3 LPUs per Rack
Memory / Bandwidth128 GB SRAM / 40 PB/s SRAM Bandwidth
Availability StatusMass Production / RFQ
About brand
NVIDIA is the leader in GPU computing and AI infrastructure for high-performance workloads.

Recently Viewed Products