NVIDIA Groq 3 LPX inference rack with 256 Groq 3 LPU accelerators, high-speed SRAM, and low-latency token generation designed to extend Vera Rubin NVL72 for agentic AI.
Fast Shipping
100k+ SKUs
8/5 Support
Certified
NVIDIA Groq 3 LPX Interactive AI Inference Accelerator Rack
NVIDIA Groq 3 LPX Interactive AI Inference Accelerator (GROQ3-LPX)
NVIDIA Groq 3 LPX is an interactive AI inference accelerator rack designed for low-latency, high-throughput token generation in agentic AI systems. Each LPX rack integrates 256 Groq 3 LPU accelerators and is designed to work alongside NVIDIA Vera Rubin NVL72, complementing Rubin GPU compute with deterministic, SRAM-centric decode acceleration for long-context and latency-sensitive inference.
Technical Specifications
| SKU | GROQ3-LPX |
|---|---|
| Product Type | Interactive AI Inference Accelerator Rack |
| Accelerators | 256 NVIDIA Groq 3 LPU accelerators per LPX rack |
| Rack SRAM | 128 GB SRAM |
| Rack DDR5 Memory | 12 TB DDR5 |
| SRAM Bandwidth | 40 PB/s per rack |
| Scale-Up Bandwidth | 640 TB/s across the LPX rack |
| Primary Platform Pairing | NVIDIA Vera Rubin NVL72 |
| Availability Status | Mass Production / RFQ |
Deployment and Compatibility Considerations
- Plan LPX as an extension to a Vera Rubin inference architecture rather than as a generic PCIe accelerator.
- Validate high-speed connectivity between LPX and Vera Rubin NVL72 racks for the intended inference topology.
- Confirm rack power, cooling, network, management, and facility requirements before deployment.
- Size the architecture around context length, decode latency, concurrency, and token-throughput objectives.
Common Use Cases
- Low-latency agentic AI token generation
- Interactive coding and tool-using agents
- Long-context reasoning and inference
- Premium latency-sensitive AI services
- Vera Rubin token-factory deployments
Frequently Paired Infrastructure
- NVIDIA Vera Rubin NVL72
- NVIDIA Spectrum-X Ethernet or Quantum-X800 InfiniBand
- NVIDIA BlueField-4 infrastructure processing
- AI-native context-memory storage
- High-density power and cooling infrastructure
Official NVIDIA Documentation
Frequently Asked Questions
What role does Groq 3 LPX play next to Vera Rubin NVL72?
LPX is designed to accelerate low-latency decode and token generation while Rubin GPUs handle high-bandwidth model computation, allowing a co-designed inference architecture for agentic workloads.
How many LPU accelerators are in an LPX rack?
NVIDIA documents 256 Groq 3 LPU accelerators per LPX rack.
| product type | AI Inference Accelerator Rack |
|---|---|
| Architecture / Series | NVIDIA Groq 3 LPX |
| Accelerator Configuration | 256 Groq 3 LPUs per Rack |
| Memory / Bandwidth | 128 GB SRAM / 40 PB/s SRAM Bandwidth |
| Availability Status | Mass Production / RFQ |
