Particle.news

NVIDIA Puts Groq 3 LPX Racks Into Full Production

The company says the rack-scale systems are built to speed token-level inference and will be paired with Vera CPUs and Rubin GPUs on cloud platforms to capture higher-margin, low-latency services.

Overview

  • NVIDIA announced that Groq 3 LPX racks are in full-scale production and that it is accelerating chip output and customer deliveries, according to company statements.
  • Each LPX rack houses 256 Groq 3 chips that include 500MB of on-die high-speed SRAM to cut memory bottlenecks in decoding and token-generation workloads.
  • NVIDIA said Groq racks will be deployed alongside Vera CPUs and Rubin GPUs on the Nebius cloud platform later this year, and the company framed the hardware as a complement to GPUs rather than a replacement.
  • A vendor-reported benchmark from Artificial Analysis, cited by NVIDIA, shows about 3,400 tokens per second for the LPX rack; NVIDIA noted this figure is a supplier test and has not been independently verified.
  • The production rollout follows NVIDIA’s $20 billion acquisition of Groq assets and comes as rivals like AMD with Cerebras-backed systems and OpenAI’s ‘Ultrafast’ mode push competing low-latency inference options.