Overview
- NVIDIA announced that Groq 3 LPX racks are in full-scale production and that it is accelerating chip output and customer deliveries, according to company statements.
- Each LPX rack houses 256 Groq 3 chips that include 500MB of on-die high-speed SRAM to cut memory bottlenecks in decoding and token-generation workloads.
- NVIDIA said Groq racks will be deployed alongside Vera CPUs and Rubin GPUs on the Nebius cloud platform later this year, and the company framed the hardware as a complement to GPUs rather than a replacement.
- A vendor-reported benchmark from Artificial Analysis, cited by NVIDIA, shows about 3,400 tokens per second for the LPX rack; NVIDIA noted this figure is a supplier test and has not been independently verified.
- The production rollout follows NVIDIA’s $20 billion acquisition of Groq assets and comes as rivals like AMD with Cerebras-backed systems and OpenAI’s ‘Ultrafast’ mode push competing low-latency inference options.