SiliconANGLE
Important
NvidiaGroqInferenceHardwareNvidia Ships Groq 3 LPX Inference Rack, Nebius First Customer
August 24, 20263 min read
Nvidia has put its Groq 3 LPX dedicated inference accelerator into full production as a rack-scale extension of the Vera Rubin platform, with Nebius as the first cloud customer. The system is quoted at up to 3,400 output tokens per second on long contexts.
Why it matters
The launch shows Nvidia rapidly productizing the Groq acquisition into a high-throughput inference offering aimed at agentic and long-context workloads.
Nvidia has begun full production shipments of the Groq 3 LPX, a dedicated inference accelerator developed from its earlier Groq-related acquisition activity. The LPX is designed as a rack-scale extension of the Vera Rubin platform and can support up to 256 accelerators per rack. Nebius will be the first cloud provider to deploy the systems.
Nvidia is quoting performance of approximately 3,400 output tokens per second on 100,000-token contexts for certain workloads, positioning the LPX as a high-throughput option for agentic and long-context inference. The product reflects Nvidia’s push to cover both training and specialized inference with tightly integrated systems.
Early cloud adoption will be an important signal of whether the architecture can win share against existing GPU-based inference clusters.