Nvidia Accelerates AI Inference with Full Production of Groq 3 LPX Racks
Nvidia announced Monday that its Groq 3 LPX rack is now in full production, marking a significant commercial milestone for the technology acquired in its largest-ever deal. The Groq rack, designed for high-speed, low-latency AI inference, will be deployed alongside Nvidia’s Vera central processors and Rubin graphics processors at Nebius, with systems going live later this year, according to Dion Harris, Nvidia’s senior director.
This strategic push underscores the escalating demand for near-instantaneous AI responses, particularly for sophisticated applications like coding assistants where prolonged delays can significantly degrade user experience. Nvidia emphasizes that this capability allows cloud providers to offer premium service tiers to customers prioritizing responsiveness. “For folks who are serving tokens, it unlocks the ability to offer premium tiers of service for those users and those customers who actually demand the most latency-sensitive service agreements,” Harris stated.
The move follows Nvidia’s acquisition of assets from AI chip startup Groq in December for an estimated $20 billion, its most substantial acquisition to date. The Groq architecture is notable for its integrated 500 megabytes of high-speed SRAM directly on the chip die, a design intended to mitigate memory-related bottlenecks that can impede inference speed. While Groq chips are manufactured by Samsung, Nvidia’s own GPUs are produced by Taiwan Semiconductor Manufacturing.
Each Groq 3 LPX rack houses 256 individual Groq 3 chips. Nvidia reports that these racks can achieve a throughput of 3,400 tokens per second, citing benchmark data from Artificial Analysis. This performance metric positions the Groq solution in a competitive landscape.
Rival chipmakers are also intensifying their focus on low-latency inference. Advanced Micro Devices (AMD), for instance, announced earlier this year a strategy to integrate its rack-scale systems with chips from Cerebras, a company that recently went public and specializes in AI hardware for low-latency workloads. Notably, OpenAI’s recently introduced Ultrafast mode, which promises 750 tokens per second, is reportedly powered by Cerebras technology.
It is crucial to understand that these low-latency specialized chips are not designed to replace the ubiquitous GPU. GPUs remain the workhorses of AI, adept at both training and inference, and possess the flexibility to adapt to evolving models and technologies. Chips like Groq primarily target the “decode” phase of model serving, a critical but specific segment of the overall AI inference process.
“This isn’t about replacing GPUs,” Harris clarified. “It’s about using the right price, right processor for the right part of the workload.”
Nvidia is currently scaling up shipments of its Vera Rubin systems, which commenced production earlier this year. During the unveiling of the Vera Rubin and Groq 3 LPX systems in March, Nvidia CEO Jensen Huang projected cumulative sales of $1 trillion between the current-generation Blackwell chips and the new Vera Rubin systems through 2027. At the time, Huang indicated that a quarter of the data center space allocated for coding applications would be dedicated to Groq chips, with the remainder fully utilized by Vera Rubin.
Nvidia is scheduled to report its latest quarterly earnings on Wednesday.
Original article, Author: Tobias. If you wish to reprint this article, please indicate the source:http://aicnbc.com/25096.html