NVIDIA has begun full-scale mass production of the accelerator "Groq 3 LPX (NVIDIA Groq 3 LPX)," specialized for agentic artificial intelligence (AI) inference. Aimed at the characteristics of agentic AI, which must generate tokens quickly while handling large-scale context, the product expands the inference performance of the next-generation AI platform "Vera Rubin."

Grock 3/Nvidia

NVIDIA said on the 25th (local time) at the Hot Chips semiconductor conference in the United States that it has started mass production of the Groq 3 LPX.

Agentic AI, unlike conventional Generative AI that answers a user's question once, reasons on its own, calls tools, and repeats tasks. In complex work, reasoning unfolds over hundreds to thousands of steps, making the ability to handle large-scale context and the speed of rapid token generation important.

Groq 3 LPX focuses on boosting token generation speed. According to NVIDIA, Groq 3 LPX recorded 3,400 output tokens per second on the Artificial Analysis benchmark applying a 100,000-token context to the open-source agentic model "Gemma 4 31B." NVIDIA said this is the fastest performance to date for that model.

It also said it delivers up to four times faster response than comparable alternative platforms for agentic AI tasks such as coding. The goal is to shorten processing time for AI agents that repeat multiple tasks, from file review to code writing and testing, tool calls, and result verification.

Jensen Huang, NVIDIA's chief executive officer (CEO), said, "Inference is the engine of AI growth," adding, "Vera Rubin extends this vision with an AI factory configuration optimized for workloads for the agentic AI era, and pushes performance boundaries further with LPX for ultrafast token generation."

Groq 3 LPX expands inference performance in combination with Vera Rubin NVL72, NVIDIA's next-generation AI platform. While Vera Rubin NVL72 handles AI computation such as large-scale context processing, Groq 3 LPX is used to boost token generation speed.

NVIDIA explained that agentic AI simultaneously demands two computing challenges: efficiently handling a vast context and generating tokens quickly with low latency. Groq 3 LPX is specialized for increasing speed in the generation stage that delivers results to the user.

AI cloud companies are also moving to adopt it. Nebius plans to introduce Groq 3 LPX into its inference platform "Nebius Token Factory." According to NVIDIA, it is the first AI cloud company to introduce Groq 3 LPX into a production environment.

Danila Shtan, Nebius' chief technology officer (CTO), said, "Generation is the inference stage that determines the actual response speed of AI systems, and NVIDIA Groq 3 LPX is designed to accelerate this very process," adding, "It provides an experience where every step of the agent loop happens instantly while developers continue using existing APIs without needing to migrate to a new stack."

Groq, a cloud company specialized in AI inference, is also expected to be among the early adopters.

NVIDIA is building Vera Rubin as an AI factory platform co-designed with seven chips and five custom-designed racks. By combining Vera Rubin NVL72 and Groq 3 LPX with the BlueField-4 DPU, Vera CPU, and Spectrum-6 Ethernet, the strategy is to increase the throughput and response speed of multi-agent systems.

※ This article has been translated by AI. Share your feedback here.