Brief IA

Nvidia's Groq 3 LPX: 3,400 t/s, Comparison with Cerebras Discussed

💼 Business & Startups·Tom Levy·

Nvidia's Groq 3 LPX: 3,400 t/s, Comparison with Cerebras Discussed

Nvidia's Groq 3 LPX: 3,400 t/s, Comparison with Cerebras Discussed
Key Takeaways
13,400 tokens/s measured on a Groq 3 LPX rack with Gemma 4 31B
2The Register notes that at least 64 chips are required on the Nvidia side, compared to 1 to 2 on the Cerebras side
3Production announced and launch expected this year, Nebius aims for a cloud offering
💡Why it mattersThe accelerator is entering production with a high token throughput for agentic use cases, but performance comparisons depend on the sizing and scope of the tests.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Nvidia announces the production of Groq 3 LPX and a record of 3,400 tokens per second on Gemma 4 31B. However, experts point out comparison biases with Cerebras, related to the number of chips and the testing scenario. The chip is integrated into the Vera Rubin platform, with a launch planned for later this year.

The number of chips and the use case shape the comparison

Several factors temper the claimed gap with Cerebras. The Register highlights that the comparison overlooks the number of chips deployed, with Cerebras using only one or two units while Nvidia requires at least 64. The Register also adds that the new CS-4 generation from Cerebras is not taken into account. According to the same source, the test conducted with Gemma 4 31B represents a favorable case, as the dense model can fit into a single rack, and how the architecture would adapt to larger mixture-of-experts models remains unknown. For example, DeepSeek V3 would require 1,342 accelerators, which corresponds to just over five racks. The Register also mentions an architecture based on LPUs, each equipped with 500 MB of memory, which is 576 times less capacity than a Rubin GPU at 288 GB, organized through Ethernet partitioning and task specialization: pre-filling is done on GPU and decoding on LPU. Groq relies on a data flow rich in SRAM. Experts warn of a comparison bias favoring Nvidia, particularly related to the size of the installation used.

High measurements on Gemma 4 31B and a claim of x4

Artificial Analysis observed a speed of 3,400 tokens per second on an LPX rack using the open-source model Gemma 4 31B with a context window of 100,000 tokens, processing 50 consecutive requests. Performance remained consistent for inputs ranging from 10,000 to 100,000 tokens, and this result is presented as the best achieved for this model. According to Nvidia, this places Groq 3 LPX at a level four times higher than the Cerebras chip, which is reported at 882 tokens per second. These figures fall within a range of independent results deemed high for token generation.

Product, integration into Rubin, and market launch timeline

At Hot Chips 2026, Nvidia stated that Groq 3 LPX is in full production, having officially started large-scale production. The accelerator, described as dedicated to interactive AI inference, integrates into the Vera Rubin platform and aims for very fast token generation for agentic systems. Nvidia plans a launch later this year.

Context: license acquisition, agentic uses, and initial deployments

At the end of December, Nvidia invested approximately $20 billion to acquire the Groq license and integrated Jonathan Ross and Sunny Madra. Groq designs processors focused on inference rather than training. In agentic uses, the speed of token generation matters because agents process very large volumes over hundreds to thousands of steps: acceleration allows for more reasoning and tool calls within the same perceived acceptable timeframe, with more iterations, file checks, and write-test cycles possible. According to Nvidia, these advancements enable certain coding tasks to be executed in "minutes instead of hours." Regarding availability, Nebius announces its intention to be the first cloud provider to offer the chip through its Token Factory, while Groq is already among the first users.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.