Nvidia Launches Groq 3 LPX Inference Chip into Full Production
Reported by The Decoder: Nvidia's dedicated inference accelerator Groq 3 LPX enters full production to supercharge AI agents - SiliconANGLE. Analysis and context written by TickrWire.
Nvidia has moved its Groq 3 LPX inference accelerator into full production, reporting 3,400 tokens per second on Gemma 4 31B, four times faster than Cerebras, but the comparison depends on accelerator count.

- Nvidia’s Groq 3 LPX is now in full production and claims 3,400 tps on Gemma 4 31B, four times faster than Cerebras’ 882 tps.
- The speed advantage depends on using at least 64 LPUs; Cerebras needs only one or two for the same model.
- Scaling to large MoE models would require thousands of LPUs, raising questions about cost and practicality.
- The benchmark used a dense model; performance on more complex architectures remains untested.
- Nebius will be the first cloud provider to offer the chip, indicating early commercial interest.
Nvidia announced at the Hot Chips 2026 conference that its Groq 3 LPX inference accelerator has entered full production. The chip, built on the Vera Rubin platform, is marketed as an interactive AI inference accelerator designed to deliver ultrafast token generation for agentic AI systems. Nvidia plans to roll the technology out later this year.
The move follows Nvidia’s late‑December acquisition of Groq for roughly $20 billion, bringing founder Jonathan Ross and president Sunny Madra into the Nvidia fold. Groq has long focused on processors tuned for inference rather than training, and the 3 LPX is the latest iteration of that line.
In an independent benchmark conducted by Artificial Analysis, the Groq 3 LPX achieved 3,400 tokens per second on the open‑source Gemma 4 31B model with a 100,000‑token context window. The test ran 50 back‑to‑back requests and maintained steady performance across input lengths from 10,000 to 100,000 tokens. Nvidia claims this figure is the highest ever recorded for the model and that it is four times faster than the next best option, Cerebras, which achieved 882 tokens per second.
However, the raw speed numbers hide a key scaling factor. Each Groq 3 LPX processing unit (LPU) contains only 500 MB of SRAM, about 576 times less than a Rubin GPU’s 288 GB. To run Gemma 4 31B, the system requires at least 64 LPUs, whereas Cerebras needs only one or two. For larger mixture‑of‑experts models such as DeepSeek V3, the Groq architecture would need 1,342 LPUs, or a little over five racks, compared to Cerebras’ single or dual‑chip requirement.
The architecture relies on a SRAM‑heavy dataflow design. Models are split across multiple LPUs over Ethernet, with GPUs handling the compute‑heavy prefill phase and the LPUs handling the bandwidth‑heavy decode phase. This hybrid setup can deliver high throughput, but it also introduces complexity in coordination and potential latency.
The comparison also omits Cerebras’ newer CS‑4 generation, which could narrow the performance gap. Moreover, the benchmark used a dense model that fits entirely in one rack, leaving open how the Groq 3 LPX scales with larger, more complex MoE models. The cost of deploying dozens or hundreds of LPUs, as well as the operational overhead of managing a multi‑rack system, remains unclear.
Nebius plans to be the first cloud provider to offer the Groq 3 LPX through its Token Factory, and Groq itself is among the early adopters. The production launch positions Nvidia to compete directly with Cerebras in the high‑speed inference market, but the real‑world advantage will depend on how quickly the technology can be integrated into agentic AI workflows and how the cost‑performance balance plays out.
What to watch next includes the rollout timeline, pricing models, and performance on a broader set of models, especially those with large MoE layers. Investors and developers will also be interested in how the chip’s architecture affects power consumption, latency, and ease of integration with existing AI pipelines.
Provides a new high‑speed inference option that could reduce latency in agentic AI pipelines.
May lower operational costs for real‑time AI services if deployment scales efficiently.
Signals Nvidia’s push into inference hardware, potentially affecting market share against Cerebras.
Illustrates practical trade‑offs between raw speed and hardware scaling in AI accelerators.
Highlights the complexity behind headline performance claims in AI hardware.
- Groq 3 LPX
- An inference accelerator built on the Vera Rubin platform, designed for fast token generation in agentic AI.
- Token generation
- The process of producing individual tokens (words or subwords) during language model inference.
- Mixture‑of‑Experts (MoE)
- A model architecture that routes input to a subset of expert sub‑models to improve efficiency.
- SRAM-heavy dataflow architecture
- A design that relies on static RAM for fast data movement, often requiring many small units.
- LPU
- LPU stands for “Logic Processing Unit,” the core compute element in Groq chips.
AI bias estimate: The comparison emphasizes Nvidia’s speed figures while downplaying the higher accelerator count and omitting Cerebras’ newer CS‑4 generation. (Automated estimate, not a definitive judgement.)
AI ToolsAccel-backed Keenable is indexing the web for AI agents
AI Tools‘The world seems to be ready’: An interview with OpenAI head of product Thibault Sottiaux
AI ToolsAn AI boss fired its first employee but only after humans reminded it of its own rules
AI ToolsVercel Introduces ‘Is Agentic’, a Free Agent-Readiness Scoring Tool That Audits Public Websites Using Ora’s 100+ Checks
AI ToolsDecoding AI’s Open-Source Course Maps Three Ways to Run an Agent Loop and the Provider Economics Behind Each
HardwareOpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show
OpenAI shared initial benchmark data for its Jalapeño inference chip at the Hot Chips conference, demonstrating higher token generation and energy efficiency than current market standards.
SecurityUkraine opens its massive labeled battlefield dataset to British firms in a landmark AI weapons partnership
Ukraine has granted the United Kingdom access to its Avengers Labs platform, providing foreign tech companies with millions of annotated combat images to train military artificial intelligence.
RoboticsI spent a day at a robot “carnival” in Shanghai. Here’s what I saw.
A recent robotics festival in Shanghai highlighted China's rapid commercial progress in humanoid systems, where local firms now dominate global delivery numbers.
SecurityTaiwanese cybersecurity firm warns that AI tools have more than doubled Chinese state-backed cyberattacks
Taiwanese cybersecurity firm TeamT5 reports that Chinese state-backed hacking groups have more than doubled their attack frequency after adopting AI models such as DeepSeek.

Jalapeño’s first results show industry-leading speed and efficiency in AI inference
OpenAI’s Jalapeño custom inference chip delivers 1.5‑3.6× better power efficiency and lower latency across major language models, outperforming commercial competitors.

Disrupting a new covert influence campaign from Russia
OpenAI blocked Russian ChatGPT accounts that were used to push a fabricated Israeli think‑tank and a pro‑Russia sovereignty index, exposing a covert influence operation.