The Groq 3 LPU is an SRAM-based AI accelerator chip made by Nvidia, unveiled at GTC 2026 in San Jose, emerging from Nvidia’s $20 billion licensing and talent deal with Groq announced last year. Each accelerator packs 500 MB of on-chip SRAM, delivering 150 TB/s SRAM bandwidth per chip and 2.5 TB/s scale-up bandwidth, built on Samsung’s 4nm process. This is Nvidia’s answer to inference becoming the next battleground in AI production systems.
Key Takeaways
- Groq 3 LPU delivers up to 35x higher throughput per megawatt for trillion-parameter models versus prior systems
- Each LPU contains 500 MB on-chip SRAM, not HBM like traditional accelerators
- Groq 3 LPX racks hold 256 interconnected LPUs with 128 GB total SRAM and 40 PB/s SRAM bandwidth per rack
- Integrates with Nvidia Vera Rubin platform to handle latency-sensitive token generation while GPUs manage prompt processing
- Optimized for agentic AI systems that consume up to 15x more tokens than standard inference
Why SRAM Beats HBM for Inference
The Groq 3 LPU abandons the HBM (high-bandwidth memory) approach that dominates GPU inference accelerators. Instead, it uses SRAM—fast, on-chip memory that delivers bandwidth advantages in decode operations, where the model generates one token at a time in response to user prompts. A Vera Rubin GPU, by contrast, carries 288 GB of HBM4 memory at 22 TB/s—massive capacity but slower per-operation throughput for the token-generation bottleneck.
This architectural choice matters because inference in 2026 is split into two phases: prefill (processing the user’s input prompt) and decode (generating the response token by token). Prefill is compute-bound; decode is memory-bound. The Groq 3 LPU targets decode, where SRAM’s 150 TB/s per chip shines. When you stack 256 LPUs in a single Groq 3 LPX rack, you get 40 PB/s of SRAM bandwidth—a figure that reshapes what “inference at scale” means.
Nvidia CEO Jensen Huang has called AI data centers “factories,” and the Groq 3 LPU is the production tool for the decode factory. It handles millions of concurrent users generating tokens in parallel, each expecting sub-second latency. The SRAM-first design eliminates the memory latency that plagued prior inference systems.
Integration with Vera Rubin: Prefill Meets Decode
The Groq 3 LPU does not stand alone. It integrates into Nvidia’s Vera Rubin platform—the NVL72 liquid-cooled supercomputer that pairs Blackwell GPUs with specialized inference accelerators. In this ecosystem, Vera Rubin GPUs handle prefill (the heavy compute work of processing your prompt), while Groq 3 LPUs handle decode (the latency-sensitive token generation).
This split-personality approach unlocks efficiency gains that neither chip alone can deliver. A Vera Rubin system without Groq 3 LPUs must use CPX—Nvidia’s GDDR7-based inference accelerator—to manage decode. Nvidia VP Buck hinted that the Groq 3 LPU may reduce the role of CPX, focusing instead on tight Groq 3 LPX integration with Rubin. The result is a system optimized for the token-per-second-per-user metric that matters in production AI: higher throughput at lower latency than GPU-only inference.
Vera Rubin systems will begin arriving in the second half of 2026, with Groq 3 LPX racks announced for integration at that time. This is not a product you buy today; it is the foundation of next-generation AI data centers arriving in months.
Performance Claims and Real-World Impact
Nvidia claims the Groq 3 LPU delivers up to 35x higher throughput per megawatt for trillion-parameter models and million-token contexts compared to prior systems. For agentic AI systems—where agents spawn multiple reasoning steps and consume up to 15x more tokens than standard inference—this advantage compounds. A million-token context is no longer theoretical; it is the operating environment for reasoning-heavy AI in 2026.
The throughput gain matters because inference is expensive. Every token generated costs compute, memory bandwidth, and power. If you can generate the same token 35 times faster per watt, you reduce the cost of running an AI factory by orders of magnitude. This is why Nvidia positions the Groq 3 LPU as a “milestone in accelerated computing” and a “new class of inference performance”—the language of a paradigm shift.
That said, these are vendor assertions, not independently verified benchmarks. Nvidia has not published detailed test suites comparing the Groq 3 LPU against competing inference accelerators under controlled conditions. The 35x figure assumes specific workloads (trillion-parameter models, million-token contexts, decode-heavy inference) and may not apply to smaller models or prefill-dominated scenarios. Skepticism is warranted until third-party benchmarks emerge.
Samsung 4nm: The Bedrock
The Groq 3 LPU is built on Samsung’s 4nm process, a critical detail often overlooked. This process node enables the dense, power-efficient SRAM arrays that make the chip tick. Samsung’s 4nm is mature, proven, and capable of delivering the yield and consistency required for data center hardware. It is not latest (TSMC and Samsung both have sub-3nm nodes), but it is the right choice for a memory-heavy design where SRAM density and power efficiency matter more than raw logic speed.
What This Means for the Inference Landscape
The Groq 3 LPU signals a shift in how Nvidia thinks about inference. For years, GPUs dominated—they were flexible, powerful, and owned the market. But GPUs are generalists; they excel at both training and inference, at prefill and decode. The Groq 3 LPU is a specialist. It sacrifices flexibility for speed in the one task that matters most in production: generating tokens fast and cheap.
This specialist approach mirrors trends in semiconductor design. Custom silicon (like Google’s TPUs or Meta’s custom training chips) outperforms general-purpose processors when the workload is predictable and large-scale. Inference in AI factories is predictable—you know the models, you know the token generation patterns, you can optimize for them. The Groq 3 LPU proves that Nvidia understands this shift and is willing to cannibalize some GPU inference revenue to own the inference factory of the future.
How does the Groq 3 LPU compare to Vera Rubin GPUs for inference?
The Groq 3 LPU and Vera Rubin GPUs are complementary, not competing. Vera Rubin GPUs excel at prefill (processing long input prompts with high compute throughput), while Groq 3 LPUs excel at decode (generating tokens with low latency and high bandwidth). Together in a Vera Rubin system, they deliver higher tokens-per-second-per-user than either alone.
When will Groq 3 LPX racks be available?
Groq 3 LPX racks were announced at GTC 2026 for integration with Vera Rubin NVL72 systems, with Vera Rubin systems arriving in the second half of 2026. There is no separate consumer or early-access availability; these are data center products for large-scale AI infrastructure deployments.
Can the Groq 3 LPU replace GPUs for inference?
No. The Groq 3 LPU is optimized for decode—the token-generation phase of inference. It still requires GPU compute for prefill, the initial processing of input prompts. A complete inference system uses both, which is why Nvidia designed the Vera Rubin platform to pair them together.
The Groq 3 LPU represents a maturation of inference as a distinct workload. For years, inference rode on the back of training-optimized GPU architectures. Now, with trillion-parameter models and million-token contexts becoming standard, Nvidia is purpose-building silicon for the inference factory. The $20 billion Groq deal was expensive, but it bought Nvidia the IP and talent to reshape how AI systems generate responses at scale. By late 2026, Vera Rubin systems with Groq 3 LPUs will define the cost-per-token economics of production AI. Everything else will have to compete on that metric.
Edited by the All Things Geek team.
Source: Tom's Hardware


