NVIDIA Groq 3 LPX Achieves 3,431 TPS Benchmark on Vera Rubin

Bybit
Coinmama




Luisa Crawford
Aug 24, 2026 17:11

NVIDIA’s Groq 3 LPX sets a new standard in AI inference, delivering 3,431 tokens/second on a 100K context benchmark and redefining high-interactivity workloads.



NVIDIA Groq 3 LPX Achieves 3,431 TPS Benchmark on Vera Rubin

NVIDIA’s Groq 3 LPX has redefined performance standards for AI inference, achieving a world-class 3,431 tokens per second (TPS) on a 100K context benchmark, according to an official blog post on August 24, 2026. Benchmarked by Artificial Analysis using the Gemma 4 31B model, this performance underscores Groq 3 LPX’s ability to handle high-interactivity and long-context AI workloads on NVIDIA’s advanced Vera Rubin platform.

The Groq 3 LPX is built around NVIDIA’s LP30 accelerators and rack-scale architecture, offering 315 PFLOPS of FP8 inference compute and 128 GB of SRAM. The system’s focus on deterministic execution, low-latency token generation, and fine-grained scheduling enables it to excel in multiturn agentic sessions and interactive workloads where context grows significantly with each user interaction. For comparison, traditional models often struggle to maintain interactivity at this scale, especially with long input contexts.

Artificial Analysis ran the benchmark with a 100K input context length, measuring Groq 3 LPX’s ability to generate outputs at an unprecedented speed. This is particularly crucial for agentic AI tasks, such as coding or reasoning workflows, where models must process large volumes of accumulated context. NVIDIA reports that such speeds translate to generating 5,000 tokens in just 1.5 seconds, compared to 50 seconds at 100 TPS—an order-of-magnitude leap in efficiency.

Notably, the system also performed well on a 10K context benchmark, delivering 3,382 TPS with minimal variation in latency. For coding tasks, NVIDIA’s SPEED-Bench tests showed a median speed of 4,767 output tokens per second, with 20% of tasks exceeding 5,500 TPS, highlighting LPX’s versatility across different use cases.

okex

Implications for AI Factories and Workloads

The Groq 3 LPX is positioned as a critical component within NVIDIA’s Vera Rubin platform, particularly in AI factories where high-interactivity serving tiers are essential. Its ability to manage 100K+ context tokens while maintaining low latency could transform industries reliant on complex, multiturn AI interactions, such as customer service automation, large-scale coding assistants, and real-time decision-making systems.

Key to this capability is the LPX’s compiler-scheduled workload planning, which minimizes communication overhead between its 256 interconnected LPUs. By overlapping computation and communication at a fine-grained level, the system achieves unmatched efficiency, even at small batch sizes where traditional tensor parallelism techniques often falter.

NVIDIA’s Strategic Position

This announcement comes at a pivotal time for NVIDIA, which has cemented its leadership in AI hardware. While the Groq 3 LPX is not a standalone cryptocurrency-related product, its implications for AI-driven industries are immense. NVIDIA’s Vera Rubin platform, now enhanced by Groq 3 LPX, positions the company to dominate high-demand AI workloads, from real-time inferencing to generative AI applications.

Shares of NVIDIA (NVDA) recently traded at $210.18, down 2.11% in the last 24 hours, amid broader market softness. However, the company’s advancements in AI inference technology reinforce its long-term growth narrative, particularly in high-margin enterprise solutions. Recent rumors about a China-specific LPU product were denied by NVIDIA, clarifying that no such roadmap exists, potentially easing geopolitical concerns for investors.

What’s Next?

The Groq 3 LPX’s demonstrated performance opens the door to new AI applications requiring both speed and scale. NVIDIA has hinted at maintaining these speeds even with multi-hundred-thousand token contexts, which could further revolutionize agentic AI use cases. Given its robust performance metrics and strategic integration with Vera Rubin, the Groq 3 LPX could set the benchmark for high-interactivity AI systems in the years to come.

Image source: Shutterstock



Source link

Binance

Be the first to comment

Leave a Reply

Your email address will not be published.


*