GTC 2026 Groq LPX + Vera Rubin: token tiers and the 35x claim

Updated 2026-10-01 · Video: NVIDIA, published 2026-03-16

This segment runs from 1:20:00 to 1:34:40. It is the business case for Vera Rubin. Jensen sorts tokens into price tiers, shows how each GPU generation moves the throughput curve, and then explains why NVIDIA bolted Groq’s LPU onto its own platform. It is the most important quarter-hour of the keynote for anyone who prices inference, runs an inference service, or models NVIDIA revenue. If you have not seen the tokens-per-watt framing yet, start with our token economics notes.

Key takeaways

  1. Tokens will be sold in tiers, like any maturing commodity. Jensen sketches free, $3, $6, $45 and $150 per million token tiers, each tied to model size, context length and speed (1:21:37). These are illustrative prices from a slide, not NVIDIA products.
  2. Stage claim: Vera Rubin earns about 5x Blackwell’s revenue per gigawatt in a simple model that splits power evenly across four tiers (1:26:02). This number is not in NVIDIA’s press release.
  3. GPUs cannot max out both FLOPS and bandwidth on one die, so NVIDIA adds a second chip type rather than compromise (1:26:38).
  4. Dynamo splits decode in two. Attention runs on Rubin, the feed-forward network runs on Groq, and Ethernet links them (1:31:53).
  5. Shipping: Samsung builds the chip, and LPX is expected around Q3 2026 (1:33:01). Jensen’s advice is to give roughly 25% of a data center to Groq if you serve a lot of coding or other high-value tokens.

Chapter notes

1:19:43 – 1:23:30 From one curve to price tiers

Jensen reuses a slide from last year’s keynote: throughput per megawatt on the vertical axis, tokens per second per user on the horizontal. Every AI factory sits somewhere on that curve. Move left and you serve many users slowly. Move right and you serve fewer users quickly, with bigger models and longer contexts.

The new part is the price overlay. The slide maps tiers to example models: a free tier on a 235B-parameter Qwen 3 with 32K context, a $3 tier on Kimi K2.5, $6 and $45 tiers on a 2-trillion-parameter GPT-class MoE, and an “Ultra” tier at $150 per million tokens with 400K context. Jensen’s back-of-envelope: a researcher burning 50 million tokens a day at $150 per million costs $7,500 a day, which he calls trivial for a research team.

How to read this: NVIDIA is not setting token prices. It is telling inference providers that the money is in the right-hand side of the chart, where speed and model size justify higher prices, and that its hardware roadmap aims there. Whether customers will actually pay $150 per million tokens is an open question. Current frontier API prices sit far below that for most models.

1:23:30 – 1:26:38 Hopper to Blackwell to Rubin on the same chart

Hopper occupies a small corner of the chart. Grace Blackwell lifts the whole curve, and Jensen claims about 35x more throughput in the tier where providers make most of their money (1:23:46). Vera Rubin lifts it again: roughly 2x at the low-speed end and, by his account, about 10x in the highest-priced tier (1:25:03).

He then runs a toy revenue model. Take 1 GW and split it 25/25/25/25 across free, medium, high and premium tiers. With the same power budget, he says Blackwell yields about 5x Hopper’s revenue and Vera Rubin about 5x Blackwell’s again. The slide shows roughly $30B a year per gigawatt for Blackwell against $150B for Rubin.

This model hides a lot. It assumes demand at every tier is unlimited and that the illustrative prices hold. The useful insight is structural. Because power is fixed, a faster chip does not just lower cost. It moves capacity into tiers that sell for more. That is why NVIDIA argues the newest system is the cheapest per dollar of revenue, even at a higher sticker price.

1:26:38 – 1:29:18 Why add Groq at all

Here Jensen concedes a limit of his own architecture. High throughput needs lots of floating-point math. Low latency at very high tokens per second needs memory bandwidth. A single die has fixed area, so optimizing for both at once is a contradiction.

The NVL72 handles most of the chart well, but past roughly 400 tokens per second per user it runs out of bandwidth. He says serving 1,000 tokens per second per user is out of reach for NVL72 alone. That gap is where Groq’s SRAM-heavy chip fits. NVIDIA acquired the team that built Groq’s chips and licensed the technology. The slide claims another 35x in the most valuable tier once LPX is added.

Throughput per megawatt versus interactivity chart comparing Hopper, Blackwell NVL72, Rubin NVL72 and Rubin plus LPX across Free, Medium, High, Premium and Ultra token tiers
1:28:18 — The full throughput curve. Rubin NVL72 is 2–3x Blackwell at low speeds. The Rubin + LPX line keeps going past 400 tokens per second per user, where Blackwell falls to near zero, and reaches 1,000+. That gap is the "35x" label. It is a ratio against a near-zero baseline in one tier, not an average speed-up.

1:29:18 – 1:33:01 Disaggregated inference through Dynamo

This is the technical core. Groq’s chip is a deterministic dataflow processor. A compiler schedules every operation ahead of time, so data and compute arrive together and nothing is scheduled dynamically. It holds about 500 MB of SRAM per chip, against 288 GB of HBM4 on a Rubin GPU. That is a big reason Groq never went mainstream on its own: holding a trillion-parameter model plus its KV cache would take an enormous number of chips.

NVIDIA’s answer is to split the work with Dynamo, its inference-serving software:

The two racks exchange activations over Ethernet. Jensen says a special mode cuts that latency roughly in half (1:32:20). Splitting at the attention/FFN boundary is a finer cut than the prefill/decode split Dynamo launched with in 2025. It only works because the network between the racks is fast enough. That is also why it is hard for a competitor to copy with off-the-shelf parts. For a slower explanation, see disaggregated inference explained.

1:33:01 – 1:34:40 Production status

Samsung manufactures the Groq LP30 chip. Jensen expects LPX shipments in the second half of 2026, probably around the third quarter. He adds that Vera Rubin sampling went far more smoothly than Grace Blackwell’s. Microsoft Azure already had its first Vera Rubin rack running. The supply chain can build thousands of these systems a week, which he equates to several gigawatts of AI factory per month.

Stage claims vs. official text

Claim On stage In NVIDIA’s official text
Vera Rubin + LPX speed-up 35x in the premium tier “Up to 35x higher inference throughput per megawatt” for trillion-parameter models
Revenue uplift Rubin ≈ 5x Blackwell; Rubin + LPX ≈ 10x on the slide “Up to 10x more revenue opportunity for trillion-parameter models”
LPX hardware 8 Groq chips per tray; LP30 is third generation 256 LPUs per rack, 128 GB SRAM, 640 TB/s scale-up bandwidth
Attention/FFN split Explained in detail Developer blog confirms LPX and NVL72 pairing; the split detail comes mainly from the keynote
Ship date Around Q3 2026 “Second half of this year”
Rubin ≈ 5x Blackwell revenue per GW Stated with a slide Not found in the press release

What changed since GTC 2025

GTC 2025 launched Dynamo as open-source inference software. Its pitch then was disaggregated serving: prefill and decode on different GPUs. NVIDIA’s March 2025 release claimed Dynamo raised tokens per GPU by more than 30x on DeepSeek-R1 on GB200 NVL72, and doubled Llama throughput on Hopper. Jensen called it the operating system of an AI factory, and he uses the same phrase again this year.

The throughput-versus-interactivity chart was also introduced in 2025. Jensen says so on stage, and that is why he shows “last year’s slide”.

What is new in 2026:

Sources: Dynamo launch release (March 2025), GTC 2025 live updates, Vera Rubin Opens Agentic AI Frontier.

Skip list

Glossary

FAQ

What is NVIDIA Groq 3 LPX?

Groq 3 LPX is a rack of 256 Groq LPU chips that NVIDIA added to the Vera Rubin platform. NVIDIA's official figures are 128 GB of on-chip SRAM per rack and 640 TB/s of scale-up bandwidth. It is built for the fastest, most latency-sensitive part of token generation.

How do Groq LPX and Vera Rubin work together?

NVIDIA Dynamo splits inference across the two. Prefill and the attention part of decode run on Vera Rubin, which holds the KV cache, while the feed-forward part of decode runs on Groq chips, which hold the model weights in SRAM. The two exchange activations over Ethernet.

Is the 35x performance claim official?

Yes, but it is narrow. NVIDIA's press release says Vera Rubin plus LPX delivers up to 35x higher inference throughput per megawatt and up to 10x more revenue opportunity for trillion-parameter models, compared with Blackwell. On stage the 35x applies to the premium, high-interactivity tier, not to all workloads.

When does Groq LPX ship?

In the keynote Jensen said Samsung manufactures the Groq LP30 chip and that LPX should ship in the second half of 2026, probably around the third quarter. NVIDIA's March press release says the second half of the year.