GTC 2026 Groq LPX + Vera Rubin: token tiers and the 35x claim
Watch the original video · 139 min
This segment runs from 1:20:00 to 1:34:40. It is the business case for Vera Rubin. Jensen sorts tokens into price tiers, shows how each GPU generation moves the throughput curve, and then explains why NVIDIA bolted Groq’s LPU onto its own platform. It is the most important quarter-hour of the keynote for anyone who prices inference, runs an inference service, or models NVIDIA revenue. If you have not seen the tokens-per-watt framing yet, start with our token economics notes.
Key takeaways
- Tokens will be sold in tiers, like any maturing commodity. Jensen sketches free, $3, $6, $45 and $150 per million token tiers, each tied to model size, context length and speed (1:21:37). These are illustrative prices from a slide, not NVIDIA products.
- Stage claim: Vera Rubin earns about 5x Blackwell’s revenue per gigawatt in a simple model that splits power evenly across four tiers (1:26:02). This number is not in NVIDIA’s press release.
- GPUs cannot max out both FLOPS and bandwidth on one die, so NVIDIA adds a second chip type rather than compromise (1:26:38).
- Dynamo splits decode in two. Attention runs on Rubin, the feed-forward network runs on Groq, and Ethernet links them (1:31:53).
- Shipping: Samsung builds the chip, and LPX is expected around Q3 2026 (1:33:01). Jensen’s advice is to give roughly 25% of a data center to Groq if you serve a lot of coding or other high-value tokens.
Chapter notes
1:19:43 – 1:23:30 From one curve to price tiers
Jensen reuses a slide from last year’s keynote: throughput per megawatt on the vertical axis, tokens per second per user on the horizontal. Every AI factory sits somewhere on that curve. Move left and you serve many users slowly. Move right and you serve fewer users quickly, with bigger models and longer contexts.
The new part is the price overlay. The slide maps tiers to example models: a free tier on a 235B-parameter Qwen 3 with 32K context, a $3 tier on Kimi K2.5, $6 and $45 tiers on a 2-trillion-parameter GPT-class MoE, and an “Ultra” tier at $150 per million tokens with 400K context. Jensen’s back-of-envelope: a researcher burning 50 million tokens a day at $150 per million costs $7,500 a day, which he calls trivial for a research team.
How to read this: NVIDIA is not setting token prices. It is telling inference providers that the money is in the right-hand side of the chart, where speed and model size justify higher prices, and that its hardware roadmap aims there. Whether customers will actually pay $150 per million tokens is an open question. Current frontier API prices sit far below that for most models.
1:23:30 – 1:26:38 Hopper to Blackwell to Rubin on the same chart
Hopper occupies a small corner of the chart. Grace Blackwell lifts the whole curve, and Jensen claims about 35x more throughput in the tier where providers make most of their money (1:23:46). Vera Rubin lifts it again: roughly 2x at the low-speed end and, by his account, about 10x in the highest-priced tier (1:25:03).
He then runs a toy revenue model. Take 1 GW and split it 25/25/25/25 across free, medium, high and premium tiers. With the same power budget, he says Blackwell yields about 5x Hopper’s revenue and Vera Rubin about 5x Blackwell’s again. The slide shows roughly $30B a year per gigawatt for Blackwell against $150B for Rubin.
This model hides a lot. It assumes demand at every tier is unlimited and that the illustrative prices hold. The useful insight is structural. Because power is fixed, a faster chip does not just lower cost. It moves capacity into tiers that sell for more. That is why NVIDIA argues the newest system is the cheapest per dollar of revenue, even at a higher sticker price.
1:26:38 – 1:29:18 Why add Groq at all
Here Jensen concedes a limit of his own architecture. High throughput needs lots of floating-point math. Low latency at very high tokens per second needs memory bandwidth. A single die has fixed area, so optimizing for both at once is a contradiction.
The NVL72 handles most of the chart well, but past roughly 400 tokens per second per user it runs out of bandwidth. He says serving 1,000 tokens per second per user is out of reach for NVL72 alone. That gap is where Groq’s SRAM-heavy chip fits. NVIDIA acquired the team that built Groq’s chips and licensed the technology. The slide claims another 35x in the most valuable tier once LPX is added.
1:29:18 – 1:33:01 Disaggregated inference through Dynamo
This is the technical core. Groq’s chip is a deterministic dataflow processor. A compiler schedules every operation ahead of time, so data and compute arrive together and nothing is scheduled dynamically. It holds about 500 MB of SRAM per chip, against 288 GB of HBM4 on a Rubin GPU. That is a big reason Groq never went mainstream on its own: holding a trillion-parameter model plus its KV cache would take an enormous number of chips.
NVIDIA’s answer is to split the work with Dynamo, its inference-serving software:
- Prefill (reading the prompt) runs on Vera Rubin. It is compute-heavy and easy to batch.
- Decode attention also runs on Vera Rubin, because it needs the large KV cache that lives in HBM.
- Decode feed-forward runs on Groq chips. The model weights sit in SRAM, so the bandwidth-bound part of generating each token is fast.
The two racks exchange activations over Ethernet. Jensen says a special mode cuts that latency roughly in half (1:32:20). Splitting at the attention/FFN boundary is a finer cut than the prefill/decode split Dynamo launched with in 2025. It only works because the network between the racks is fast enough. That is also why it is hard for a competitor to copy with off-the-shelf parts. For a slower explanation, see disaggregated inference explained.
1:33:01 – 1:34:40 Production status
Samsung manufactures the Groq LP30 chip. Jensen expects LPX shipments in the second half of 2026, probably around the third quarter. He adds that Vera Rubin sampling went far more smoothly than Grace Blackwell’s. Microsoft Azure already had its first Vera Rubin rack running. The supply chain can build thousands of these systems a week, which he equates to several gigawatts of AI factory per month.
Stage claims vs. official text
| Claim | On stage | In NVIDIA’s official text |
|---|---|---|
| Vera Rubin + LPX speed-up | 35x in the premium tier | “Up to 35x higher inference throughput per megawatt” for trillion-parameter models |
| Revenue uplift | Rubin ≈ 5x Blackwell; Rubin + LPX ≈ 10x on the slide | “Up to 10x more revenue opportunity for trillion-parameter models” |
| LPX hardware | 8 Groq chips per tray; LP30 is third generation | 256 LPUs per rack, 128 GB SRAM, 640 TB/s scale-up bandwidth |
| Attention/FFN split | Explained in detail | Developer blog confirms LPX and NVL72 pairing; the split detail comes mainly from the keynote |
| Ship date | Around Q3 2026 | “Second half of this year” |
| Rubin ≈ 5x Blackwell revenue per GW | Stated with a slide | Not found in the press release |
What changed since GTC 2025
GTC 2025 launched Dynamo as open-source inference software. Its pitch then was disaggregated serving: prefill and decode on different GPUs. NVIDIA’s March 2025 release claimed Dynamo raised tokens per GPU by more than 30x on DeepSeek-R1 on GB200 NVL72, and doubled Llama throughput on Hopper. Jensen called it the operating system of an AI factory, and he uses the same phrase again this year.
The throughput-versus-interactivity chart was also introduced in 2025. Jensen says so on stage, and that is why he shows “last year’s slide”.
What is new in 2026:
- A non-NVIDIA chip inside the platform. GTC 2025 did not mention Groq. Disaggregation was GPU-to-GPU. In 2026 it spans two chip architectures.
- A finer split. In 2025 the cut was between prefill and decode. Now decode itself splits into attention on Rubin and FFN on Groq.
- Explicit token tiers with prices. The 2025 curve showed a trade-off. The 2026 version attaches dollar figures and a revenue model to it.
Sources: Dynamo launch release (March 2025), GTC 2025 live updates, Vera Rubin Opens Agentic AI Frontier.
Skip list
- 1:22:56 – 1:23:14 The researcher-cost arithmetic. It is summarized above.
- 1:28:21 – 1:28:28 Applause pause.
- 1:33:24 – 1:33:40 Product beauty shot of the LPX rack with no new information.
Glossary
- Token tier — a price level for AI output, set by model size, context length and generation speed.
- Interactivity (TPS/user) — tokens per second delivered to one user; higher means faster, more responsive answers.
- Throughput per megawatt — total tokens a factory produces per unit of power; see tokens per watt explained.
- Dynamo — NVIDIA’s open-source inference-serving software that routes parts of a request to different hardware.
- Groq LPU — a deterministic, compiler-scheduled chip that keeps model weights in on-chip SRAM for very fast token generation.
- Attention / FFN — the two main blocks of a transformer layer; attention reads the KV cache, the feed-forward network applies the model’s weights.
FAQ
What is NVIDIA Groq 3 LPX?
Groq 3 LPX is a rack of 256 Groq LPU chips that NVIDIA added to the Vera Rubin platform. NVIDIA's official figures are 128 GB of on-chip SRAM per rack and 640 TB/s of scale-up bandwidth. It is built for the fastest, most latency-sensitive part of token generation.
How do Groq LPX and Vera Rubin work together?
NVIDIA Dynamo splits inference across the two. Prefill and the attention part of decode run on Vera Rubin, which holds the KV cache, while the feed-forward part of decode runs on Groq chips, which hold the model weights in SRAM. The two exchange activations over Ethernet.
Is the 35x performance claim official?
Yes, but it is narrow. NVIDIA's press release says Vera Rubin plus LPX delivers up to 35x higher inference throughput per megawatt and up to 10x more revenue opportunity for trillion-parameter models, compared with Blackwell. On stage the 35x applies to the premium, high-interactivity tier, not to all workloads.
When does Groq LPX ship?
In the keynote Jensen said Samsung manufactures the Groq LP30 chip and that LPX should ship in the second half of 2026, probably around the third quarter. NVIDIA's March press release says the second half of the year.