Disaggregated Inference Explained: GTC 2026 Rubin + Groq Split

Updated 2026-10-01

Disaggregated inference means splitting one language-model request into stages and running each stage on the hardware that suits it best. At GTC 2026, NVIDIA took the idea one step further than before. Its Dynamo software now splits the decode step itself between Vera Rubin GPUs and Groq 3 LPX accelerators. Jensen Huang explains it from 1:29:18 to 1:34:40.

The stages being split

An LLM request has two phases. Prefill reads the whole prompt in one parallel pass. It is compute-heavy and runs efficiently on GPUs. Decode then produces the answer one token at a time. Each decode step does two different jobs. The attention layers read the KV cache, the stored record of everything in the context, which can reach many gigabytes for long agent sessions. The feed-forward (FFN) layers, which in mixture-of-experts models are the expert blocks, apply the model’s weights. At small batch sizes, that work is limited by how fast weights can be read from memory.

The first generation of disaggregation, which Dynamo introduced in 2025, put prefill and decode on separate GPU pools so they would not slow each other down. The 2026 version is called attention–FFN disaggregation. It also splits decode.

Why Rubin plus Groq

The two chips are opposites. Jensen contrasts them at 1:30:16: a Groq LP30 has about 500 MB of SRAM, and a Rubin GPU has 288 GB. SRAM is very fast but small, and HBM is large but slower per byte. Earlier, at 1:26:37, he names the underlying conflict: high throughput needs huge FLOPS, low latency needs huge bandwidth, and a die of fixed size cannot maximize both.

The design shown at 1:31:55 assigns work by those strengths:

Stage Runs on Why
Prefill Vera Rubin NVL72 Parallel, compute-bound
Decode attention over KV cache Vera Rubin NVL72 Needs large memory capacity
Decode FFN / MoE experts Groq 3 LPX Latency-bound weight reads benefit from SRAM speed

Dynamo routes the prefill and runs the loop. For every token, intermediate activations move from GPU to LPU and back. Jensen says the link is Ethernet in a special mode that cuts latency roughly in half, a detail not found in NVIDIA’s written materials. A trillion-parameter model needs many LPUs just to hold its weights. That is why an LPX rack contains 256 of them.

What it means for latency and cost

The benefit appears at the fast end of the curve. At 1:27:56, Jensen says NVL72 alone runs short of bandwidth beyond roughly 400 tokens per second per user. Adding LPX extends service toward 1,000 tokens per second, which suits premium coding and research tiers. NVIDIA’s release claims up to 35x more throughput per megawatt and up to 10x more revenue opportunity for trillion-parameter models, compared with Blackwell.

The cost is complexity and a second chip to buy. Jensen’s own guidance at 1:28:33 is to use Vera Rubin alone for throughput-heavy work and to put about 25% of the floor on Groq if your workload is mostly high-value coding tokens. In other words, disaggregation pays off only where faster tokens can be sold at a higher price. The tokens per watt explainer covers that pricing logic.

For the hardware details, see Groq LPX and disaggregated inference and the Vera Rubin platform page.

FAQ

What is disaggregated inference?

It means running different stages of one LLM request on different hardware instead of one GPU doing everything. The classic split is prefill on one pool and decode on another. NVIDIA's GTC 2026 design also splits decode itself, putting attention and feed-forward layers on different chips.

Why does NVIDIA pair Vera Rubin with Groq LPX?

Rubin GPUs have large HBM capacity for the KV cache, and Groq LPUs have fast on-chip SRAM for low-latency math. Splitting each decode step lets each chip do the part it is best at. NVIDIA says the pairing raises throughput per megawatt up to 35x for trillion-parameter models.