GTC 2026 Vera Rubin platform: 7 chips, 5 racks, what's new

Updated 2026-10-01 · Video: NVIDIA, published 2026-03-16

This page covers the hardware tour of the GTC 2026 keynote, from 1:07:56 to 1:20:00: a ten-year recap from DGX-1 to Blackwell, then the Vera Rubin rack, Vera CPU, NVLink 6, the co-packaged-optics Spectrum-X switch, BlueField-4 STX storage and a first physical look at Rubin Ultra in the Kyber rack. Watch it if you buy, plan or analyze data-center capacity. The economics that justify all this hardware sit in the next segment, covered in our Groq LPX and disaggregated inference notes.

Key takeaways

  1. The unit of sale is now the AI factory, not the chip. Vera Rubin is pitched as seven chips and five rack types that NVIDIA designs together (1:11:12). That framing matters more than any single spec, because it is how NVIDIA argues against cheaper standalone accelerators.
  2. Vera CPU becomes a product in its own right. Jensen describes a CPU built for single-thread speed and claims about 2x the performance per watt of other CPUs (1:16:09). He also says standalone CPU sales are already a multi-billion-dollar business. Neither number appears in NVIDIA’s press releases.
  3. Assembly time is a selling point. The Vera Rubin rack is 100% liquid-cooled and cable-free inside the tray, which Jensen says cuts installation from two days to two hours (1:14:00). It runs on 45 °C warm water.
  4. Optics enter the switch package. The Spectrum-X switch with co-packaged optics (CPO), built on TSMC’s COUPE process, is described as in full production (1:15:29).
  5. Rubin Ultra changes the rack shape. The Kyber rack stands compute nodes vertically against a midplane and links 144 GPUs in one NVLink domain (1:17:58).

Chapter notes

1:07:56 – 1:10:04 Ten years from DGX-1 to Blackwell

The segment opens with a video recap. It starts with DGX-1 in April 2016: eight Pascal GPUs on first-generation NVLink at 170 teraflops. It then steps through the Volta NVLink Switch, the Mellanox acquisition and the A100 SuperPOD, Hopper’s FP8 Transformer Engine, and Blackwell NVL72 with 130 TB/s of all-to-all bandwidth across 72 GPUs.

The pattern is the point. Each generation added one more layer of the data center to what NVIDIA designs itself: first the board, then the GPU-to-GPU switch, then the network, and now storage and the CPU. If you want to know why NVIDIA keeps saying “extreme co-design”, this two-minute montage shows how it got there. The recap closes with a claim of a 40-million-fold increase in compute over ten years (1:11:20). Treat that as a marketing number: it compares a 2016 single box with a 2026 multi-rack system, so it mixes per-chip gains with scale-out gains.

1:10:04 – 1:11:36 The Vera Rubin platform in one slide

The video describes Vera Rubin NVL72 at 3.6 exaflops with 260 TB/s of NVLink all-to-all bandwidth. NVIDIA’s official materials confirm the 260 TB/s figure (3.6 TB/s per GPU) and 50 PFLOPS of NVFP4 inference per Rubin GPU. Seventy-two of those works out to the quoted 3.6 exaflops, so that number is an NVFP4 inference figure, not general-purpose FP64 compute.

Around the GPU rack sit four more rack types:

The official release lists the same five racks and calls them all in full production.

NVIDIA Vera Rubin slide: 7 chips, 5 rack systems, with a table comparing a 1 GW AI factory built on x86 plus Hopper against Vera Rubin
1:35:26 — Jensen returns to the platform at the start of the next segment with this summary slide. It compares a 1 GW factory on x86 + Hopper with one on Vera Rubin: 600K vs 300K GPUs, 1.2 vs 16 ZFLOPS, and 2M vs 700M tokens per second. The table is the clearest single view of the platform's chips and trays. Its 350x token claim is a stage figure; we did not find it in the official press release.

The slide above is the best reference for the platform, even though it appears a few minutes after this segment ends. Read the table carefully. Halving the GPU count while multiplying AI FLOPS by about 13x tells you most of the gain is per-GPU and in low-precision formats. The “Memory BW-per-Domain” row explicitly counts Groq SRAM, so the 100 EB/s figure only applies to a factory that includes LPX racks.

1:11:36 – 1:14:43 Why agents need a different system

Jensen’s argument for redesigning the whole rack, not just the GPU, is about agents. An agent does three things: it thinks (LLM inference), it remembers (KV cache, structured data through cuDF, vector search through cuVS), and it calls tools (browsers and virtual machines in the cloud). Each of those hits a different part of the system. Thinking is GPU-bound, memory is storage-bound, and tool calls are CPU-bound and latency-sensitive.

That is the justification for the Vera CPU. Jensen describes it as built for single-thread performance, the only data-center CPU using LPDDR5 memory, with high performance per watt. NVIDIA’s Vera release gives 88 custom Olympus cores, up to 1.2 TB/s of LPDDR5X bandwidth and 1.8 TB/s of NVLink-C2C to the GPU. The practical reading is simple: when an AI agent waits on a slow CPU, the expensive GPU sits idle, so NVIDIA wants to own that part too.

The physical rack demo at 1:14:00 is more interesting for operators than for investors. Moving from two days to two hours of installation per rack, if real at scale, shortens the time between delivery and revenue. Warm-water cooling at 45 °C means less power spent on chillers, which in a power-capped site turns directly into more power for compute.

This is a rapid walk past hardware on stage:

Jensen also shows a 256-node liquid-cooled Ethernet rack using the same structured-cabling approach as the NVLink rack (1:17:21). The takeaway: NVIDIA is taking rack mechanics it built for NVLink and reusing them across networking.

1:17:39 – 1:20:00 Rubin Ultra and the Kyber rack

Rubin Ultra is the one genuinely new physical object in this segment. Standard Rubin compute trays slide into the Oberon rack horizontally. Rubin Ultra nodes stand vertically in the new Kyber rack and plug into a midplane, with NVLink switches mounted on the back side. One Kyber rack links 144 GPUs in a single NVLink domain.

Why it matters: a bigger NVLink domain lets a single model instance spread across more GPUs without going over the slower scale-out network. That is what very large mixture-of-experts models need. The trade-off is a new rack form factor that data centers must plan for, which is why NVIDIA keeps Oberon available for buyers who do not want to change anything (see the roadmap notes).

What changed since GTC 2025

At GTC 2025 (March 18, 2025), Vera Rubin was a roadmap item. NVIDIA’s live blog said Vera Rubin systems would arrive in the second half of 2026 and Rubin Ultra systems in the second half of 2027. It also introduced the Kyber rack design as part of the Rubin Ultra plan.

Topic GTC 2025 said GTC 2026 said
Vera Rubin timing Systems in H2 2026 “Full production”; Azure has the first rack running; partner products from H2 2026
Platform scope GPU + Vera CPU + NVLink Seven chips, five racks, including a Groq LPU that did not exist in the 2025 plan
Rack naming “Vera Rubin NVL144” “Vera Rubin NVL72” for what appears to be the same 72-package rack; NVIDIA now counts GPU packages, not dies
Rubin Ultra H2 2027, Kyber rack Physical Kyber node shown on stage; chip said to be taped out
Co-packaged optics Spectrum-X and Quantum-X photonics switches announced Spectrum-X CPO described as in full production

Two items stand out. First, the schedule held: what was promised for the second half of 2026 was in production by March 2026. Second, the scope grew. The Groq LPU and the standalone Vera CPU rack were not part of last year’s Vera Rubin story.

Sources: GTC 2025 live updates, Vera Rubin Opens Agentic AI Frontier, Rubin launch at CES 2026.

Skip list

Glossary

FAQ

What is the NVIDIA Vera Rubin platform?

Vera Rubin is NVIDIA's 2026 data-center platform. NVIDIA describes it as seven chips (Rubin GPU, Vera CPU, NVLink 6 Switch, ConnectX-9, BlueField-4, Spectrum-6 with co-packaged optics, and the Groq 3 LPU) packaged into five rack-scale systems that are sold and tuned as one AI factory.

How fast is the Vera Rubin NVL72 rack?

In the keynote video NVIDIA quotes 3.6 exaflops for the NVL72 rack and 260 TB/s of all-to-all NVLink bandwidth. The 260 TB/s figure (3.6 TB/s per GPU) appears in NVIDIA's official materials; the 3.6 exaflops figure is consistent with 72 GPUs at 50 PFLOPS NVFP4 each.

What is the difference between Rubin and Rubin Ultra?

Rubin NVL72 uses the existing Oberon rack with compute trays slid in horizontally. Rubin Ultra moves to the new Kyber rack, where compute nodes stand vertically and plug into a midplane, putting 144 GPUs in a single NVLink domain.

When does Vera Rubin ship?

At GTC 2026 Jensen said Vera Rubin was in full production and that Microsoft Azure had the first rack running. NVIDIA's press release says partner products arrive starting in the second half of 2026.