Nvidia has laid out full specifications for "Rubin," the GPU architecture that follows Blackwell, and the numbers are aimed squarely at two problems hyperscalers keep raising in public: the electricity bill for running large language models, and the memory bottleneck that slows down long-context inference. Rubin GPUs, built around a chip Nvidia calls the R100, are expected to reach cloud providers in the second half of 2026.
What's actually in the chip
The R100 is built on TSMC's 3-nanometer process and packs 336 billion transistors across a dual-die design, up from 208 billion on Blackwell. It pairs with up to 288GB of next-generation HBM4 memory delivering as much as 22 terabytes per second of bandwidth — nearly triple the 8 TB/s of Blackwell's HBM3e. That memory jump matters more than it might sound: HBM4 doubles the interface width per stack compared with HBM3e, and bandwidth is what determines how fast a GPU can stream the "KV cache" that long-context conversations and agentic workloads depend on. Nvidia's published figures put Rubin's inference throughput at 50 petaflops of FP4 compute per GPU and training throughput at 35 petaflops, versus roughly 10 petaflops each on Blackwell — a claimed 5x inference and 3.5x training improvement generation-over-generation. NVLink 6 interconnect bandwidth doubles to 3.6 terabytes per second per GPU.
None of this comes free on the power bill: individual Rubin GPUs are rated around 1,800–2,300 watts, up from about 1,200–1,400 watts on Blackwell, and Nvidia has said there will be no air-cooled configuration — every Rubin deployment requires liquid cooling. At the rack level, Nvidia's "Vera Rubin NVL72" configuration pairs 72 Rubin GPUs with 36 of the company's new Vera CPUs (built on custom Arm-based "Olympus" cores) to deliver up to 3.6 exaflops of FP4 inference performance and 20.7TB of pooled HBM4 memory in a single liquid-cooled rack.
Where the efficiency argument comes from
Nvidia and several independent analysts have framed Rubin's real selling point as tokens generated per watt rather than raw compute, since that is the metric that actually determines a data center's power draw for a given amount of AI work delivered. Nvidia's own public materials describe up to roughly 10x higher inference throughput per watt compared with Blackwell at the rack-scale level, driven by the combination of faster compute, denser memory, and a redesigned rack architecture rather than any single component. Third-party estimates of the resulting drop in per-token serving cost vary widely — from roughly 2.5x to 10x depending on the workload and whether the comparison is chip-to-chip or full-rack-to-full-rack — so treat any single "reduces AI costs by X%" headline with some skepticism until real-world deployments produce audited numbers.
Beyond the standard GPU: Rubin CPX and Kyber racks
Alongside the standard R100, Nvidia has introduced Rubin CPX, a variant it describes as a new class of GPU purpose-built for "massive context" inference — the kind of workload where a model has to hold hundreds of thousands of tokens of conversation, code, or documents in memory at once. A rack built around the CPX variant, referred to as NVL144 CPX, is spec'd for up to 100TB of aggregate memory and roughly 1.7 petabytes per second of bandwidth, according to Nvidia's own product materials. For the next rung up the product line, Nvidia has also detailed a "Kyber" rack architecture intended for a future Rubin Ultra chip, designed to support up to 144 GPUs in a vertical configuration with a newer NVLink 7.0 interconnect. None of that hardware ships in 2026; it's a roadmap commitment, and roadmap commitments from any chipmaker should be read as directional rather than guaranteed.
Why hyperscalers care right now
Data center electricity demand tied to AI has become one of the industry's most politically sensitive topics in 2026, with utilities, regulators, and local communities in the US and Europe pushing back on new substation construction and water use tied to hyperscale campuses. Every major cloud provider — Amazon, Google, Microsoft, and Oracle among them — has said publicly it plans to deploy Rubin-based systems once they ship, and the chip is central to the huge compute contracts AI labs have been signing throughout 2026 to secure capacity years in advance.
- 336 billion transistors on TSMC's 3nm process (vs. 208 billion on Blackwell)
- Up to 288GB of HBM4 memory at up to 22 TB/s bandwidth per GPU
- 50 petaflops FP4 inference / 35 petaflops FP4 training per GPU
- NVLink 6 at 3.6 TB/s per GPU; liquid cooling required, no air-cooled option
- Production targeted for the second half of 2026
Our take
Nvidia clearly wants Rubin judged on efficiency rather than raw speed, and that's a sign the company knows power, not silicon supply, is becoming the ceiling on how much AI compute the world can actually plug in. The headline "5x" and "10x" figures are real Nvidia claims, not fabrications, but they come from Nvidia's own benchmarking and comparisons that mix chip-level and rack-level measurements — exactly the kind of numbers that deserve independent verification once real customers get their hands on the hardware later this year.
What to watch next
The things to track between now and Rubin's actual shipment in H2 2026: whether TSMC's 3nm capacity can keep up with demand from Nvidia alongside Apple and other customers, how quickly cloud providers can build out the liquid-cooling infrastructure Rubin requires, and whether independent benchmarks confirm anything close to Nvidia's efficiency claims once the chips are running production workloads rather than vendor demos.