Skip to content
FrankX.AI
Research Hub/NVIDIA Blackwell & Rubin GPU Architecture

NVIDIA Blackwell & Rubin GPU Architecture

NVL72 rack-scale systems, 4-bit floating point (FP4) Tensor Cores, and 5th-gen NVLink interconnects

TL;DR

NVIDIA Blackwell transitions AI hardware from individual discrete GPUs to unified rack-scale computers. The GB200 NVL72 connects 72 Blackwell GPUs and 36 Grace CPUs into a single massive GPU via NVLink 5 (1.8 TB/s per GPU bidirectional bandwidth), delivering a 30x inference throughput acceleration and 25x energy efficiency increase on trillion-parameter MoE models.

Updated 2026-08-186 source references4 claims indexed

Research briefs like this, when the evidence is ready. Source links, limitations, and open questions.

Subscribe

1.44 ExaFLOPs

FP4 AI compute per GB200 NVL72 rack

NVIDIA Blackwell Whitepaper

1.8 TB/s

Bidirectional NVLink 5 bandwidth per GPU

NVIDIA Hardware Spec

30x

Inference throughput speedup over H100 on MoE models

MLPerf / NVIDIA Evals

100% Liquid

Direct-to-chip liquid cooling enabling 120kW+ per rack

Datacenter Infrastructure Reports
01

Two-Die Monolithic Reticle Packaging

Blackwell merges two full-reticle limit dies into a single unified GPU over a 10 TB/s high-density chip-to-chip (NV-HBI) interface, containing 208 billion transistors fabricated on custom TSMC 4NP process technology.

Unified Cache Coherency

Memory

Both dies share a single coherent L2 cache and 192GB of ultra-fast HBM3e memory at 8 TB/s bandwidth.

Zero Software Fragmentation

CUDA

Presents to CUDA drivers and compilers as a single seamless unified GPU without multi-GPU NUMA complexities.

High-Density Packaging (CoWoS-L)

Packaging

Advanced 2.5D packaging with multiple passive silicon interposers for sub-picosecond signal integrity.

02

Second-Gen Transformer Engine & FP4 Precision

Blackwell introduces native micro-tensor FP4 arithmetic, doubling compute throughput and halving memory bandwidth demands compared to FP8 without degrading model convergence.

Micro-Tensor Scaling

FP4

Applies fine-grained dynamic quantization scaling factors across small 16-element sub-vectors to prevent underflow.

Autonomous Precision Switching

Engine

Dynamically switches between FP4, FP8, and FP16 across attention layers during forward and backward passes.

Lossless MoE Quantization

MoE

Enables 600B+ parameter MoE models to execute fully within single NVL72 rack memory pools.

03

NVL72 Rack-Scale Computing & Liquid Cooling

The GB200 NVL72 is not 72 separate servers; it is a single liquid-cooled 120kW supercomputer. It replaces traditional copper PCIe switches with a massive 5000-copper-cable NVLink spine.

Direct-to-Chip Liquid Cooling

Cooling

Liquid coolant circulates at 25°C directly over CPU and GPU cold plates, removing 100% of thermal heat with zero fans.

Passive Copper NVLink Spine

Cabling

Replaces expensive optical transceivers with passive cartridge copper cables, saving 20kW of power per rack.

All-to-All Trillion-Parameter Serving

Routing

All 72 GPUs communicate at full crossbar bandwidth, eliminating network stalls in distributed MoE expert routing.

Key Findings

1

The GB200 NVL72 acts as a single 1.44 ExaFLOP GPU with 13.8 TB of unified fast memory, eliminating inter-node networking bottlenecks for MoE models.

2

Native FP4 precision cuts inference power consumption by 25x on massive reasoning models like DeepSeek-R1 and GPT-5.

3

Direct-to-chip liquid cooling enables compute densities exceeding 120kW per rack while lowering datacenter PUE to under 1.10.

4

Passive copper NVLink spine cabling reduces datacenter networking transceiver failures by over 90%.

5

The next-generation Rubin architecture integrates HBM4 3D-stacked memory with 36-die NVLink 6 switches, scaling bandwidth by an additional 2.5x.

Research Transparency

Limitations

  • Deploying NVL72 racks requires specialized datacenter facilities capable of 120kW+ power delivery and liquid cooling CDU loops.
  • Global supply constraints on advanced CoWoS packaging and HBM3e/HBM4 memory limit immediate mass availability.

What We Don't Know

  • ?The long-term silicon degradation rates under continuous 24/7 FP4 thermal cycling at 120kW power densities.
  • ?Optimal scheduling compilers for dynamically partitioning heterogeneous CPU-GPU workloads across multi-rack clusters.
Evidence Grade:Grade A(Backed by NVIDIA Blackwell Architecture Technical Whitepaper, Hot Chips 2024 presentations, and MLPerf verified hardware benchmarks.)

Frequently Asked Questions

Blackwell is NVIDIA's next-generation GPU architecture featuring 208 billion transistors across dual-die packaging, 4-bit Floating Point (FP4) Tensor Cores, and NVLink 5 interconnects, designed to train and serve trillion-parameter AI models.

From research to practice

Learn these tools hands-on

The research maps the landscape. These portals curate the videos, docs, and experts to actually build with the platforms it covers.