What runs underneath?
Compute & infrastructure
Silicon, inference systems, AI factories and the physical constraints on intelligence.
14 research briefs 84 distinct sources
Research briefs
NVIDIA Blackwell & Rubin GPU Architecture
NVL72 rack-scale systems, 4-bit floating point (FP4) Tensor Cores, and 5th-gen NVLink interconnects
Language Processing Units (LPUs) & SRAM Silicon
Groq LPUs, deterministic tensor streaming, SRAM-first architectures, and ultra-high-speed inference
Wafer-Scale Engines & Cerebras CS-3 Systems
4-trillion transistor monoliths, 44GB on-wafer SRAM, 21 PB/s bandwidth, and cluster-scale supercomputing
AI Factories, Megawatt Datacenters & Grid Infrastructure
100MW–1GW datacenter topologies, liquid cooling CDUs, high-voltage power distribution, and PUE optimization
High-Speed AI Network Fabrics: InfiniBand, RoCEv2 & Optical Switching
InfiniBand Quantum-X, RoCEv2 (Ultra Ethernet), non-blocking fat-tree topologies, and Optical Circuit Switches
Oracle Cloud (OCI) Superclusters & Sovereign Cloud Architecture
Bare-metal GPU nodes, non-blocking RoCEv2 fabrics, distributed data sovereignty, and multi-cloud interconnects
AI Energy Economics, Nuclear Power & SMR Micro-Grids
Small Modular Reactors (SMRs), geothermal, grid queues, carbon-free baseload, and gigawatt power purchase agreements
On-Device Edge AI Silicon & Neural Processing Units (NPUs)
Apple Neural Engine, Qualcomm Snapdragon X, Intel Core Ultra, and 4-bit edge inference runtimes
Open Silicon & RISC-V AI Accelerators (Tenstorrent)
Wormhole, Blackhole, open-source ISA architectures, chiplet scaling, and decoupling AI from closed hardware ecosystems
Application-Specific Transformer ASICs (Etched Sohu)
Hardwired transformer architectures, zero general-purpose overhead, and 10x throughput per dollar
Vector Database Infrastructure & Distributed Indexing
HNSW, DiskANN, IVF-PQ, GPU-accelerated similarity search, and hybrid vector-relational engines
AI Inference Optimization Runtimes & Serving Engines
vLLM, TensorRT-LLM, SGLang, PagedAttention, continuous batching, chunked prefill, and speculative decoding
Confidential Computing & Hardware-Attested GPU Security
Trusted Execution Environments (TEEs), NVIDIA Hopper/Blackwell CC, remote attestation, and data clean rooms
High-Throughput AI Storage & Distributed Parallel Filesystems
GPUDirect Storage (GDS), NVMe-over-Fabrics, high-throughput checkpointing, and parallel filesystems (Lustre, WEKA, VAST)