Open textbook derives AI infrastructure from hardware limits

Bojie Li has released ai-infra-book, an open-source textbook that explains AI systems through compute, memory, bandwidth, and communication constraints. Within weeks, the project passed 3,400 GitHub stars. Its Apache 2.0 repository includes the manuscript, a PDF, a calculation CLI, and reproducible experiments.

Published in Chinese as Understanding AI Infra: Quantitative Analysis and System Design, the book follows the quantitative tradition of Hennessy and Patterson’s Computer Architecture: A Quantitative Approach. It is a companion to Li’s earlier AI agent book, which has passed 45,000 GitHub stars.

Start with the bottleneck

Li’s method begins by defining the task and quality target, listing the required compute, storage, communication, and dependencies, then comparing order-of-magnitude estimates with hardware capacity, bandwidth, and throughput. This process exposes common sizing errors, including counting model-weight reads while omitting the KV cache, projecting performance from peak FLOPs when memory cannot supply data fast enough, and distributing work across accelerators without budgeting for interconnect traffic.

Five recurring questions organize the analysis: what data moves, how much moves, how often it moves, which path it takes, and which components wait. The book applies that framework to operator fusion, runtime scheduling, multi-accelerator servers, and clusters containing thousands of GPUs.