John L. Hennessy & David A. Patterson
References
Every chapter closes with a "grounded in" line. This is the same list turned inside out: each source A–Z, and the chapters built on it.
22 sources · 20 chaptersMemoryGPU
Computer Architecture: A Quantitative Approach, 6th ed.
BookComputer Organization and Design, RISC-V ed.
BookDavid A. Patterson & John L. Hennessy
Computer Systems: A Programmer's Perspective, 3rd ed.
BookRandal E. Bryant & David R. O'Hallaron
Docs
Talk
Lectures
Dissecting the NVIDIA Volta GPU Architecture via Microbenchmarking (2018)
PaperZhe Jia, Marco Maggioni, Benjamin Staiger & Daniele P. Scarpazza
Gallery of Processor Cache Effects
ArticleIgor Ostrovsky
General-Purpose Graphics Processor Architectures (2018)
BookTor M. Aamodt, Wilson W. L. Fung & Timothy G. Rogers
GPU Glossary
GlossaryModal
Hitting the Memory Wall: Implications of the Obvious (1995)
PaperWm. A. Wulf & Sally A. McKee
Latency Numbers Every Programmer Should Know
ArticleJeff Dean's numbers · Colin Scott's interactive version
Article
NVIDIA A100 Tensor Core GPU Architecture (Ampere whitepaper)
WhitepaperNVIDIA
NVIDIA H100 Tensor Core GPU Architecture (Hopper whitepaper)
WhitepaperNVIDIA
- 1.2SIMT: The Warp· Table 4 · Fig. 7
- 1.3Inside the SM· SM architecture · Table 4
- 1.4Latency Hiding· Table 4 · Fig. 7
- 1.5The Memory System· Table 3 · SXM5
- 1.6The Full Chip· Architecture In-Depth
- 1.7Host & Device· Table 3 · NVLink · PCIe
- 1.8Roofline· Table 1 · Table 3
- 1.XCapstone: Read the Spec Sheet· the datasheet lines; 132/144 binning gap
NVIDIA Hopper Architecture In-Depth
ArticleNVIDIA technical blog
NVIDIA's H100: Funny L2, and Tons of Bandwidth
ArticleChips and Cheese
Operating Systems: Three Easy Pieces
BookRemzi H. Arpaci-Dusseau & Andrea C. Arpaci-Dusseau
Performance Analysis and Tuning on Modern CPUs, 2nd ed.
BookDenis Bakhvalov
Programming Massively Parallel Processors, 4th ed.
BookWen-mei W. Hwu, David B. Kirk & Izzat El Hajj
The VisualGPU roadmap
This siteThis curriculum's own map — every chapter and what it builds on
Sources for the CUDA and Inference topics will join the list as those chapters are built.