Wafer's AI Performance Engineering Resources Part 10: A Practical Guide to the CUDA Programming Guide
Part 10 of Wafer's AI performance engineering resource series curates the NVIDIA CUDA Programming Guide for practitioners. It points kernel tuners to the chapte
Wafer (@wafer_ai, verified account) published part 10 of its AI performance engineering resource series on October 8, 2026. The topic is NVIDIA's CUDA Programming Guide, framed around one starting point for making an LLM faster: understanding where execution is waiting.

Image source: @wafer_ai via X
The series context is clear. The author describes the collection as "the most comprehensive AI performance engineering repo in the world" (the author's own claim) and is publishing one post per resource. Because the root post body is truncated at the thread preview, this article covers only what part 10 verifiably contains.
What part 10 covers: the mechanics behind the waits
The post contrasts two kinds of waiting: an attention kernel stalled while it waits for its next tile of data, and a serving runtime leaving the GPU idle by failing to hand over the next launch on time. NVIDIA's CUDA Programming Guide is introduced as the document that explains the mechanisms behind those waits.
The practical hook is specific. The named audience is AI performance engineers working in PyTorch, Triton, or CUDA, and the post points to the SIMT and CUDA Tile chapters from the angle of tuning attention, matmul, and normalization kernels — connecting layouts and data reuse to GPU execution. Anything beyond that visible portion of the post cannot be confirmed.
Verified paths: official docs and the curation repo
Two links are verified through follow-up posts in the thread:
- NVIDIA's official CUDA Programming Guide: https://docs.nvidia.com/cuda/cuda-programming-guide/
- Wafer's curation repository: https://github.com/wafer-ai/gpu-perf-engineering-resources
The repository's code contents, commits, and license cannot be confirmed from thread-level evidence alone, so treat it strictly as a curated collection. For readers who touch kernel code, the practical move is to open the official guide first and bookmark the series repo.
Who should read it
This is a fitting entry-level curation for performance engineers who want to separate kernel-side data stalls from runtime-side scheduling gaps when diagnosing LLM serving latency. If you tune attention or matmul kernels, or narrow bottlenecks at the Triton/CUDA level, look at the SIMT and tile chapters first.
Sources
- Wafer (@wafer_ai) original thread: part 10: CUDA Programming Guide
- NVIDIA official docs: CUDA Programming Guide
- Wafer curation repo: gpu-perf-engineering-resources