
A Systems Engineering Guide to Measuring and Improving Model Training and Inference With GPUs
by Kenji Watanuki
Your model works. It is also three times slower and five times more expensive than it needs to be, and nobody on the team can say why.
Most AI performance problems are not solved by trying random optimisations until something improves. They are solved by building a model of the workload, measuring where time and memory actually go, and then applying the right fix to the right bottleneck. This book gives you that method. You will learn to construct a roofline for a real workload, classify whether it is compute-bound, memory-bound, or communication-bound, and only then reach for a fix. That discipline separates teams that ship reliable, cost-effective systems from teams that keep guessing.
Starting from GPU architecture and profiling fundamentals, the book moves through the full stack of production AI systems. You will profile kernels and input pipelines, adopt reduced precision without silently losing accuracy, fuse operators, and choose intelligently between data, tensor, pipeline, and sharded-optimiser parallelism once communication starts dominating. On the serving side, you will implement continuous batching, manage the key-value cache, apply quantisation and distillation with measured trade-offs, and size capacity against real latency objectives. Every technique is presented with the measurement that tells you whether it worked.
What you will learn:
For machine learning engineers, platform and systems engineers, and technical leads running training or inference at scale, this is the systems engineering guide that turns performance work from folklore into a repeatable engineering practice.