Skip to content
Chiify
Scaling Production-Ready AI Workloads

Scaling Production-Ready AI Workloads

A Systems Engineering Guide to Measuring and Improving Model Training and Inference With GPUs

by Kenji Watanuki

Your model works. It is also three times slower and five times more expensive than it needs to be, and nobody on the team can say why.

Most AI performance problems are not solved by trying random optimisations until something improves. They are solved by building a model of the workload, measuring where time and memory actually go, and then applying the right fix to the right bottleneck. This book gives you that method. You will learn to construct a roofline for a real workload, classify whether it is compute-bound, memory-bound, or communication-bound, and only then reach for a fix. That discipline separates teams that ship reliable, cost-effective systems from teams that keep guessing.

Starting from GPU architecture and profiling fundamentals, the book moves through the full stack of production AI systems. You will profile kernels and input pipelines, adopt reduced precision without silently losing accuracy, fuse operators, and choose intelligently between data, tensor, pipeline, and sharded-optimiser parallelism once communication starts dominating. On the serving side, you will implement continuous batching, manage the key-value cache, apply quantisation and distillation with measured trade-offs, and size capacity against real latency objectives. Every technique is presented with the measurement that tells you whether it worked.

What you will learn:

  • Build a performance model that predicts whether a workload is compute-, memory-, or communication-bound before you change a line of code
  • Profile GPU kernels and data pipelines to find the real bottleneck instead of the assumed one
  • Keep accelerators fed with input pipelines that do not stall on I/O or preprocessing
  • Apply mixed precision and numerical techniques without degrading model quality
  • Optimise at the kernel level through operator fusion and memory-aware implementation
  • Choose between data, tensor, pipeline, and sharded-optimiser parallelism as communication begins to dominate
  • Measure and improve interconnect and scaling efficiency across multi-GPU and multi-node training
  • Design inference serving with continuous batching, key-value cache management, and quantisation or distillation trade-offs
  • Size capacity and cost against explicit latency and service objectives, then operate the system in production

For machine learning engineers, platform and systems engineers, and technical leads running training or inference at scale, this is the systems engineering guide that turns performance work from folklore into a repeatable engineering practice.

$79.99