Performance work starts with a small set of measurements, then grows into a model of queues, workloads, resources, and trade-offs. This series follows that progression and ends with the full LLM inference stack, where hardware, kernels, hosts, runtimes, and orchestration can all contribute to the bottleneck.

The articles share a common goal: connect a local optimization to its end-to-end effect, then choose the next experiment based on expected return rather than familiarity with a particular layer.