Learning: Inference
Be able to stand up an inference server for a real model, on GPU and on CPU/edge in turn, and defend the latency and throughput numbers it produces instead of quoting whatever the framework's defaults happen to give you.
Latest lesson: 25. Reviewing a Serving Configuration
Success looks like
- Stand up a serving stack (vLLM on GPU, then llama.cpp on CPU/edge) for a given model and get it answering requests.
- Quote a p99 latency budget for a given batch size and model, and defend the number from the KV cache, batching and quantization choices that produced it.
- Review someone else's serving configuration and name specifically what it costs (a knob mistuned against the stated workload's bottleneck, a topology chosen without a bottleneck that justifies it, a rollout that breaks a caller mid-flight), rather than saying it feels wrong.
Constraints
- Assumes comfort running a Python ML environment and basic familiarity with what a language model is; no prior serving-infrastructure experience required.
- Core stacks: vLLM (GPU) and llama.cpp (CPU/edge), covered in that order. Other stacks (TGI, TensorRT-LLM) are mentioned only where a concept transfers differently.
Out of scope
- Training or fine-tuning a model or adapter: see
llm/finetuning(lessons 0025, serving adapters, and 0026, cost, latency and throughput), linked back to rather than restated. - Judging output quality or building an eval for a served model: see
llm/evalsfor that boundary.
The arc
Fourteen stages, first request to a defended, fleet-scale, production-monitored, and reviewed serving configuration. A stage takes several lessons and the boundaries are soft; what makes a stage done is the capability, not the lesson count.
| Stage | Lessons | Covers | Done when |
|---|---|---|---|
| 1. The KV cache | 0001 to 0003 | Why autoregressive generation is expensive, what the cache trades memory for, how cache size grows with context and batch | Can compute a model's KV cache memory footprint for a given context length and batch size |
| 2. Batching | 0004 to 0006 | Static vs continuous batching, request scheduling, the throughput/latency trade-off | Can defend a batching configuration for a stated workload |
| 3. Quantization at serve time | 0007 to 0009 | int8/int4/AWQ/GPTQ at inference time, accuracy vs speed vs memory | Can pick a serving-time quantization scheme and defend the trade-off |
| 4. vLLM in practice | 0010 to 0012 | Standing up vLLM, PagedAttention, the tuning knobs that matter | A vLLM stack is running and answering real requests |
| 5. llama.cpp on CPU/edge | 0013 to 0015 | GGUF, llama.cpp's architecture, what changes off-GPU | A llama.cpp stack is running on CPU and answering real requests |
| 6. The latency budget | 0016 to 0017 | p99 measurement methodology, tying the number back to cache, batching and quantization choices | Can quote and defend a p99 latency budget end to end |
| 7. Speculative decoding | 0018 | Draft-model, n-gram, Medusa, and EAGLE speculative decoding | Can explain why parallel verification is nearly free and pick a speculative-decoding variant for a workload |
| 8. Sampling and constrained decoding at serve time | 0019 | Per-request sampling parameters in a shared batch, FSM-indexed constrained decoding | Can explain why constrained decoding's guarantee is structural and why its per-token cost stays small |
| 9. Prefix caching | 0020 | Automatic prefix caching as cross-request block sharing, long-document and multi-turn workloads, its prefill-only limit | Can explain when prefix caching helps a workload and why it never speeds up decode |
| 10. Multi-GPU and multi-node parallelism | 0021 | Tensor parallelism's actual mechanism past lesson 12's single flag, pipeline parallelism, and a decision procedure for which one (or both) a deployment needs | Can explain what a tensor-parallel combine step and a pipeline bubble each cost, and choose the right split for a given interconnect and node count |
| 11. Fleet-level serving | 0022 | Request routing informed by cache locality, queue-depth-driven autoscaling, and why a cold start's model load time bounds how fast a new replica actually helps | Can explain how routing, autoscaling, and cold-start latency have to be reasoned about as one pipeline rather than tuned separately |
| 12. Production observability and cost | 0023 | Continuously exported metrics, SLIs/SLOs/error budgets that turn a metric into a target, and cost per million tokens derived from measured throughput | Can defend a production configuration by citing its p99, the SLO and error budget it's held to, and its cost per million tokens |
| 13. Long context at serve time | 0024 | RoPE scaling versus sliding-window attention and attention sinks, and what each actually does to lesson 2's cache-growth formula | Can explain whether a long-context technique extends usable length, caps cache size to a constant, or fixes a capped cache's failure mode |
| 14. Judgment | 0025 | Reviewing someone else's configuration for mismatched knobs, safe rollouts through what others depend on, and saying when speculative decoding is not a free win | Trusted to review a serving configuration and name what each choice costs, not just that it feels wrong |
Lessons
Work through these in order.
| # | Lesson | Teaches |
|---|---|---|
| 0001 | The KV Cache | Why generation gets expensive, and what caching buys back |
| 0002 | Capacity and Batch Size | How head sharing (GQA/MQA) and batch size change the cache's real footprint |
| 0003 | Growth, Prefill, Decode, and Precision | How the cache grows across prefill and decode, and what lowering its precision buys back |
| 0004 | Static vs Continuous Batching | Why batching requests together helps throughput, and why continuous batching beats the static kind |
| 0005 | Request Scheduling | How a continuous-batching scheduler picks the next request, and why a large prefill can stall everyone else |
| 0006 | The Throughput/Latency Trade-off | How batch size trades throughput against per-token latency, and how to defend a batching configuration against a stated latency budget |
| 0007 | Quantization Schemes at Serve Time | What int8, int4, GPTQ and AWQ actually do to a model's weights, and why the naive version of low-bit quantization needs a fix |
| 0008 | Accuracy, Speed, and Memory Trade-offs | Why memory, speed, and accuracy don't move together when a model is quantized, and how each is actually measured |
| 0009 | Picking and Defending a Quantization Scheme | A decision framework for choosing a serving-time quantization scheme, and the three numbers that defend it |
| 0010 | Standing Up vLLM | Installing vLLM, launching its server, and finding the flags that carry the concepts already taught |
| 0011 | PagedAttention | The block table mechanism behind vLLM, its block-size trade-off, and how it lets sequences share a common prefix's cache |
| 0012 | The vLLM Tuning Knobs That Matter | Two more flags, gpu-memory-utilization and tensor-parallel-size, plus a decision procedure for which knob a symptom actually points at |
| 0013 | GGUF | What llama.cpp's single-file model format bundles together, its quantization naming, and why it can be memory-mapped instead of loaded |
| 0014 | llama.cpp's Architecture | The ggml tensor library underneath llama.cpp, its backend abstraction, and how CPU threading differs from GPU batching |
| 0015 | What Changes Off-GPU | Standing up llama.cpp's server, and how batching, cache management, and quantization each look different at CPU/edge scale |
| 0016 | p99 Latency Measurement Methodology | Why p99 beats an average, why TTFT and inter-token latency need separate numbers, and what makes a p99 measurement trustworthy |
| 0017 | Defending a Latency Budget End to End | A diagnostic order for a missed p99 budget, and the three things a final defended configuration must cite |
| 0018 | Speculative Decoding | Verifying several candidate tokens in one parallel forward pass costs about what generating one token normally costs, since decoding is bottlenecked on moving weights, not on the arithmetic, which is the one fact every variant of speculative decoding is built on |
| 0019 | Sampling and Constrained Decoding at Serve Time | A continuous batch already runs many requests through one shared forward pass, and giving each one its own sampling settings, or forcing one of them to only ever produce valid JSON, both have to happen per-row inside that same shared pass without slowing everyone else down |
| 0020 | Prefix Caching | Lesson 11 let parallel samples of the same request share one prompt's cached blocks; prefix caching is the same sharing applied across otherwise unrelated requests, and it speeds up exactly the phase that redundant prefill wastes, never the phase that generates the answer |
| 0021 | Tensor and Pipeline Parallelism | Lesson 12 named tensor-parallel-size as a flag that splits a model across GPUs; this lesson opens up what that split actually does inside a layer, introduces pipeline parallelism as a different split of the same problem, and gives a decision procedure for which one (or both) a workload actually needs |
| 0022 | Fleet-Level Serving | Lesson 5's scheduler picks the next request for one server's queue; a fleet adds a decision before that, which server gets the request at all, a decision about when to add more servers, and a cold-start cost that makes a freshly added server useless for tens of seconds |
| 0023 | Production Observability and Cost | Lesson 17 defended one measured p99 figure for one configuration; this lesson covers what has to be continuously exported to catch a regression automatically, the SLO and error budget that turn "good enough" into a target instead of a vibe, and the cost-per-million-tokens number that has to be defended alongside p99, not instead of it |
| 0024 | Long Context at Serve Time | Lesson 2's capacity formula charges memory linearly with context length; RoPE scaling extends how long a context can be without changing that charge at all, while sliding-window attention and attention sinks change the formula itself by capping context length to a constant |
| 0025 | Reviewing a Serving Configuration | Review someone else's serving configuration and name specifically what it costs (a knob mistuned against the stated workload's bottleneck, a topology chosen without a bottleneck that justifies it, a rollout that breaks a caller mid-flight), rather than saying it feels wrong |
Reference
- Glossary: canonical terms for this topic
- Resources: trusted sources
- KV cache: per-token cost, head sharing, growth across a request, and capacity planning
- Batching: static vs continuous batching, request scheduling, and the throughput/latency trade-off
- Quantization at Serve Time: int8/int4/GPTQ/AWQ, why memory, speed and accuracy don't move together, and a decision framework for picking a scheme
- vLLM: standing up vLLM, PagedAttention, and the tuning knobs that matter
- llama.cpp: GGUF, llama.cpp's architecture, and what changes off-GPU
- Latency Budget: p99 measurement methodology, and tying a missed budget back to cache, batching and quantization choices
How this works
Each lesson is short and self-contained. Answer keys are collapsed: recall first, then open them. The real-world reps matter more than the reading, and spacing them out is the point. Anything still unclear at the end of a lesson is worth chasing to its primary source before moving on.