Skip to content
teach

Learning: Inference

Be able to stand up an inference server for a real model, on GPU and on CPU/edge in turn, and defend the latency and throughput numbers it produces instead of quoting whatever the framework's defaults happen to give you.

Latest lesson: 25. Reviewing a Serving Configuration

Success looks like

  • Stand up a serving stack (vLLM on GPU, then llama.cpp on CPU/edge) for a given model and get it answering requests.
  • Quote a p99 latency budget for a given batch size and model, and defend the number from the KV cache, batching and quantization choices that produced it.
  • Review someone else's serving configuration and name specifically what it costs (a knob mistuned against the stated workload's bottleneck, a topology chosen without a bottleneck that justifies it, a rollout that breaks a caller mid-flight), rather than saying it feels wrong.

Constraints

  • Assumes comfort running a Python ML environment and basic familiarity with what a language model is; no prior serving-infrastructure experience required.
  • Core stacks: vLLM (GPU) and llama.cpp (CPU/edge), covered in that order. Other stacks (TGI, TensorRT-LLM) are mentioned only where a concept transfers differently.

Out of scope

  • Training or fine-tuning a model or adapter: see llm/finetuning (lessons 0025, serving adapters, and 0026, cost, latency and throughput), linked back to rather than restated.
  • Judging output quality or building an eval for a served model: see llm/evals for that boundary.

The arc

Fourteen stages, first request to a defended, fleet-scale, production-monitored, and reviewed serving configuration. A stage takes several lessons and the boundaries are soft; what makes a stage done is the capability, not the lesson count.

Stage Lessons Covers Done when
1. The KV cache 0001 to 0003 Why autoregressive generation is expensive, what the cache trades memory for, how cache size grows with context and batch Can compute a model's KV cache memory footprint for a given context length and batch size
2. Batching 0004 to 0006 Static vs continuous batching, request scheduling, the throughput/latency trade-off Can defend a batching configuration for a stated workload
3. Quantization at serve time 0007 to 0009 int8/int4/AWQ/GPTQ at inference time, accuracy vs speed vs memory Can pick a serving-time quantization scheme and defend the trade-off
4. vLLM in practice 0010 to 0012 Standing up vLLM, PagedAttention, the tuning knobs that matter A vLLM stack is running and answering real requests
5. llama.cpp on CPU/edge 0013 to 0015 GGUF, llama.cpp's architecture, what changes off-GPU A llama.cpp stack is running on CPU and answering real requests
6. The latency budget 0016 to 0017 p99 measurement methodology, tying the number back to cache, batching and quantization choices Can quote and defend a p99 latency budget end to end
7. Speculative decoding 0018 Draft-model, n-gram, Medusa, and EAGLE speculative decoding Can explain why parallel verification is nearly free and pick a speculative-decoding variant for a workload
8. Sampling and constrained decoding at serve time 0019 Per-request sampling parameters in a shared batch, FSM-indexed constrained decoding Can explain why constrained decoding's guarantee is structural and why its per-token cost stays small
9. Prefix caching 0020 Automatic prefix caching as cross-request block sharing, long-document and multi-turn workloads, its prefill-only limit Can explain when prefix caching helps a workload and why it never speeds up decode
10. Multi-GPU and multi-node parallelism 0021 Tensor parallelism's actual mechanism past lesson 12's single flag, pipeline parallelism, and a decision procedure for which one (or both) a deployment needs Can explain what a tensor-parallel combine step and a pipeline bubble each cost, and choose the right split for a given interconnect and node count
11. Fleet-level serving 0022 Request routing informed by cache locality, queue-depth-driven autoscaling, and why a cold start's model load time bounds how fast a new replica actually helps Can explain how routing, autoscaling, and cold-start latency have to be reasoned about as one pipeline rather than tuned separately
12. Production observability and cost 0023 Continuously exported metrics, SLIs/SLOs/error budgets that turn a metric into a target, and cost per million tokens derived from measured throughput Can defend a production configuration by citing its p99, the SLO and error budget it's held to, and its cost per million tokens
13. Long context at serve time 0024 RoPE scaling versus sliding-window attention and attention sinks, and what each actually does to lesson 2's cache-growth formula Can explain whether a long-context technique extends usable length, caps cache size to a constant, or fixes a capped cache's failure mode
14. Judgment 0025 Reviewing someone else's configuration for mismatched knobs, safe rollouts through what others depend on, and saying when speculative decoding is not a free win Trusted to review a serving configuration and name what each choice costs, not just that it feels wrong

Lessons

Work through these in order.

# Lesson Teaches
0001 The KV Cache Why generation gets expensive, and what caching buys back
0002 Capacity and Batch Size How head sharing (GQA/MQA) and batch size change the cache's real footprint
0003 Growth, Prefill, Decode, and Precision How the cache grows across prefill and decode, and what lowering its precision buys back
0004 Static vs Continuous Batching Why batching requests together helps throughput, and why continuous batching beats the static kind
0005 Request Scheduling How a continuous-batching scheduler picks the next request, and why a large prefill can stall everyone else
0006 The Throughput/Latency Trade-off How batch size trades throughput against per-token latency, and how to defend a batching configuration against a stated latency budget
0007 Quantization Schemes at Serve Time What int8, int4, GPTQ and AWQ actually do to a model's weights, and why the naive version of low-bit quantization needs a fix
0008 Accuracy, Speed, and Memory Trade-offs Why memory, speed, and accuracy don't move together when a model is quantized, and how each is actually measured
0009 Picking and Defending a Quantization Scheme A decision framework for choosing a serving-time quantization scheme, and the three numbers that defend it
0010 Standing Up vLLM Installing vLLM, launching its server, and finding the flags that carry the concepts already taught
0011 PagedAttention The block table mechanism behind vLLM, its block-size trade-off, and how it lets sequences share a common prefix's cache
0012 The vLLM Tuning Knobs That Matter Two more flags, gpu-memory-utilization and tensor-parallel-size, plus a decision procedure for which knob a symptom actually points at
0013 GGUF What llama.cpp's single-file model format bundles together, its quantization naming, and why it can be memory-mapped instead of loaded
0014 llama.cpp's Architecture The ggml tensor library underneath llama.cpp, its backend abstraction, and how CPU threading differs from GPU batching
0015 What Changes Off-GPU Standing up llama.cpp's server, and how batching, cache management, and quantization each look different at CPU/edge scale
0016 p99 Latency Measurement Methodology Why p99 beats an average, why TTFT and inter-token latency need separate numbers, and what makes a p99 measurement trustworthy
0017 Defending a Latency Budget End to End A diagnostic order for a missed p99 budget, and the three things a final defended configuration must cite
0018 Speculative Decoding Verifying several candidate tokens in one parallel forward pass costs about what generating one token normally costs, since decoding is bottlenecked on moving weights, not on the arithmetic, which is the one fact every variant of speculative decoding is built on
0019 Sampling and Constrained Decoding at Serve Time A continuous batch already runs many requests through one shared forward pass, and giving each one its own sampling settings, or forcing one of them to only ever produce valid JSON, both have to happen per-row inside that same shared pass without slowing everyone else down
0020 Prefix Caching Lesson 11 let parallel samples of the same request share one prompt's cached blocks; prefix caching is the same sharing applied across otherwise unrelated requests, and it speeds up exactly the phase that redundant prefill wastes, never the phase that generates the answer
0021 Tensor and Pipeline Parallelism Lesson 12 named tensor-parallel-size as a flag that splits a model across GPUs; this lesson opens up what that split actually does inside a layer, introduces pipeline parallelism as a different split of the same problem, and gives a decision procedure for which one (or both) a workload actually needs
0022 Fleet-Level Serving Lesson 5's scheduler picks the next request for one server's queue; a fleet adds a decision before that, which server gets the request at all, a decision about when to add more servers, and a cold-start cost that makes a freshly added server useless for tens of seconds
0023 Production Observability and Cost Lesson 17 defended one measured p99 figure for one configuration; this lesson covers what has to be continuously exported to catch a regression automatically, the SLO and error budget that turn "good enough" into a target instead of a vibe, and the cost-per-million-tokens number that has to be defended alongside p99, not instead of it
0024 Long Context at Serve Time Lesson 2's capacity formula charges memory linearly with context length; RoPE scaling extends how long a context can be without changing that charge at all, while sliding-window attention and attention sinks change the formula itself by capping context length to a constant
0025 Reviewing a Serving Configuration Review someone else's serving configuration and name specifically what it costs (a knob mistuned against the stated workload's bottleneck, a topology chosen without a bottleneck that justifies it, a rollout that breaks a caller mid-flight), rather than saying it feels wrong

Reference

  • Glossary: canonical terms for this topic
  • Resources: trusted sources
  • KV cache: per-token cost, head sharing, growth across a request, and capacity planning
  • Batching: static vs continuous batching, request scheduling, and the throughput/latency trade-off
  • Quantization at Serve Time: int8/int4/GPTQ/AWQ, why memory, speed and accuracy don't move together, and a decision framework for picking a scheme
  • vLLM: standing up vLLM, PagedAttention, and the tuning knobs that matter
  • llama.cpp: GGUF, llama.cpp's architecture, and what changes off-GPU
  • Latency Budget: p99 measurement methodology, and tying a missed budget back to cache, batching and quantization choices

How this works

Each lesson is short and self-contained. Answer keys are collapsed: recall first, then open them. The real-world reps matter more than the reading, and spacing them out is the point. Anything still unclear at the end of a lesson is worth chasing to its primary source before moving on.

Table of contents