KV Cache
The deployment problem is concrete: a 70B model fits across 8×80 GiB accelerators, but one 128K conversation adds 40 GiB of state and eight add 320 GiB before allocator overhead. You will derive where those bytes come from, distinguish persistent cache from transient GQA expansion, decide whether weights or cache dominate each decode operating point, and select among paging, quantization, eviction, reuse, GQA, and MLA by the exact term each changes. Published claims, deterministic arithmetic, and our interpretations remain visibly separate.
Core path ~31 min: 01 → 04 then 07 · Systems deep dive +27 min: 05 → 06
Discover the wasted work
After this beat you can: You can state, in cost terms, why naive autoregressive decoding is untenable — because you found the waste yourself.
Training a decoder runs one forward pass over a whole sequence: every position computed at once, the quadratic attention cost paid once. Generation is a different regime. You produce one token, append it, and run the model again — T times for T tokens.
Before any complexity classes, work the trace below: a five-token prompt, one token just generated, and one question about what the model is about to do.
Prompt: [The] [capital] [of] [France] [is] — the model just emitted "Paris". To generate the next token, it runs attention again. Which quantities computed from "The capital of France is" have changed since the previous decode step?
Caching K and V removes which cost — and leaves which untouched?