Flagship module · compute → memory → bandwidth · 7 beats

KV Cache

The deployment problem is concrete: a 70B model fits across 8×80 GiB accelerators, but one 128K conversation adds 40 GiB of state and eight add 320 GiB before allocator overhead. You will derive where those bytes come from, distinguish persistent cache from transient GQA expansion, decide whether weights or cache dominate each decode operating point, and select among paging, quantization, eviction, reuse, GQA, and MLA by the exact term each changes. Published claims, deterministic arithmetic, and our interpretations remain visibly separate.

Core path ~31 min: 01 → 04 then 07 · Systems deep dive +27 min: 05 → 06

Beat 01 / 07CONCEPT~5 min

Discover the wasted work

After this beat you can: You can state, in cost terms, why naive autoregressive decoding is untenable — because you found the waste yourself.

Training a decoder runs one forward pass over a whole sequence: every position computed at once, the quadratic attention cost paid once. Generation is a different regime. You produce one token, append it, and run the model again — T times for T tokens.

Before any complexity classes, work the trace below: a five-token prompt, one token just generated, and one question about what the model is about to do.

Work the trace before reading on

Prompt: [The] [capital] [of] [France] [is] — the model just emitted "Paris". To generate the next token, it runs attention again. Which quantities computed from "The capital of France is" have changed since the previous decode step?

Prove it before continuing

Caching K and V removes which cost — and leaves which untouched?

Answer the question above to continue — committing unlocks, right or wrong.