KV Cache

When a model reads text it computes “keys” and “values” for every token. The KV cache stores these so the model does not have to redo the work for text it has already seen.

A cache hit makes prefill hundreds of times faster. Caching is also why benchmarks must be careful: if answers are cached, a model can look perfect for the wrong reason.