Prefill & Decode
Running an AI model has two phases. Prefill processes your entire prompt in parallel – fast, measured in tokens per second. Decode then generates the answer one token at a time, which is slower and grows heavier as context grows.
That is why a long conversation feels slower: more context to carry means slower decode and prefill degradation.
