Speculative Decoding

Speculative decoding is a trick to make models generate text faster. A small, quick model (the drafter) guesses the next few tokens. The large model then checks all of those guesses in a single pass.

If the guesses are accepted, you effectively get several tokens for the price of one step – a big speed-up. If a guess is wrong, its tail is thrown away and resampled. In our benchmark this produced roughly a 2.8x speed-up, but only when the GPU had spare power left.