Throughput (tokens/second)
Throughput is simply how many tokens a model generates per second (tok/s). Higher means faster answers and snappier AI agents.
It is helped by quantization and speculative decoding, and hurt by long context (see prefill & decode).
