Quantization
Quantization means storing a model’s numbers (its weights) with less precision. Instead of full 16-bit values (BF16) you might keep 8-bit or 4-bit ones. The payoff: the model becomes much smaller, uses less memory and usually runs faster.
The trade-off is quality. Think of it like saving a photo as a smaller JPEG – most of the time it looks fine, but detail can degrade.
The catch: not every quantization is equal. One made for a specific model and engine keeps almost all the accuracy; a random file from the internet can silently cost a lot. In our benchmark, the same base model scored 228 vs 330 in math reasoning purely because of a different quantization.
