BF16 (bfloat16)

BF16 (bfloat16) is a 16-bit floating-point format that keeps enough precision for training and high-quality inference. It is the reference point: when we compare a quantized model to “full quality”, we mean BF16.

BF16 uses more memory and runs slower than NVFP4, but it is the safest choice when accuracy matters most.