Inference Engine (SGLang, vLLM)

An inference engine is the software that actually runs a model on your hardware efficiently. Popular options include SGLang and vLLM (sometimes labeled “VLM” in benchmarks).

The engine matters: the same model can behave and perform differently on a different engine, and specific quantizations are built for specific engines.