AI Glossary
Plain-English definitions of the AI terms used across our articles. No jargon, no maths – just what each concept means and why it matters.
Pick a term below to read more.
- Token: The smallest unit of text a model reads or writes (roughly ~4 characters).
- Quantization: Shrinking a model by storing its numbers at lower precision – smaller, faster, but rarely free.
- NVFP4: A 4-bit floating-point format from NVIDIA – about 4x smaller than BF16 with much of the accuracy kept.
- BF16 (bfloat16): The standard 16-bit “brain float” format – the full-quality baseline models are compared against.
- Speculative Decoding: A small fast model guesses several tokens ahead; the big model verifies them all at once.
- Drafter (Draft Model): The small fast helper model that proposes tokens in speculative decoding.
- D-flash: A speculative-decoding method that proposes up to ~7 tokens per step.
- MTP (Multi-Token Prediction): A simpler way to emit 2-3 tokens per step instead of one.
- Acceptance Length: How many drafted tokens per step the big model actually accepts – higher means faster.
- Prefill & Decode: Prefill reads the whole prompt at once; decode generates the answer token by token.
- KV Cache: Stored values that stop the model from recomputing earlier text – the reason long chats stay fast.
- Context Window: The maximum amount of text (in tokens) the model can consider at once.
- Throughput (tokens/second): How many tokens the model produces per second – the speed of generation.
- Needle-in-a-Haystack: A test that hides a small fact in a huge text and checks whether the model can retrieve it.
- Tool Use (Agentic AI): When a model calls external tools – search, code, files – and chains steps to finish a task.
- Inference Engine (SGLang, vLLM): The software that runs a model efficiently on a GPU – and it is not neutral.
- Uncensored Model: A model with fewer safety filters – often re-quantized by third parties, which can quietly lower quality.
- Mixture of Experts (MoE): A model with many expert sub-networks where only a few activate per token – fast for its size.
- Dense Model: A standard model that uses all its parameters for every token – smaller, but often more consistent.
- Battle Arena & Elo: A simulated tournament where models build armies and fight – Elo is the ranking score.
