AI Glossary

Plain-English definitions of the AI terms used across our articles. No jargon, no maths – just what each concept means and why it matters.

Pick a term below to read more.

  • Token: The smallest unit of text a model reads or writes (roughly ~4 characters).
  • Quantization: Shrinking a model by storing its numbers at lower precision – smaller, faster, but rarely free.
  • NVFP4: A 4-bit floating-point format from NVIDIA – about 4x smaller than BF16 with much of the accuracy kept.
  • BF16 (bfloat16): The standard 16-bit “brain float” format – the full-quality baseline models are compared against.
  • Speculative Decoding: A small fast model guesses several tokens ahead; the big model verifies them all at once.
  • Drafter (Draft Model): The small fast helper model that proposes tokens in speculative decoding.
  • D-flash: A speculative-decoding method that proposes up to ~7 tokens per step.
  • MTP (Multi-Token Prediction): A simpler way to emit 2-3 tokens per step instead of one.
  • Acceptance Length: How many drafted tokens per step the big model actually accepts – higher means faster.
  • Prefill & Decode: Prefill reads the whole prompt at once; decode generates the answer token by token.
  • KV Cache: Stored values that stop the model from recomputing earlier text – the reason long chats stay fast.
  • Context Window: The maximum amount of text (in tokens) the model can consider at once.
  • Throughput (tokens/second): How many tokens the model produces per second – the speed of generation.
  • Needle-in-a-Haystack: A test that hides a small fact in a huge text and checks whether the model can retrieve it.
  • Tool Use (Agentic AI): When a model calls external tools – search, code, files – and chains steps to finish a task.
  • Inference Engine (SGLang, vLLM): The software that runs a model efficiently on a GPU – and it is not neutral.
  • Uncensored Model: A model with fewer safety filters – often re-quantized by third parties, which can quietly lower quality.
  • Mixture of Experts (MoE): A model with many expert sub-networks where only a few activate per token – fast for its size.
  • Dense Model: A standard model that uses all its parameters for every token – smaller, but often more consistent.
  • Battle Arena & Elo: A simulated tournament where models build armies and fight – Elo is the ranking score.