Signal & Noise
Menu

Decoding

The token-by-token generation phase of inference, after the prompt is processed. Sequential and memory-bound, it's where most serving latency lives.

Related terms

← Back to the full glossary