Signal & Noise
Menu

Quantization

Storing model weights at lower precision (4–8 bit instead of 16) to cut memory 2–4x with small quality loss. What makes local inference practical.

Related terms

← Back to the full glossary