Quantization
Storing model weights at lower precision (4–8 bit instead of 16) to cut memory 2–4x with small quality loss. What makes local inference practical.
Storing model weights at lower precision (4–8 bit instead of 16) to cut memory 2–4x with small quality loss. What makes local inference practical.