Size & speed

quantization

Shrinking a model by storing its weights at lower precision (e.g. 4-bit instead of 16-bit) so it fits a smaller machine, at a small quality cost.

Return to all terms