Size & speed
quantization
Shrinking a model by storing its weights at lower precision (e.g. 4-bit instead of 16-bit) so it fits a smaller machine, at a small quality cost.
Size & speed
Shrinking a model by storing its weights at lower precision (e.g. 4-bit instead of 16-bit) so it fits a smaller machine, at a small quality cost.