Training
Quantization
Reducing the numerical precision a model uses (for example from 32-bit to 4-bit) to make it faster and smaller in memory.
Reducing the numerical precision a model uses (for example from 32-bit to 4-bit) to make it faster and smaller in memory.