https://en.wikipedia.org/wiki/Quantization_(signal_processing)
Roughly speaking: The AI equivalent of reducing bitrate. Works quite well if you're only running them in inference mode and don't want to train them as the networks are quite noise-resistant (rounding all weights is, in essence, introducing noise).