Quantization at serve time — AI Dictionary

Compressing weights to fewer bits before deployment, shrinking load time and memory footprint at some accuracy cost.