Operator fusion computes several consecutive operations (a matrix product, a bias, an activation, a rescaling) in one pass over the data. Quantisation represents weights and activations as small integers with scale factors, so that inference runs in integer arithmetic (Jacob and co-authors, 2018). Post-training quantisation derives the integers and scales from a trained float model and a calibration sample; quantisation-aware training fine-tunes the model with the rounding simulated in the forward pass, so that it learns weights that survive it (Nagel and co-authors, 2021).
| bits per weight and activation | |||
| rank IC on 20 000 new observations | 8 | 6 | 4 |
| quantised after training | 0.182 | 0.181 | 0.143 |
| quantisation-aware fine-tuning, then quantised | 0.179 | 0.176 | 0.164 |
ml_infer.accuracy.