सभी किताबें

पेशेवर

ऐप्स परिचय Coach लॉग इन पढ़ना शुरू करें

Quantitative Finance · शब्दावली

Operator fusion, quantisation, post-training quantisation, quantisation-aware training क्या है?

अन्य नाम: operator fusion · quantisation · post-training quantisation · quantisation-aware training

Definition 26.3 Machine Learning for Markets · अध्याय 26 — Low-Latency Inference

Operator fusion computes several consecutive operations (a matrix product, a bias, an activation, a rescaling) in one pass over the data. Quantisation represents weights and activations as small integers with scale factors, so that inference runs in integer arithmetic (Jacob and co-authors, 2018). Post-training quantisation derives the integers and scales from a trained float model and a calibration sample; quantisation-aware training fine-tunes the model with the rounding simulated in the forward pass, so that it learns weights that survive it (Nagel and co-authors, 2021).

bits per weight and activation
rank IC on 20 000 new observations864
quantised after training0.1820.1810.143
quantisation-aware fine-tuning, then quantised0.1790.1760.164
Table 26.1. The network’s information coefficient after quantisation (float: 0.182; the forest: 0.224). Data: ml_infer.accuracy.
अध्याय में पढ़ें →