Todos os livros

Profissional

Apps Sobre Coach Entrar Começar a ler

Quantitative Finance · Glossário

O que é Operator fusion, quantisation, post-training quantisation, quantisation-aware training?

Também chamado de: operator fusion · quantisation · post-training quantisation · quantisation-aware training

Definition 26.3 Machine Learning for Markets · Capítulo 26 — Low-Latency Inference

Operator fusion computes several consecutive operations (a matrix product, a bias, an activation, a rescaling) in one pass over the data. Quantisation represents weights and activations as small integers with scale factors, so that inference runs in integer arithmetic (Jacob and co-authors, 2018). Post-training quantisation derives the integers and scales from a trained float model and a calibration sample; quantisation-aware training fine-tunes the model with the rounding simulated in the forward pass, so that it learns weights that survive it (Nagel and co-authors, 2021).

bits per weight and activation
rank IC on 20 000 new observations864
quantised after training0.1820.1810.143
quantisation-aware fine-tuning, then quantised0.1790.1760.164
Table 26.1. The network’s information coefficient after quantisation (float: 0.182; the forest: 0.224). Data: ml_infer.accuracy.
Ler no capítulo →