Alle boeken

Professioneel

Apps Over Coach Inloggen Begin met lezen

Quantitative Finance · Begrippenlijst

Wat is Attention mechanism, transformer?

Ook bekend als: attention mechanism · transformer

Definition 8.4 Machine Learning for Markets · Hoofdstuk 8 — Sequence Models on Order Books

An attention mechanism lets each position of a sequence form its output as a weighted average of all positions: with queries Q=ZAQ\mathsf Q = ZA_{\mathsf Q}, keys K=ZAK\mathsf K = ZA_{\mathsf K} and values V=ZAV\mathsf V = ZA_{\mathsf V} computed from the sequence ZZ (T×hT\times h), the output is softmax(QK⊤/h)V\mathrm{softmax}(\mathsf Q\mathsf K^\top/\sqrt h)\mathsf V. A transformer layer combines several attention heads, a position-wise network, residual connections and layer normalisation; positions are encoded by adding learned or fixed vectors to the inputs (Vaswani et al., 2017).

Information coefficient of ( up) - ( down) with the forward change, averaged over eight test sessions (and two seeds for the sequence models); bars show the standard deviation across sessions and seeds. The logistic regression reads ten hand-made features at the decision time; the sequence models read the raw ten-level book over ten seconds. Data: ml_lobseq.table.
Figure 8.2. Information coefficient of P(up)−P(down)\P(\text{up}) - \P(\text{down}) with the forward change, averaged over eight test sessions (and two seeds for the sequence models); bars show the standard deviation across sessions and seeds. The logistic regression reads ten hand-made features at the decision time; the sequence models read the raw ten-level book over ten seconds. Data: ml_lobseq.table.
Lees in het hoofdstuk →