جميع الكتب

مهني

1 Markets I: The Ecosystem and Exchange-Traded Marketsالأسواق عبر الإنترنت 2 Markets II: Rates, FX and Creditالأسواق عبر الإنترنت 3 Markets III: Commodities, Energy and Cryptoالأسواق عبر الإنترنت 4 Quantitative Methodsالأساليب عبر الإنترنت 5 Derivatives and Volatilityالمشتقات عبر الإنترنت 6 Rates, Credit, XVA and Riskالفائدة والائتمان والمخاطر عبر الإنترنت 7 Research Craft: Predictors, Backtests, Measurement, Portfoliosالبحث عبر الإنترنت 8 Strategies I: Equities and Futuresالاستراتيجيات عبر الإنترنت 9 Strategies II: Volatility, Relative Value, Macro and the Bank Desksالاستراتيجيات عبر الإنترنت 10 Microstructure and Executionالتنفيذ عبر الإنترنت 11 Market Making and High-Frequency Tradingصناعة السوق عبر الإنترنت 12 Machine Learning for Marketsتعلم الآلة عبر الإنترنت 13 Low-Latency Softwareالتكنولوجيا عبر الإنترنت 14 Networks, Hardware and Trading Infrastructureالتكنولوجيا عبر الإنترنت 15 Research, Data and Risk Platformsالتكنولوجيا عبر الإنترنت 16 The Desk and the Firmالشركة عبر الإنترنت 17 The Industry: Firms, Roles and Careersالمسارات المهنية عبر الإنترنت 18 The Interview Bookالمسارات المهنية عبر الإنترنت
التطبيقات حول المدرب تسجيل الدخول ابدأ القراءة

Quantitative Finance · المسرد

ما معنى Attention mechanism, transformer؟

يُعرف أيضًا باسم: attention mechanism · transformer

Definition 8.4 Machine Learning for Markets · الفصل 8 — Sequence Models on Order Books

An attention mechanism lets each position of a sequence form its output as a weighted average of all positions: with queries Q=ZAQ\mathsf Q = ZA_{\mathsf Q}, keys K=ZAK\mathsf K = ZA_{\mathsf K} and values V=ZAV\mathsf V = ZA_{\mathsf V} computed from the sequence ZZ (T×hT\times h), the output is softmax(QK⊤/h)V\mathrm{softmax}(\mathsf Q\mathsf K^\top/\sqrt h)\mathsf V. A transformer layer combines several attention heads, a position-wise network, residual connections and layer normalisation; positions are encoded by adding learned or fixed vectors to the inputs (Vaswani et al., 2017).

Information coefficient of ( up) - ( down) with the forward change, averaged over eight test sessions (and two seeds for the sequence models); bars show the standard deviation across sessions and seeds. The logistic regression reads ten hand-made features at the decision time; the sequence models read the raw ten-level book over ten seconds. Data: ml_lobseq.table.
Figure 8.2. Information coefficient of P(up)−P(down)\P(\text{up}) - \P(\text{down}) with the forward change, averaged over eight test sessions (and two seeds for the sequence models); bars show the standard deviation across sessions and seeds. The logistic regression reads ten hand-made features at the decision time; the sequence models read the raw ten-level book over ten seconds. Data: ml_lobseq.table.
اقرأ في الفصل →