Alle boeken

Professioneel

Apps Over Coach Inloggen Begin met lezen

Quantitative Finance · Begrippenlijst

Wat is Training–serving skew?

Ook bekend als: training--serving skew

Definition 24.3 Machine Learning for Markets · Hoofdstuk 24 — Data and Feature Stores

Training–serving skew is any difference between the feature values a model was trained on and those it receives in production for the same moment: different code, different clocks, different inputs, different arithmetic or different data timing.

RMS difference as a share of the feature’s standard deviation
featureone event aheadone-second barsfloat32late trades
spread0.380.9200
imbalance0.240.730.000
depth imbalance0.110.410.000
weighted mid minus mid0.300.790.000
OFI 5 s0.030.2400
OFI 30 s0.010.0700
volume 30 s0.010.0700.07
signed volume 30 s0.020.1000.11
trades 10 s0.020.1800.17
mid change 5 s0.090.3900
mid change 30 s0.030.1300
messages 10 s0.000.0400
Table 24.1. Size of each planted skew, feature by feature, over 7 714 decisions (0.00: a difference smaller than half a hundredth; 0: none). Data: ml_featstore.feature_skew.
decisionsdetectedrank IClive P&L
skewmismatched(5-sample check)backtestlive(ticks)
none000.4110.4110.089
research one event ahead100%100%0.3950.4180.091
research on one-second bars96%100%0.4380.3600.078
production float32 storage92%100%0.4110.4110.089
production trade prints 200 ms late100%100%0.4110.4110.089
Table 24.2. Parity test (tolerance 10−910^{-9}) and a ridge forecast of the mid’s change over the next second, trained on four sessions with each research pipeline and tested on four others: its backtest on research features and its live result on production features, rank IC and mean P&L per decision of trading the forecast’s sign. Data: ml_featstore.parity_table, model_effects.
Lees in het hoofdstuk →