Machine Learning for Markets · Machine learning
1Why Financial Machine Learning Is Different
A new model forecasts next month’s return of 500 stocks. Out of sample, over ten years, its R-squared is 0.57%, and the desk is pleased. It should be: in the simulated market the model was trained on, where the expected returns are known because they were planted, the best forecast that can exist scores 0.60%. The same gradient-boosted trees, given a task whose signal is strong, explain 92% of the variance. Nothing is wrong with either number. The model did as well on the stocks as it did on the easy task; the stocks simply contain almost nothing to learn, and what they contain changes, is shared by every stock at once, and is traded away by whoever learns it first. This chapter fixes the vocabulary of machine learning the rest of the book uses, and measures, on data whose truth is known, the four ways market data differ from the data the field grew up on.
1.1 Learning from data: the vocabulary
Definition 1.1 (Supervised and unsupervised learning)
Supervised learning fits a function from features to a target on examples , , by minimising an empirical loss , so that predicts for new . Unsupervised learning has no target: it describes the distribution of itself (clusters, factors, a lower-dimensional representation, a density).
The target is Book 7’s prediction target (chapter 6): a forward return over a forecast horizon, or a label built from one (chapter 2 of this book). The loss always takes two arguments here; is always subscripted. The learning rate of later chapters is written , a symbol declared local to this book (the series keeps for the vol-of-vol elsewhere).
Definition 1.2 (Training, validation and test sets; hyperparameter)
The training set is the data on which is fitted. A hyperparameter is a choice the fit does not make itself (a penalty, a tree depth, a number of layers, a learning rate); the validation set is data held out from the fit on which hyperparameters are chosen. The test set is data used once, after every choice is made, to estimate how the chosen model will perform.
On market data the three sets are consecutive stretches of time, in that order (Figure 1.1), because the model will be used on a future it has not seen. The test set is Book 7’s holdout set (chapter 1): touched once. A test set consulted twice has become a validation set.
Definition 1.3 (Generalisation error, overfitting, bias–variance decomposition)
The generalisation error of a fitted model is its expected loss on a new draw from the distribution the model will meet. Overfitting is the fitting of features of the training sample that do not recur, so that the in-sample loss understates the generalisation error. For the squared loss, the bias–variance decomposition splits the expected generalisation error at a point into noise, squared bias and variance (Proposition 1.4).
Proposition 1.4 (Bias–variance decomposition)
Let with and , and let be fitted on a training sample independent of . Then
where and of are over training samples.
Proof. Write . The three terms are uncorrelated given : is independent of the training sample and has mean zero, and has mean zero over training samples while is a constant. Square and take expectations. ∎
The noise is the floor no model goes below. In most of machine learning the noise is small and the fight is between bias (a model too rigid for ) and variance (a model that follows the sample). On market data is almost all of the loss, and a model with any variance at all spends it on noise: the balance tips towards rigid, regularised models (chapter 4).
1.2 Signal-to-noise near zero and the R-squared ceiling
Definition 1.5 (Out-of-sample R-squared, predictability ceiling)
The out-of-sample R-squared of forecasts of returns on a test set is
measured against a forecast of zero, not against the sample mean. The predictability ceiling of a dataset is of the true conditional expectation : the best any model of those features can score.
The zero benchmark follows Gu, Kelly and Xiu, whose study of 30 000 stocks over 60 years uses it because the historical mean of a stock’s return is so noisy a forecast that beating it is too easy. Their monthly out-of-sample stock-level R-squared runs from 0.16% for a three-variable linear model to 0.33–0.40% for trees and neural networks, with 0.40% for the best network. Those are the numbers of a field where half a per cent is a strong result.
Proposition 1.6 (The ceiling bounds every model; R-squared and correlation)
- For any forecast built from the features, : in expectation, no model’s exceeds the ceiling.
- If and a forecast have mean zero, , and the forecast is used as , then , maximal at where it equals , and zero at .
Proof. (1) because is uncorrelated with every function of . (2) Expand and maximise the quadratic in . ∎
The second part links R-squared to the information coefficient of Book 7 (chapter 6): a forecast whose correlation with returns is earns at most , and a correlation of 0.075 is an R-squared of 0.56%. It also says how a good forecast scores badly: a forecast with the right direction and twice the right size scores zero.
The book’s sandbox, firm.mlsynth, draws a monthly cross-section of 500 stocks with twenty characteristics published as cross-sectional ranks in , as the empirical literature ranks them. Six carry signal, linearly and through an interaction, a square, a threshold and an absolute value; fourteen are noise. Returns add a market factor, ten industry factors and heavy-tailed specific shocks, and the planted expected return is scaled so that the ceiling is 0.7%. Ridge regression, LightGBM’s gradient-boosted trees and a small multilayer perceptron are trained on twenty years and tested on the next ten (Table 1.1); for comparison, the same three models are fitted to a task with a Bayes-optimal R-squared of 0.95 (20 000 examples, a smooth nonlinear function of five of twenty features).
| 500 stocks, ceiling 0.60% out of sample | strong signal, Bayes 0.95 | ||||
| model | in sample | out of sample | corr. with the truth | in sample | out of sample |
| ridge regression | 0.39% | 0.38% | 0.74 | 0.72 | 0.72 |
| boosted trees | 1.45% | 0.57% | 0.92 | 0.92 | 0.92 |
| neural network | 0.64% | 0.28% | 0.74 | 0.94 | 0.93 |
firm.mlsynth’s 500 stocks, and 20 000 training examples of a task with a strong signal. R-squared against zero for the stocks, against the mean for the generic task. Data: ml_why.compare and ml_why.strong_signal.Table 1.1 is the chapter in miniature. On the strong task every model’s in-sample and out-of-sample scores agree and the flexible models win. On the stocks, the boosted trees’ in-sample R-squared is 1.45%, twice the ceiling of the training years (0.72%): the excess is noise it has fitted. Its out-of-sample 0.57% is 95% of the ceiling, the best of the three; the linear model reaches 64% of it, and the network, trained here without the care of chapter 7, less than half. The forecasts’ correlations with the truth (0.74 to 0.92) are high; their correlations with the realised returns are below 0.08.
ml_why.compare on firm.mlsynth.1.3 Non-stationarity and adversarial feedback
The distribution a model meets is not the one it was trained on. Some change is exogenous (a new regulation, a new venue, a crisis). Some is caused by the model’s own trade and its competitors’: a predictor that is public, or that several firms have found, is traded until its return is smaller than its cost. Book 7 measured the result on real anomalies as post-publication decay (chapter 13) and its cause as crowding (chapter 28). In the physical sciences the data do not read the paper.
firm.mlsynth can age a signal. From the first test month the strongest linear characteristic’s effect decays with a half-life of two years, and five years into the test the interaction between two characteristics changes sign. The boosted trees, trained once on the first twenty years, keep their R-squared through the decay, which the other characteristics hide, and lose it when the interaction flips (Figure 1.3): years 6 to 10 average 0.01%, while the ceiling, which moves with the new truth, averages 0.83%. Nothing in the model’s inputs announced the change; its errors did. Chapter 12 detects it and chapter 27 builds the alarms.
ml_why.drift_by_year.1.4 Small effective samples
The training set holds stock-months. It is not 120 000 independent observations. The stocks’ unexpected returns share a market and an industry: their average pairwise correlation is 0.27, and the effective number of independent stocks in a month, , is 3.7 out of 500. Whatever depends on the market’s path is estimated from 240 draws, not 120 000.
The intercept is the clearest case. Trained on raw returns, a model’s intercept is the training window’s average market return. The true premium in the sandbox is zero; the window happened to average a month, about two standard errors of a 240-month mean at a 4.5% monthly volatility. The three models trained on raw returns score , and out of sample, each of them worse than a forecast of zero. Trained on returns minus each month’s cross-sectional mean, they learn which stocks do better, not where the market goes, and score the table’s numbers.
Method 1.7 (Before fitting a model of returns)
- Compute or bound the ceiling: on synthetic data exactly; on real data, from the best published R-squared or information coefficient of comparable work (Proposition 1.6).
- Decide what the model is asked: the cross-section (demean or rank the target by date), or the level.
- Count the effective sample: dates, not rows, for anything common to all assets; overlapping labels (chapter 2) divide it again.
- Fix the test period before looking at it, and plan how long a live track must be to tell the model from zero and from the simpler model.
Length is the other effective sample. The boosted trees’ monthly R-squared has a mean of 0.62% and a standard deviation of 0.64%: 4.3 months of live trading put its mean two standard errors from zero. Telling it from ridge regression takes longer: the monthly difference has a mean of 0.21% and a standard deviation of 0.47%, which needs 21 months. With less history the flexible model loses its advantage (Figure 1.4): trained on the last two years only, the trees score 0.18% and ridge 0.29%; they cross between four and eight years. The more flexible the model, the more of its training data goes to learning what the noise is not.
ml_why.learning_curve.1.5 What carries over from the rest of machine learning, and what does not
The algorithms carry over: least squares and its penalties (Book 4, chapter 16), trees, networks, gradient methods (Book 4, chapter 24). So does the discipline of held-out evaluation. What does not carry over is the default settings and the intuitions trained on images and text: shuffled cross-validation (Book 7, chapter 20, and chapter 3), a validation score read as a fact rather than a draw, hyperparameters tuned until the validation set agrees, an accuracy of 55% taken for a weak model when it may be close to the ceiling, and a model fitted once and trusted for years. Every chapter of this book is one of these corrections, measured.
1.6 Tutorial: the 0.57% model
Goal. Measure the ceiling of a synthetic stock panel, compare three models in and out of sample, and repeat on a task with a strong signal. End state: the table, Figures 1.2, 1.3 and 1.4.
The panel. Characteristics are ranks of persistent latent processes; the truth is a linear and a nonlinear part, scaled so that hits the ceiling.
non = (sgn * 1.5 * C[:, 0] * C[:, 3] + 1.2 * (C[:, 4] ** 2 - 1 / 3) + 0.8 * ((C[:, 5] > 0.5) - 0.25) - 0.8 * (np.abs(C[:, 1]) - 0.5)) return lin, non def panel(cfg: PanelConfig | None = None) -> Panel: cfg = cfg or PanelConfig() rng = np.random.default_rng(cfg.seed) T, n, k = cfg.months, cfg.n, cfg.k phi = np.linspace(0.60, 0.99, k) rng.shuffle(phi[6:]) # noise characteristics: mixed persistence phi[:6] = [0.95, 0.90, 0.80, 0.97, 0.85, 0.70] L = rng.standard_normal((n, k)) X = np.empty((T, n, k)) for t in range(T): L = phi * L + np.sqrt(1 - phi**2) * rng.standard_normal((n, k)) X[t] = _ranks(L.T).T beta = np.clip(1.0 + 0.3 * rng.standard_normal(n), 0.2, 2.0) industry = rng.integers(0, cfg.n_ind, n) svol = cfg.spec_vol * np.exp(0.35 * rng.standard_normal(n) - 0.5 * 0.35**2) mkt = cfg.mkt_vol * rng.standard_normal(T) ind = cfg.ind_vol * rng.standard_normal((T, cfg.n_ind)) tsc = math.sqrt((cfg.t_df - 2) / cfg.t_df) eps = svol * tsc * rng.standard_t(cfg.t_df, (T, n)) noise = beta * mkt[:, None] + ind[:, industry] + eps if cfg.style_vol > 0: # characteristic-sorted books carry factor risk fs = cfg.style_vol * rng.standard_normal((T, k)) noise = noise + np.einsum("tnk,tk->tn", X, fs) g = np.empty((T, n)) for t in range(T): lin, non = _alpha_shape(X[t], t, cfg) lz = lin / lin.std() nz = non / non.std()Listing 1.1. The synthetic stock panel with a planted expected return. code/firm/mlsynth/firm_mlsynth.py The target and the three models. Train on returns minus each month’s cross-sectional mean; score against zero.
@functools.lru_cache(maxsize=4) def compare(seed=1, target="xs"): """In- and out-of-sample R-squared of the three models on the stock panel, and the ceiling. target 'xs' trains on cross-sectionally demeaned returns, 'raw' on the returns themselves (the intercept is then the training window's average market return).""" P = stock_panel(seed) X, r, mu, _ = P.flat(TRAIN) Xt, rt, mut, mt = P.flat(TEST) y = xs_target(P, TRAIN) if target == "xs" else r out = {"ceiling_is": ceiling(P, TRAIN), "ceiling_oos": ceiling(P, TEST), "months_oos": {}} for name, fit in FITS.items(): m = fit(X, y) p, pt = predict(m, X), predict(m, Xt) out[name] = {"is": r2_oos(r, p), "oos": r2_oos(rt, pt), "corr_truth": float(np.corrcoef(pt, mut)[0, 1])} out["months_oos"][name] = np.array([r2_oos(rt[mt == t], pt[mt == t]) for t in TEST]) out["months_truth"] = np.array([r2_oos(rt[mt == t], mut[mt == t]) for t in TEST]) return outListing 1.2. In- and out-of-sample R-squared, and the ceiling. code/ml/01-why-financial-machine-learning-is-different/python/ml_why.py - Run
compare(),compare(target=’raw’),strong_signal(),learning_curve(),effective_stocks(),drift_by_year()andfig_why.py.
What to change next. Lower the strong task’s signal-to-noise ratio until the Bayes R-squared is 1% and see which model wins with 20 000 examples (exercise 7); raise nonlin to 0.9 and see whether ridge still reaches half the ceiling.
1.7 Build: synthetic learning tasks
Purpose. Every model of this book is scored against the best predictor that exists for its data.
Interface. task(n, p, snr, kind, seed) with its Bayes R-squared; PanelConfig(n, months, k, seed, ceiling, nonlin, premium, mkt_vol, ind_vol, spec_vol, t_df, decay_from, decay_half_life, regime_at); panel(cfg) -> Panel with X (months, stocks, characteristics), r, mu, beta, industry, mkt, and Panel.flat(months); r2_oos, ceiling, months_to_detect.
Rules. Characteristics at month are known at its end and predict the return of month ; the truth is never an input; seeded and deterministic.
Acceptance tests. code/firm/mlsynth/tests/: the truth scores the Bayes R-squared on the generic task; the ceiling is hit on average over seeds; the noise characteristics carry (almost) nothing; decay and flip change the truth; the R-squared and detection formulas on small cases.
Stretch. Characteristics observed with error, so that the ceiling from the features is below the truth’s; missing values and listings that start and end, as in firm.synthmkt.
Sources and further reading
- S. Gu, B. Kelly and D. Xiu, “Empirical asset pricing via machine learning”, Review of Financial Studies 33(5), 2020 (NBER Working Paper 25398, 2018, revised 2019, Table 1).
- S. Geman, E. Bienenstock and R. Doursat, “Neural networks and the bias/variance dilemma”, Neural Computation 4(1), 1992.
- T. Hastie, R. Tibshirani and J. Friedman, The Elements of Statistical Learning, 2nd ed., Springer, 2009.
- R. Israel, B. Kelly and T. Moskowitz, “Can machines ‘learn’ finance?”, Journal of Investment Management, 2020 (SSRN 3624052).
- M. López de Prado, Advances in Financial Machine Learning, Wiley, 2018.
1.8 Exercises
Exercise 1.1 ★
A forecast of next month’s returns has a correlation of 0.05 with them and is optimally scaled. What is its out-of-sample R-squared? What if it is used at three times the optimal scale?
Solution
Solution of Exercise 1.1.
. At , ; for , .
Exercise 1.2 ★
The unexpected returns of 500 stocks have an average pairwise correlation of 0.27. What is the effective number of independent stocks in a month? With 0.05?
Solution
Solution of Exercise 1.2.
; with 0.05, .
Exercise 1.3 ★
A model’s monthly R-squared has a mean of 0.5% and a standard deviation of 0.8%. How many months of independent results put the mean two standard errors from zero? Three?
Solution
Solution of Exercise 1.3.
months; for three standard errors .
Exercise 1.4 ★★
The boosted trees score 1.45% in sample, where the ceiling is 0.72%. How can a model beat the best possible predictor in sample, and what does the gap say about its out-of-sample score?
Solution
Solution of Exercise 1.4.
The ceiling is the R-squared of the truth; a model fitted to the same sample can also fit that sample’s noise, which the truth does not, so its in-sample R-squared can exceed the ceiling. The excess (0.73 points) is fitted noise: it is the variance term of Proposition 1.4 seen from inside the sample, and it will not recur. Out of sample the trees cannot beat 0.60% in expectation and score 0.57%.
Exercise 1.5 ★★
The training window’s average market return was a month, and the mean squared monthly stock return is . What is the standard error of a 240-month mean at a monthly volatility of 4.5%, and how many points of R-squared does an intercept error of cost?
Solution
Solution of Exercise 1.5.
: the window’s is 1.9 standard errors from the true zero. A constant error adds to the mean squared error: points of R-squared. The test decade’s returns averaged , so the cross term points doubles the cost: about the 0.67 points between ridge’s 0.38% and .
Exercise 1.6 ★★
Find the flaw. “We trained our model on raw monthly returns and its out-of-sample R-squared is negative, so the characteristics carry no information about returns.”
Solution
Solution of Exercise 1.6.
A model of raw returns learns an intercept equal to the training window’s average return, which is the market’s path, estimated from 240 draws; that error alone can make R-squared negative (exercise 5). Train on returns minus each month’s cross-sectional mean, or score the cross-section separately: the same characteristics then give 0.38–0.57%.
Exercise 1.7 ★★★
Coding. Run ml_why.strong_signal with a signal-to-noise ratio that gives a Bayes R-squared of 1% (snr = 0.0101/0.9899), 20 000 training examples. Report the three models’ out-of-sample R-squared and explain the ranking.
Solution
Solution of Exercise 1.7.
strong_signal(snr=0.0101/0.9899): Bayes 1.01%; ridge 0.67%, boosted trees 0.42% (5.1% in sample), network 0.12%. The Friedman function is nonlinear, so ridge is biased, but with 20 000 examples and 1% of signal the variance of the flexible models costs more than ridge’s bias: at low signal-to-noise the rigid model wins until the sample is large (the stock panel’s trees needed eight years of 500 stocks).
Exercise 1.8 ★★★
Prove part 2 of Proposition 1.6 and use it to explain why a model whose forecasts correlate at 0.92 with the truth can still score below zero out of sample.
Solution
Solution of Exercise 1.8.
Expand with ; the quadratic in peaks at with value , and . A forecast that correlates at 0.92 with the truth correlates with returns at about ; if its scale is more than twice (a model fitted to raw returns carrying a spurious intercept, or one whose predictions spread as if the signal were strong), although the ranking of stocks is good. The information coefficient, which ignores scale, is the better measure of the ranking; R-squared also judges the size.
1.9 Problem: The 0.57% Model
Problem 1.1
Weekend problem — a model near its ceiling
The chapter’s panel: 500 stocks, twenty characteristics, twenty years of training and ten of test, and a planted truth.
Part I — The ceiling.
- What is the predictability ceiling in the training and in the test years, and why do they differ?
- What correlation with returns does an R-squared of 0.60% correspond to?
- Why is the R-squared measured against zero rather than the sample mean?
- What share of the months does the truth itself lose to a forecast of zero, from Figure 1.2?
Part II — The models.
- Give the three models’ in- and out-of-sample R-squared.
- What share of the ceiling does each reach out of sample?
- Why does the ranking on the strong task differ from the ranking on the stocks?
- Which term of the bias–variance decomposition explains the trees’ in-sample 1.45%?
Part III — Effective samples.
- How many independent stocks is a month worth, and why?
- What do the models score when trained on raw returns, and why?
- How long a training window does boosting need to beat ridge regression?
- How many months of live results tell the trees from zero, and from ridge?
Part IV — The verdict.
- What happens to the trees when the interaction flips sign, and what did the inputs show?
- State the named result: the ceiling, the model’s share of it, and the months needed to tell it from zero and from ridge.
- Would you trade the trees or ridge regression, with ten years of history? With two?
- What should a monitoring rule watch, given part IV’s first question?
- How would crowding appear in the sandbox?
- Why is the ceiling unknowable on real data, and what replaces it?
- What does Gu, Kelly and Xiu’s best monthly R-squared of 0.40% suggest about real ceilings?
- In one sentence: why is financial machine learning different?
Solution
Solution of Problem 1.1.
- 0.72% in the training years, 0.60% in the test years: the ceiling is a sample statistic of the truth against realised returns, and the noise (market and specific) differs from decade to decade.
- .
- The historical mean of a stock’s return is so noisy that beating it is easy; zero is the stricter benchmark (Gu, Kelly and Xiu).
- 19% of the months (23 of 120); the trees lose 18%.
- Ridge 0.39% and 0.38%; trees 1.45% and 0.57%; network 0.64% and 0.28%.
- 64%, 95% and 47%.
- On the strong task noise is small and bias dominates, so flexible models win and score the same in and out of sample; on the stocks noise dominates and variance decides.
- The variance term: the trees have fitted training noise.
- 3.7: the market and industry factors make the stocks’ shocks correlate at 0.27 on average.
- , , : their intercepts equal the training window’s average return, .
- Between 48 months (trees 0.29%, ridge 0.33%) and 96 (0.47% against 0.37%).
- 4.3 months from zero; 21 months from ridge.
- Its R-squared falls to an average of 0.01% in years 6–10 while the ceiling averages 0.83%; the inputs, ranks, look the same before and after.
- Named result: ceiling 0.60%, the trees reach 95% of it (0.57%), 4.3 months to tell them from zero and 21 to tell them from ridge.
- With ten years, the trees (with a live track long enough to confirm them); with two years, ridge (0.29% against 0.18%).
- The realised R-squared or IC with a sequential test, since the errors, not the inputs, revealed the change.
- As
decay_from: a characteristic’s effect shrinking after it is widely traded. - The truth is unobserved; bound it from the best published R-squared or IC of comparable work and from the spread of models that converge on a level.
- It is a lower bound on the real ceiling; that very different models cluster between 0.33% and 0.40% suggests the ceiling is not far above.
- Because the signal is a fraction of a per cent of the variance, changes, is shared across assets and is competed away.
1.10 Interview questions
Interview question 1.1 ★ researcher, mle
Your model of monthly stock returns has an out-of-sample R-squared of 0.5%. Is that good?
Solution
Solution of Interview question 1.1.
For individual stocks, monthly, measured against zero out of sample, yes: the best published models reach about 0.4%. Check the benchmark (zero or mean), the period, whether it is the cross-section or the level, and the correlation it implies ().
What the interviewer is looking for: calibrated expectations of R-squared in finance, and the questions that make a number meaningful.
Interview question 1.2 ★ researcher, mle
State the bias–variance decomposition. Where do models of returns sit on the trade-off, and why?
Solution
Solution of Interview question 1.2.
Expected squared error at = noise + squared bias + variance. Returns are almost all noise, so variance is expensive and a modest bias is cheap: regularised, rigid models are the norm and flexible ones need a lot of data.
What the interviewer is looking for: the decomposition stated correctly and applied to a low signal-to-noise problem.
Interview question 1.3 ★★ researcher
You have ten million rows of daily data on 2 000 stocks over 20 years. How many independent observations do you have?
Solution
Solution of Interview question 1.3.
Far fewer than ten million. For anything common to all stocks (the market, a factor) there are about 5 000 days; stocks are correlated (an effective few dozen per day at best); overlapping labels divide again; and for a slow signal the independent observations are closer to the number of its half-lives in the sample.
What the interviewer is looking for: dates versus rows, cross-sectional correlation, overlap and persistence.
Interview question 1.4 ★★ researcher
Why do return-prediction studies measure R-squared against a forecast of zero instead of the historical mean?
Solution
Solution of Interview question 1.4.
The historical mean of an individual stock’s return is a very noisy estimate, so a forecast that beats it may have done so only because the benchmark is bad; zero is a fixed, stricter reference.
What the interviewer is looking for: understanding that the benchmark is a model with estimation error.
Interview question 1.5 ★★ researcher, trader
A model had an information coefficient of 0.06 in its backtest and 0.02 in its first six months live. List the possible causes and how you would tell them apart.
Solution
Solution of Interview question 1.5.
Overfitting in the backtest (search size, leakage), drift or decay of the signal (crowding, regime), implementation differences (data timing, universe, features computed differently live), and plain noise (six months is short). Check the backtest’s trial count and deflated statistics, rerun it on the live period’s data, compare features live and offline, and compute how many months the gap needs to be significant.
What the interviewer is looking for: a structured list and a test for each item, including the possibility that nothing is wrong.
Interview question 1.6 ★★★ researcher, mle
Show that a correctly signed forecast can have a negative out-of-sample R-squared, and say what you would do about it.
Solution
Solution of Interview question 1.6.
With mean-zero and forecast , , negative when although . Shrink the forecast (calibrate its scale on validation data, chapter 11’s methods, or Book 7’s isotonic calibration), and judge the ranking by the IC.
What the interviewer is looking for: the formula and the difference between ranking skill and calibration.