---
title: "Machine Learning"
book: "The Interview Book"
subject: quant
language: en
chapter: 19
exercises: 0
source: https://one-course.com/books/quant/18/en/chapter/19-machine-learning
---

# Chapter 19 — Machine Learning

A candidate’s model scored 71% accuracy in cross-validation on the daily direction of a futures contract and 50% in its first month live. Asked why, he says “overfitting”. The interviewer asks him to name the three lines of his pipeline most likely to have leaked the future, and the interview starts there. Machine-learning questions in finance interviews are rarely about architectures. They are about the ways a model can look good and be worthless, which in markets are more numerous than in most fields: the signal is weak, the data are few and dependent, and the future does not resemble the past. The vocabulary is One Quant Book 12’s (chapters 1 to 11) and the validation discipline One Quant Book 7’s (chapters 3 and 20).

## 19.1 Bias, variance and regularisation

The bias–variance decomposition (One Quant Book 12, chapter 1) answers most “what happens if” questions. More data lowers variance and leaves bias; more features or a more flexible model lower bias and raise variance; more regularisation raises bias and lowers variance. In finance, where the signal-to-noise ratio of returns is tiny, variance dominates: the best models are usually simple, strongly regularised, and evaluated honestly.

Ridge and lasso regularise differently (One Quant Book 4, chapter 16). The ridge penalty $\lambda\|\beta\|_2^2$ shrinks all coefficients and shares weight among correlated features; the lasso penalty $\lambda\|\beta\|_1$ has corners on the axes, so its solution often sits on one, setting some coefficients exactly to zero. With nearly duplicate features the lasso picks one of them, and which one depends on noise.

**Example 19.1 (Two nearly identical features).**

Two features correlated at 0.99, a target equal to their sum plus noise, 100 observations. Over 200 samples, a lasso with $\alpha = 0.5$ sets one of the two coefficients to zero in 63% of them, choosing each about equally often; ridge gives both about 0.95 every time. The lasso’s selection is not a finding about which feature matters.

## 19.2 Trees, ensembles and networks

Gradient boosting and random forests (One Quant Book 12, chapter 5) fit interactions and non-linearities without specification, which is why they overfit noise readily: with enough trees a boosted model fits a pure-noise training set.

**Example 19.2 (Boosting on noise).**

Five noise features and a noise target, 1 000 training and 1 000 test observations, 300 trees of depth 3 at a learning rate of 0.1: training $R^2$ of 0.65 and test $R^2$ of $-0.18$. Early stopping on a validation set that respects time would have stopped near zero trees, which is the right answer.

Two further traps recur. Impurity-based feature importances favour continuous and high-cardinality features even when they are noise; use permutation importance on held-out data. And deep networks on daily returns disappoint for reasons of data, not architecture: a few thousand dependent observations of a signal with an out-of-sample $R^2$ well below 1% cannot train millions of parameters (One Quant Book 12, chapter 7).

## 19.3 Validation with time

**Method 19.3 (Finding the leak).**

1. *Timestamps* : is every feature computed from data available at the decision time, with its publication lag? (point-in-time data, One Quant Book 7, chapter 3)
2. *Labels* : do labels overlap in time (a 20-day forward return every day)? Then neighbouring observations share most of their label.
3. *Splits* : does cross-validation shuffle observations, so that neighbours of a test point sit in the training set? Purge training observations whose labels overlap the test period and add an embargo after it (One Quant Book 7, chapter 20).
4. *Preprocessing* : were scalers, feature selection or hyperparameters fitted on the whole sample?
5. *Selection* : how many models were tried before this one?

![Twenty simulated random walks with no predictability, trailing-return features and overlapping 20-day forward labels, fitted with five-nearest-neighbour regression. Shuffled cross-validation reports a positive R2 on every one (mean 0.26); a walk-forward evaluation with purging reports a negative one (mean -0.61). Data: fig_iv_leak.py.](https://one-course.com/images/onecourse/chapters/quant-18/iv-machine-learning/fig-6866719c6c58.svg)

***Figure 19.1.** Twenty simulated random walks with no predictability, trailing-return features and overlapping 20-day forward labels, fitted with five-nearest-neighbour regression. Shuffled cross-validation reports a positive $R^2$ on every one (mean 0.26); a walk-forward evaluation with purging reports a negative one (mean $-0.61$). Data: `fig_iv_leak.py`.*

The mechanism in [Figure 19.1](#fig-iv-machine-learning-leak) is leakage through overlap: slowly varying features make yesterday and tomorrow the nearest neighbours of today, and overlapping labels make their targets almost today’s target, so a shuffled split lets the model look up the answer. Nothing is predictable; the score is.

## 19.4 Design questions

**Definition 19.4 (Machine-learning design question).**

A *machine-learning design question* asks the candidate to turn a vague business goal into a learning problem: the decision the model serves, the target and its horizon, the features available at decision time, the validation scheme, the metric that reflects the decision’s payoff, and the monitoring after deployment.

**Method 19.5 (Answering a design question).**

1. Name the decision and its cost structure first; the metric follows from it, not from convention.
2. Define the target precisely (event, horizon, how it is observed, when it is known).
3. List features by source and latency; exclude anything not known at the decision time.
4. Start with a baseline (a rule, a logistic regression), then a stronger model if it earns its complexity.
5. Validate forward in time with purging; report the metric with its uncertainty.
6. Say how the model fails in production (drift, feedback from its own actions) and what is monitored (One Quant Book 12, chapter 27).

## 19.5 Worked answers

**Example 19.6 (Reading a confusion table aloud).**

*“Of 10 000 orders, 3% come from counterparties later judged toxic. The model flags 500 orders, of which 240 are toxic. How good is it?”* There are 300 toxic orders; the model catches 240, a recall of 80%, and 240 of its 500 flags are right, a precision of 48%. Against the base rate of 3%, a flagged order is 16 times more likely to be toxic than a random one: the lift. Whether 48% is good depends on the action: widening quotes for the flagged orders costs a little on the 260 false alarms and saves a lot on the 240 toxic ones, and the threshold should be moved until the marginal flag’s expected saving equals its expected cost ([Interview question 19.13](#iq-iv-machine-learning-13) does this for a trade).

**Example 19.7 (How many independent labels?).**

*“We have ten years of daily data and predict twenty-day forward returns. How many observations do we have?”* About 2 520 rows but far fewer independent labels: consecutive twenty-day returns share nineteen days, so roughly $2\,520/20 = 126$ non-overlapping ones. A $t$-statistic computed as if the 2 520 were independent is inflated by about $\sqrt{20} \approx 4.5$, and a validation fold whose labels overlap the training fold’s leaks. The fixes are standard: purge training rows whose label window overlaps the validation period, embargo a gap after it, and compute standard errors with overlap-robust (Newey–West or block bootstrap) methods ([Method 19.3](#met-iv-machine-learning-leak)).

## 19.6 Question bank

**Interview question 19.1 ★ mle, researcher • systematic fund.**

What happens to bias and to variance when you add data, add features, or increase regularisation?

**Solution of Interview question 19.1.**

More data: variance falls, bias unchanged. More features or flexibility: bias falls, variance rises. More regularisation: bias rises, variance falls. With noisy financial targets the variance term usually dominates, so the gains come from data and regularisation, not flexibility.

*What the interviewer is looking for: the three directions, and which term dominates in finance.*

**Interview question 19.2 ★ mle, researcher • any.**

Why does the lasso set coefficients exactly to zero while ridge does not?

**Solution of Interview question 19.2.**

The constraint region of the lasso ($\|\beta\|_1 \le t$) has corners on the axes, and the elliptical contours of the squared loss usually first touch it at a corner, where some coordinates are zero. The ridge region is a ball with no corners, so the contact point generically has all coordinates non-zero. Equivalently, the lasso’s subgradient at zero is an interval, so zero is optimal for small effects.

*What the interviewer is looking for: the geometric or subgradient explanation.*

**Interview question 19.3 ★ researcher, trader • market maker.**

A model predicts the next day’s direction with 55% accuracy. Is that good?

**Solution of Interview question 19.3.**

It depends on the payoffs and on the evidence. If gains and losses are symmetric and unrelated to the prediction, 55% is a large edge for daily data; if the model is right on small moves and wrong on large ones, it can lose money. Ask for the expected P&L per prediction, a proper score of the probabilities (log loss or Brier), the sample size (55% over 250 days has a standard error of about 3 points) and how many models were tried.

*What the interviewer is looking for: payoff asymmetry, uncertainty of the estimate and the right metric.*

**Interview question 19.4 ★ mle • any.**

Why standardise features before fitting ridge or lasso, and with which statistics?

**Solution of Interview question 19.4.**

The penalty treats all coefficients alike, so features on larger scales are penalised less per unit of effect; standardising puts them on equal footing. Use the mean and standard deviation of the training fold only (fitted inside each validation split), otherwise the scaling leaks information from the test period.

*What the interviewer is looking for: scale dependence of the penalty and leak-free preprocessing.*

**Interview question 19.5 ★★ mle, researcher • systematic fund.**

Two features are correlated at 0.99. A lasso keeps one and drops the other, and on a new sample it keeps the other one. What is happening, and what do you report?

**Solution of Interview question 19.5.**

With nearly collinear features the lasso’s solution is unstable: small changes in the sample flip which feature carries the effect ([Example 19.1](#ex-iv-machine-learning-lasso): one coefficient zeroed in 63% of samples, either feature equally often). Report the pair as one effect, use ridge or an elastic net to share the weight, or combine the features; never interpret the lasso’s choice between them.

*What the interviewer is looking for: instability of sparse selection under collinearity and a robust alternative.*

**Interview question 19.6 ★★ researcher, mle • systematic fund.**

A colleague predicts 20-day forward returns every day from trailing 20-, 60- and 120-day returns, with shuffled 5-fold cross-validation, and reports an $R^2$ of 0.26. What do you suspect, and how do you test it?

**Solution of Interview question 19.6.**

Leakage through overlapping labels and shuffled folds: consecutive days have almost the same features and almost the same 20-day target, so a test day’s neighbours in the training set give away its label. Test it on data with no signal: on the chapter’s noise panels the same pipeline scores 0.26 with shuffled folds and $-0.61$ with purged walk-forward evaluation ([Figure 19.1](#fig-iv-machine-learning-leak)). Re-evaluate with walk-forward splits, purge training labels that overlap the test period, and add an embargo.

*What the interviewer is looking for: the overlap mechanism and a noise-panel test of the pipeline.*

**Interview question 19.7 ★★ mle • proprietary firm.**

A gradient-boosted model has a training $R^2$ of 0.65 and a test $R^2$ of $-0.18$. What happened, and what would you change?

**Solution of Interview question 19.7.**

It fitted noise: the gap between training and test $R^2$ is the variance of an over-flexible model (the chapter’s demonstration uses pure noise and gets these numbers). Use fewer, shallower trees with a lower learning rate, subsampling, and early stopping on a time-ordered validation set; compare with a linear baseline; and check that the test set is forward in time.

*What the interviewer is looking for: recognising overfitting and the standard boosting controls.*

**Interview question 19.8 ★★ researcher, mle • systematic fund.**

A random forest’s feature importances rank a noisy continuous feature first. Why can that happen, and what importance would you use instead?

**Solution of Interview question 19.8.**

Impurity importance counts how much each split reduces training impurity, and a continuous noisy feature offers many split points, so the forest finds spurious splits on it; the importance is computed in sample. Use permutation importance on held-out, time-ordered data (the drop in the validation metric when the feature is shuffled), with repeats to measure its noise.

*What the interviewer is looking for: the bias of impurity importance and a held-out alternative.*

**Interview question 19.9 ★★ mle, researcher • multi-manager fund.**

Why does a deep network trained on ten years of daily returns usually do worse than a regularised linear model?

**Solution of Interview question 19.9.**

Ten years are about 2 500 dependent observations of a target whose predictable part is a fraction of a percent of its variance; a network with many parameters fits the noise, and its training is unstable across seeds. Non-stationarity makes old data a poor guide. A regularised linear model has little variance and captures most of what is there. Networks help where data are plentiful and structured (order books, text), not on daily returns.

*What the interviewer is looking for: sample size against signal strength, and non-stationarity.*

**Interview question 19.10 ★★★ mle, developer • market maker.**

Design a model that predicts the probability that a passive limit order will be filled within one second.

**Solution of Interview question 19.10.**

Decision: where and whether to post, trading off fill probability against adverse selection. Target: filled within one second of placement (with partial fills defined), observed from our own order records. Features known at placement: queue position and size ahead, spread, depth on both sides, recent trade rate and imbalance, time of day, volatility, the order’s price relative to the touch. Baseline: a logistic regression; then boosting. Validate forward in time on our own orders only, beware of selection (we only observe outcomes for orders we sent), and calibrate the probabilities (proper scoring rule). Monitor calibration drift and the feedback loop (the model changes where we post, which changes the data).

*What the interviewer is looking for: the full chain from decision to monitoring, including selection bias and feedback.*

**Interview question 19.11 ★★★ mle, trader • market maker.**

Design a model that flags toxic order flow for a market maker, so that quotes can be widened for some counterparties.

**Solution of Interview question 19.11.**

Define toxicity as the mark-out: the move of the mid price against us after the counterparty’s trade, over horizons such as one second and one minute. Target: the mark-out of each fill (regression) or its sign beyond a threshold. Features by counterparty and order: past mark-outs of the counterparty, order size, timing relative to news or other venues’ moves, the counterparty’s order-to- trade ratio. Validate by counterparty and forward in time; the metric is the P&L improvement from widening on the model’s flags, not accuracy. Watch fairness and regulatory constraints on discriminating between clients, and the feedback when toxic clients adapt.

*What the interviewer is looking for: mark-out as the target, counterparty-level validation and a P&L metric.*

**Interview question 19.12 ★★★ researcher, mle • systematic fund.**

A model scored 71% accuracy in cross-validation on daily direction and 50% in its first month live. List the most likely causes in order, and the check for each.

**Solution of Interview question 19.12.**

In order: (1) leakage through overlapping labels or shuffled folds: re-run with purged walk-forward evaluation; (2) look-ahead in features (a close used before it was known, restated data): audit each feature’s timestamp; (3) selection over many models or parameters: count the trials and deflate; (4) regime change: compare feature distributions live and in sample; (5) execution: the live prediction arrives later than assumed. A 71% accuracy on daily direction is itself implausible, which is the first clue.

*What the interviewer is looking for: a ranked list of causes, each with a concrete check.*

**Interview question 19.13 ★★★ researcher, trader • proprietary firm.**

A classifier outputs the probability that a trade idea is right. A right trade gains 1 and a wrong one loses 2. Above what probability do you trade? Why is accuracy the wrong metric, and which would you use?

**Solution of Interview question 19.13.**

Trade when $p \times 1 > (1 - p) \times 2$, that is $p > \tfrac23$. Accuracy weights all errors alike and uses a threshold of one half, so a model with high accuracy can lose money; evaluate the expected P&L of the decision rule, and score the probabilities with a proper rule (log loss) so that they can be thresholded correctly.

*What the interviewer is looking for: the payoff-based threshold and the case for proper scoring.*

Sources and further reading

- One Quant Book 12, chapters 1–11 and 27; One Quant Book 7, chapters 3 and 20; One Quant Book 4, chapter 16.
- M. López de Prado, *Advances in Financial Machine Learning* , Wiley, 2018, on purging and embargoes.
