---
title: "Validation"
book: "Machine Learning for Markets"
subject: quant
language: en
chapter: 3
exercises: 8
source: https://one-course.com/books/quant/12/en/chapter/3-validation
---

# Chapter 3 — Validation

A research team’s model scores an [out-of-sample R-squared](https://one-course.com/books/quant/12/en/chapter/1-why-financial-machine-learning-is-different#def-ml-why-financial-machine-learning-is-different-r2) of 1.3% on the last seven years of a monthly stock panel, three times what the best published models reach. The post-mortem finds nothing wrong with the model: the earnings surprise it uses was joined to each stock by the end of the fiscal quarter, and the quarter’s report reaches the market a month later, in the very month whose return the model predicts. On the same data joined by filing date the model scores $-0.19\%$, as it should: in the world of this chapter nothing is predictable at the moment of decision, so every positive number below is a leak. Validation estimates how a model will do on data it has not seen; this chapter is about the many ways the data it has not seen gets into it anyway, and how to catch each one.

## 3.1 What validation estimates

Validation answers two different questions, and a single number cannot answer both.

**Definition 3.1 (Model selection, hyperparameter search).**

*Model selection* chooses, among candidate models or configurations, the one expected to generalise best. A *hyperparameter search* is model selection over a set of [hyperparameter](https://one-course.com/books/quant/12/en/chapter/1-why-financial-machine-learning-is-different#def-ml-why-financial-machine-learning-is-different-sets) values: a grid, random draws (Bergstra and Bengio, 2012), or a sequential search that proposes new values from the scores of the old.

The second question is *assessment*: how well will the chosen model do? The score that won the selection is a biased answer to it.

**Proposition 3.2 (The winner’s score is optimistic).**

Let $\hat S_1,\dots,\hat S_K$ be unbiased estimates of the generalisation scores $S_1,\dots,S_K$ of $K$ candidates. Then $\E[\max_k\hat S_k] \ge \max_kS_k$, with equality only if the maximum is attained by the same candidate on almost every draw with no estimation error.

**Proof.** $\max$ is convex, so by Jensen’s inequality $\E[\max_k\hat S_k] \ge \max_k\E[\hat S_k] = \max_kS_k$. ∎

The bias grows with the number of candidates and with the noise of each score, which on market data is large. Book 4 (chapter 12) turned this into the deflated Sharpe ratio; here it is the reason a selection score must never be reported as an assessment.

## 3.2 Splits that respect time

The chapter’s world is `firm.mlsynth`’s panel with a ceiling of zero: 200 stocks, 240 months, twenty characteristics that carry nothing. Each stock reports quarterly; the report of the quarter ending at month $t$ is filed during month $t+1$ and moves that month’s return by 2% per unit of standardised surprise. The pipeline is careful everywhere else: gradient-boosted trees trained on returns demeaned by month, on the first 160 months less one, scored against zero on the last 80.

Book 7 (chapter 20) built the splits that respect time: purged $k$-fold with an embargo, combinatorial purged cross-validation and walk-forward analysis. `firm.cvsplit` wraps them as scikit-learn cross-validators, so the library’s search and scoring tools use them unchanged ([Figure 3.1](#fig-ml-validation-splits)). The label overlap they guard against is severe on panels: with three-month labels sampled monthly and 4-fold cross-validation, a shuffled fold leaves 38 080 training labels that share a month with a test label; the purged fold leaves none. Deep trees on the twenty noise characteristics score 3.47% with shuffled folds and $-3.37\%$ with purged ones. The characteristics are persistent, so a flexible model can recognise a stock in a given year and read the neighbouring months’ labels off it. A classifier asked to tell the first half of the months from the second from the characteristics alone reaches an AUC of 0.71: the features locate time, which is the warning.

![Four ways to split sixteen periods. Shuffled folds scatter test periods among training periods, so overlapping labels and persistent features leak; purging removes the training labels that overlap the test block and the embargo those just after it; walk-forward trains only on the past; nested cross-validation chooses hyperparameters on inner folds of the outer training data only.](https://one-course.com/images/onecourse/chapters/quant-12/ml-validation/fig-e54b1ffcded9.svg)

***Figure 3.1.** Four ways to split sixteen periods. Shuffled folds scatter test periods among training periods, so overlapping labels and persistent features leak; purging removes the training labels that overlap the test block and the embargo those just after it; walk-forward trains only on the past; [nested cross-validation](#def-ml-validation-nested) chooses [hyperparameters](https://one-course.com/books/quant/12/en/chapter/1-why-financial-machine-learning-is-different#def-ml-why-financial-machine-learning-is-different-sets) on inner folds of the outer training data only.*

## 3.3 Choosing hyperparameters without fooling yourself

**Definition 3.3 (Nested cross-validation).**

*Nested cross-validation* runs an inner cross-validation inside each [training set](https://one-course.com/books/quant/12/en/chapter/1-why-financial-machine-learning-is-different#def-ml-why-financial-machine-learning-is-different-sets) of an outer one: the inner folds choose the [hyperparameters](https://one-course.com/books/quant/12/en/chapter/1-why-financial-machine-learning-is-different#def-ml-why-financial-machine-learning-is-different-sets), the configuration so chosen is refitted on the outer [training set](https://one-course.com/books/quant/12/en/chapter/1-why-financial-machine-learning-is-different#def-ml-why-financial-machine-learning-is-different-sets) and scored on the outer test fold. The mean outer score assesses the whole procedure, selection included.

Selection noise is easiest to see where there is nothing to select. Thirty runs of one configuration of boosted trees, identical except for their random seed, are trained on the chapter’s first 159 months. Their R-squared on months 160 to 199 ranges from $-0.60\%$ to $-0.27\%$, with a median of $-0.40\%$; the best of them looks 0.13 points better than the median. On months 200 to 239, which the choice never saw, that run scores $-0.20\%$ and the median run $-0.19\%$ ([Figure 3.2](#fig-ml-validation-seeds)): across the thirty seeds the two periods’ scores correlate at $-0.16$. The 0.13 points were the selection’s own noise, as [Proposition 3.2](#prop-ml-validation-max) says.

![Thirty seeds of the same boosted trees on a world with nothing to predict: the score on the months used to choose among them against the score on months nobody looked at. The run chosen for its selection score (circled) is ordinary afterwards. Data: ml_validation.selection.](https://one-course.com/images/onecourse/chapters/quant-12/ml-validation/fig-e2ba09aa329d.svg)

***Figure 3.2.** Thirty seeds of the same boosted trees on a world with nothing to predict: the score on the months used to choose among them against the score on months nobody looked at. The run chosen for its selection score (circled) is ordinary afterwards. Data: `ml_validation.selection`.*

[Nested cross-validation](#def-ml-validation-nested) corrects the report. With eight such seeds as the candidates, three outer purged folds and three inner ones on every other month (so labels do not overlap), the flat search reports its best mean outer score, $-0.62\%$; the nested procedure, choosing inside each outer [training set](https://one-course.com/books/quant/12/en/chapter/1-why-financial-machine-learning-is-different#def-ml-why-financial-machine-learning-is-different-sets), scores $-0.85\%$. The honest number is the worse one. The candidates were seeds, which makes the point plainly; with real [hyperparameters](https://one-course.com/books/quant/12/en/chapter/1-why-financial-machine-learning-is-different#def-ml-why-financial-machine-learning-is-different-sets) the same bias sits on top of whatever difference the [hyperparameters](https://one-course.com/books/quant/12/en/chapter/1-why-financial-machine-learning-is-different#def-ml-why-financial-machine-learning-is-different-sets) make.

**Method 3.4 (A search that can be reported).**

1. Fix the splits (purged folds or walk-forward, from the labels’ spans) and the score before the first run.
2. Search on inner folds only; record every configuration tried in the research log (Book 7, chapter 1): the trial count is part of the result.
3. Report the nested (or walk-forward) score of the procedure, never the winning inner score.
4. Keep a final holdout that no search has touched, and use it once.

## 3.4 Leakage hunting

**Definition 3.5 (Target leakage, train–test contamination).**

*Target leakage* is information in a feature, or in any step of the pipeline, that would not have been available at the decision time the prediction stands for, typically because it was derived from the target or recorded after it. *Train–test contamination* is information about the test data reaching the fit through the data themselves: overlapping labels or duplicated records on both sides of a split, or statistics computed on the whole sample.

Book 7 named look-ahead bias (chapter 3) and survivorship bias; leakage is their general form, and the literature on data mining calls it one of the commonest mistakes in the field (Kaufman, Rosset, Perlich and Stitelman). The chapter plants five leaks in its careful pipeline ([Table 3.1](#tab-ml-validation-leaks)). Two are in the features: the period-end join of the surprise, and *target encoding*, here the stock’s industry’s average return in the same month, a statistic of the target that happens to be available in the table when the model is fitted. One is in the feature selection: screening 1 000 noise candidates on their information coefficient over the whole sample before the split, and keeping ten. One is in the splits (shuffled folds on overlapping labels), and one in the choice (the best seed by test score).

| leak | leaky $R^2_{\mathrm{oos}}$ | honest version | caught by |
| --- | --- | --- | --- |
| period-end join of the surprise | 1.27% | $-0.19\%$ (filing date) | truncation test |
| target encoding, same month | 8.97% | $-0.19\%$ (none) | canary, truncation test |
| screening on the whole sample | 0.13% | $-0.16\%$ (training months) | canary, truncation test |
| shuffled folds, overlapping labels | 3.47% | $-3.37\%$ (purged folds) | fold-overlap check |
| best of 30 seeds on the test | $-0.27\%$ | $-0.20\%$ (untouched months) | holdout, nested search |

***Table 3.1.** Five leaks in a world where nothing is predictable at decision time (200 stocks, 240 months). Every positive number is fake; the last row’s leak is a gain of 0.13 points over the median seed that does not recur. Screening uses ridge regression on the ten screened features, the others boosted trees. Data: `ml_validation.leaks`.*

**Definition 3.6 (Adversarial validation, leakage audit).**

*Adversarial validation* trains a classifier to tell training rows from test rows; an AUC well above one half means the features locate the split, so shuffled folds will leak and the test period differs from the training period. A *leakage audit* is the set of checks a pipeline passes before its score is believed: at least a truncation test of every feature, a canary run of the whole pipeline, a fold-overlap check of the splits, and adversarial validation.

The truncation test recomputes each feature on the data known at a date and compares it with the stored value: on the surprise stored by filing month, the period-end join’s feature differs by up to 4.77 standard deviations, the filing-date join’s by exactly zero; the screened features differ by 5.13, because the screen chose different candidates with less data. The canary runs the whole pipeline on a target that is noise by construction, with the real target’s shape: the target-encoding pipeline still scores 4.03% on noise (standard error 0.02%) and the screening pipeline 0.17% (0.03%), which no honest pipeline can. The canary does not catch the period-end join ($-0.44\%$), because that feature leaks about the real target, not about the pipeline’s own: the audit needs both tests. The fold-overlap check counts training labels whose spans touch the test block; the holdout catches selection. No single check catches all five.

**Method 3.7 (The leakage audit).**

1. Store every input with its knowledge time (Book 7’s bitemporal store) and recompute each feature on the data known at a sample of dates; any difference is a leak.
2. Run the whole pipeline on a canary target with the real target’s structure (same overlap, same scale); a score more than two standard errors above chance is a leak in the pipeline.
3. Check that no training label overlaps a test label; run [adversarial validation](#def-ml-validation-audit) on the features.
4. Keep a holdout outside every search, and count the trials.

## 3.5 Tutorial: five leaks

**Goal.** Plant five leaks in a pipeline on a world with nothing to predict, measure the R-squared each one fakes, catch each with a detector, and compare flat and nested searches. **End state:** [Table 3.1](#tab-ml-validation-leaks), [Figure 3.2](#fig-ml-validation-seeds), the detector numbers.

1. **The period-end join**: one line decides whether the quarter ending at $t$ appears at $t$ or at $t+1$. `def surprise_known (s, join=" filing " ): """Row t: the latest surprise as the pipeline sees it at the end of month t. 'filing' uses the quarters filed by then (ended the month before or earlier); 'period' joins by the fiscal period's end, so the quarter ending at t appears at t although it is filed during t + 1.""" if join == " period " : return _ffill(s) shifted = np.full_like(s, np.nan)` **Listing 3.1.** The surprise as the pipeline sees it, by filing date or by period end. code/ml/03-validation/python/ml_validation.py
2. **[Nested cross-validation](#def-ml-validation-nested)** over any scikit-learn estimator and any splitters. `def nested_cv (make_model, grid: dict , X, y, outer, inner, score) -> dict : """outer: a splitter over rows; inner(train_idx) -> a splitter over positions 0..len(train_idx)-1. For each outer fold, the configuration with the best mean inner score is refitted on the outer training rows and scored on the outer test rows. flat_best is the best mean score over the outer folds themselves -- the optimistic number a flat search reports.""" X, y = np.asarray(X), np.asarray(y) cands = _grid(grid) outer_scores, chosen = [], [] flat = np.zeros((len (cands), outer.get_n_splits())) for f, (tr, te) in enumerate (outer.split(X)): inner_means = [] for c in cands: s = [score(y[tr][b], make_model(**c).fit(X[tr][a], y[tr][a]).predict(X[tr][b])) for a, b in inner(tr).split(X[tr])] inner_means.append(np.mean(s)) for j, c in enumerate (cands): flat[j, f] = score(y[te], make_model(**c).fit(X[tr], y[tr]).predict(X[te])) best = int (np.argmax(inner_means)) chosen.append(cands[best]) outer_scores.append(flat[best, f]) flat_means = flat.mean(axis=1 ) return {" outer_scores " : np.array(outer_scores), " chosen " : chosen, " flat_best " : float (flat_means.max()), " flat_scores " : flat_means}` **Listing 3.2.** Nested cross-validation and the flat search’s optimistic best. code/firm/cvsplit/firm_cvsplit.py
3. **The truncation test and the canary.** `def truncation_test (feature_fn, data, checkpoints) -> float : """feature_fn(data) -> array whose first axis is time. Recompute it on data truncated after each checkpoint and compare row t with the full computation's row t. Zero means the feature at t used nothing after t.""" full = np.asarray(feature_fn(data), float ) worst = 0.0 for t in checkpoints: part = np.asarray(feature_fn(data[: t + 1 ]), float ) d = np.nanmax(np.abs(part[t] - full[t])) worst = max (worst, float (d)) return worst def canary (pipeline, make_noise_target, reps: int = 5 , seed: int = 0 ): """pipeline(y) -> out-of-sample score using the given target; make_noise_target(rng) -> a target with the real target's structure (overlap, shape) but no relation to anything. Returns the mean score and its standard error.""" rng = np.random.default_rng(seed) s = np.array([pipeline(make_noise_target(rng)) for _ in range (reps)]) return float (s.mean()), float (s.std(ddof=1 ) / np.sqrt(reps))` **Listing 3.3.** Two detectors of the leakage audit. code/firm/cvsplit/firm_cvsplit.py
4. **Run** `leaks()` , `overlap_detectors()` , `detectors()` , `nested()` and `fig_validation.py` .

**What to change next.** Replace the same-month target encoding by one computed on the previous month’s returns and check that the truncation test passes; raise the announcement effect to 4% and see the period-end join’s fake R-squared grow with its square.

## 3.6 Build: splitters and leak detectors

**Purpose.** Every model of the firm is validated with splits that respect time, selected without touching its assessment, and passed through a [leakage audit](#def-ml-validation-audit).

**Interface.** `PurgedKFold(t0, t1, n_splits, embargo)`, `CPCVSplit(t0, t1, n_folds, n_test, embargo)`, `WalkForwardSplit(n, train, test, step, anchored)` (scikit-learn cross-validators on Book 7’s `firm.overfit`); `nested_cv(make_model, grid, X, y, outer, inner, score)`; `fold_overlap`, `truncation_test`, `canary`, `adversarial_validation`.

**Rules.** Rows in time order; every label carries its span; a score used to choose is never reported as an assessment.

**Acceptance tests.** `code/firm/cvsplit/tests/`: the splitters work inside `cross_val_score` and purged folds leave no overlap; nested search picks the right ridge penalty and the flat best is never below the nested score; the truncation test flags a centred moving average and passes a causal one; the canary flags a pipeline that scores a feature made of its target; [adversarial validation](#def-ml-validation-audit) is near one half on one distribution and high on two.

**Stretch.** A dependency tracer that records which rows each feature value read; a canary that permutes the target in blocks as long as the labels.

Sources and further reading

- S. Varma and R. Simon, “Bias in error estimation when using cross-validation for model selection”, *BMC Bioinformatics* 7, 2006.
- G. C. Cawley and N. L. C. Talbot, “On over-fitting in model selection and subsequent selection bias in performance evaluation”, *Journal of Machine Learning Research* 11, 2010.
- S. Kaufman, S. Rosset, C. Perlich and O. Stitelman, “Leakage in data mining: formulation, detection, and avoidance”, *ACM Transactions on Knowledge Discovery from Data* 6(4), 2012.
- J. Bergstra and Y. Bengio, “Random search for hyper-parameter optimization”, *Journal of Machine Learning Research* 13, 2012.

## 3.7 Exercises

**Exercise 3.1 ★.**

Five [hyperparameter](https://one-course.com/books/quant/12/en/chapter/1-why-financial-machine-learning-is-different#def-ml-why-financial-machine-learning-is-different-sets) settings have true [out-of-sample R-squared](https://one-course.com/books/quant/12/en/chapter/1-why-financial-machine-learning-is-different#def-ml-why-financial-machine-learning-is-different-r2) of 0.30% each, and each validation estimate has a standard error of 0.10%, independently. Roughly what does the best validation score average? (The expected maximum of five standard normals is 1.16.)

**Solution of Exercise 3.1.**

$0.30\% + 1.16\times0.10\% = 0.42\%$: the winner’s score overstates every candidate’s true 0.30% by 0.12 points.

**Exercise 3.2 ★.**

Labels span three months and are sampled monthly; the test fold is months 100 to 139. With purging and a one-month embargo, which training months go?

**Solution of Exercise 3.2.**

The test labels span months 100 to 141. Purging drops the training labels starting at 98 and 99 (they end inside the span) and at 140 and 141 (they start inside it); the embargo also drops 142.

**Exercise 3.3 ★.**

A canary run of a pipeline gives R-squared scores of 0.21%, 0.15% and 0.18% on three noise targets. Is the pipeline leaking? Compute the mean and its standard error.

**Solution of Exercise 3.3.**

Mean 0.18%, standard deviation 0.03%, standard error $0.03/\sqrt3 = 0.017\%$: ten standard errors above zero on targets that cannot be predicted. The pipeline leaks.

**Exercise 3.4 ★★.**

Why does the canary miss the period-end join, while the truncation test catches it? Which of the two catches target encoding, and why both?

**Solution of Exercise 3.4.**

The period-end join leaks information about the real returns (the report moves them); on a noise target it carries nothing, so the canary sees no gain ($-0.44\%$). The truncation test asks when each value became known and sees the surprise one month early. Target encoding is built from the target: on a noise target it still leaks (the canary’s 4.03%), and recomputed with data truncated at a date it differs (the truncation test).

**Exercise 3.5 ★★.**

[Adversarial validation](#def-ml-validation-audit) of a model’s training and test months gives an AUC of 0.95. List two different conclusions this supports and one it does not.

**Solution of Exercise 3.5.**

The features locate the period, so shuffled folds would leak (neighbours are recognisable); and the test period’s feature distribution differs from training (drift, a new regime, a changed data source). It does not show that the model is wrong or that the features predict returns: a feature can shift and stay useful.

**Exercise 3.6 ★★.**

*Find the flaw.* “We standardise every feature with the mean and standard deviation of the whole dataset, then run purged cross-validation, so there is no leakage.”

**Solution of Exercise 3.6.**

Whole-sample statistics use the test data: the scale and centre of each feature in the training rows depend on the future. For stationary, well-behaved features the leak is small; for trending or drifting ones (volumes, prices, counts) it is not, and it is a look-ahead however small. Fit the scaler inside each training fold, or use cross-sectional ranks computed date by date.

**Exercise 3.7 ★★★.**

*Coding.* In `ml_validation`, set the announcement effect `ANN` to 4% and rerun the period-end and filing-date pipelines. By how much does the fake R-squared grow, and why about fourfold?

**Solution of Exercise 3.7.**

`announcement_effect(0.04)`: the period-end join scores 5.47% (1.27% at 2%), the filing-date join $-0.14\%$. The explained variance from a known surprise is proportional to the square of the effect, so doubling it multiplies the gain by about four; the model’s own error adds the same small negative term in both cases.

**Exercise 3.8 ★★★.**

Prove that [nested cross-validation](#def-ml-validation-nested)’s outer score is an unbiased estimate of the procedure’s generalisation score when the outer folds are independent of the procedure’s choices, and explain why the flat best is not.

**Solution of Exercise 3.8.**

In each outer fold the procedure (search, choice, refit) uses only the outer training rows, so its fitted model is independent of the outer test rows; its score there is an unbiased estimate of the generalisation score of what the procedure produced from that much data, and so is the average over folds. The flat best takes a maximum over candidates of scores measured on the same test rows, and by [Proposition 3.2](#prop-ml-validation-max) its expectation exceeds every candidate’s true score.

## 3.8 Problem: Five Leaks

**Problem 3.1.**

Weekend problem — auditing a pipeline that cannot win

The chapter’s world: 200 stocks, 240 months, characteristics that carry nothing, and earnings reports filed a month after each quarter ends.

**Part I — The honest pipeline.**

1. What does the careful pipeline score, and why is a negative number the right answer here?
2. Why is one month removed between the training and the test months?
3. What does the filing-date join put in month $t$ ’s row, and what does the period-end join put there?
4. What announcement effect makes the period-end join’s fake R-squared 1.27%?

**Part II — The leaks.**

5. Give the fake R-squared of the target encoding and explain its size.
6. What do screening on the whole sample and on the training months give?
7. How many training labels share a month with a shuffled test fold, and how many with a purged one?
8. What do deep trees score with shuffled and with purged folds, and what lets them exploit the overlap?

**Part III — Selection.**

9. What range do the thirty seeds’ selection scores cover, and how far is the best from the median?
10. What does the chosen seed score on untouched months?
11. What do the flat and the nested searches report?
12. Why does [Proposition 3.2](#prop-ml-validation-max) predict the gap?

**Part IV — The audit.**

13. Which detector catches each leak, with its number?
14. Why can no single detector catch all five?
15. What does an adversarial AUC of 0.71 on the characteristics mean for this pipeline?
16. State the *named result* : the R-squared each leak fakes on a target that cannot be predicted, and the detector that catches it.
17. Write the audit as a checklist a reviewer signs.
18. Which of the five leaks would survive a paper-trading period, and for how long?
19. How should the research log record a leak once found?
20. In one sentence: what is validation for?

**Solution of Problem 3.1.**

1. $-0.19\%$ : nothing is predictable at decision time, so the best possible score is zero and a fitted model can only lose to it.
2. The last training row’s return is that of month 160: a gap keeps the training labels from ending inside the test.
3. Filing date: the latest surprise filed by the end of month $t$ (quarter ended at $t-1$ or before). Period end: the surprise of the quarter ending at $t$ , filed during $t+1$ .
4. 2% of return per unit of standardised surprise.
5. 8.97%: the same-month industry return explains the industry factor’s share of the variance, which is large.
6. 0.13% on the whole sample, $-0.16\%$ on the training months.
7. 38 080 against none.
8. 3.47% and $-3.37\%$ ; persistent characteristics let deep trees recognise a stock in a given period and read its neighbours’ overlapping labels.
9. From $-0.60\%$ to $-0.27\%$ ; the best is 0.13 points above the median $-0.40\%$ .
10. $-0.20\%$ , against a median of $-0.19\%$ ; the two periods’ scores correlate at $-0.16$ .
11. Flat best $-0.62\%$ , nested $-0.85\%$ .
12. The maximum of noisy estimates overstates in expectation; the nested score is measured on data the choice never saw.
13. Period-end join: truncation test (4.77 against 0). Target encoding: canary (4.03%) and truncation test. Screening: canary (0.17%) and truncation test (5.13). Overlap: fold-overlap check (38 080 against 0). Selection: untouched holdout or nested search.
14. Each detector looks at one place (features, pipeline, splits, choice); a leak lives in one place.
15. The characteristics identify the period, so any shuffled split leaks and the test months differ from the training months.
16. *Named result* : on a target that cannot be predicted, the period-end join fakes 1.27%, target encoding 8.97%, whole-sample screening 0.13% ( $-0.16\%$ honest), shuffled folds on overlapping labels 3.47% ( $-3.37\%$ purged), and selecting the best seed on the test 0.13 points; truncation, canary, fold-overlap and holdout catch them in that order.
17. Features recomputed as of sample dates match; canary score within two standard errors of chance; no training label overlaps a test label; adversarial AUC recorded; holdout untouched; trial count recorded.
18. None survives: each vanishes as soon as the data arrive in real time. Selection shows up as a smaller live score; the others make live trading lose at once, which is why they are found late if paper trading is short.
19. As a finding with its measured effect, the fix, and every result it invalidates.
20. To estimate honestly how a procedure will do on data it has not seen, which requires that it really has not seen them.

## 3.9 Interview questions

**Interview question 3.1 ★ researcher, mle.**

What is data leakage? Give three examples from financial data.

**Solution of Interview question 3.1.**

Information in training or features that would not be available at the decision time. Examples: fundamentals joined by period end instead of filing date; revised macro data used as first released; universe of today’s index members (survivorship); features scaled with future statistics; overlapping labels across shuffled folds.

*What the interviewer is looking for: a precise definition tied to decision time, and concrete financial cases.*

**Interview question 3.2 ★★ mle.**

You tuned twenty [hyperparameter](https://one-course.com/books/quant/12/en/chapter/1-why-financial-machine-learning-is-different#def-ml-why-financial-machine-learning-is-different-sets) settings by 5-fold cross-validation and report the best fold-average score. What is wrong, and what do you report instead?

**Solution of Interview question 3.2.**

The best of twenty fold-averages is biased upwards (a maximum of noisy estimates). Report a [nested cross-validation](#def-ml-validation-nested) or a walk-forward score of the whole procedure, and the number of settings tried.

*What the interviewer is looking for: selection bias and the nested or holdout remedy.*

**Interview question 3.3 ★★ researcher.**

How would you test whether a feature pipeline has look-ahead, without reading its code?

**Solution of Interview question 3.3.**

Recompute each feature on data truncated at a set of dates and compare with the stored feature (any difference is look-ahead); run the pipeline on noise targets with the same structure (canary); check the knowledge times of inputs against decision times.

*What the interviewer is looking for: black-box tests that do not rely on reading code.*

**Interview question 3.4 ★★ researcher, mle.**

What is [adversarial validation](#def-ml-validation-audit) and what does it tell you about a train–test split?

**Solution of Interview question 3.4.**

A classifier trained to separate training from test rows; its AUC measures how different they are. High AUC warns that the split’s sides differ (drift) and that the features locate time, so shuffled splits leak; the features it relies on are the ones to inspect.

*What the interviewer is looking for: what it measures and the two uses.*

**Interview question 3.5 ★★ researcher.**

Why is target encoding dangerous with time series, and how do you do it safely?

**Solution of Interview question 3.5.**

Encodings computed with same-period or future targets leak the target itself. Compute them from past periods only (expanding or lagged means), inside each training fold, with smoothing toward the global mean.

*What the interviewer is looking for: lagging the statistic, and computing it inside the fold.*

**Interview question 3.6 ★★★ researcher.**

A model’s cross-validated score is excellent and its paper-trading score is at zero. Walk through how you would find out whether the cause is leakage, [overfitting](https://one-course.com/books/quant/12/en/chapter/1-why-financial-machine-learning-is-different#def-ml-why-financial-machine-learning-is-different-generalisation) or a change in the market.

**Solution of Interview question 3.6.**

Run the [leakage audit](#def-ml-validation-audit) (truncation, canary, fold overlap): a leak shows there. Then check the search: trial count, nested score, deflated statistics; a nested score near zero means [overfitting](https://one-course.com/books/quant/12/en/chapter/1-why-financial-machine-learning-is-different#def-ml-why-financial-machine-learning-is-different-generalisation). If both pass, compare feature distributions and model errors live against backtest ([adversarial validation](#def-ml-validation-audit), drift tests, chapter 12) and compute how long the paper-trading period must be for the gap to be significant.

*What the interviewer is looking for: an ordered diagnosis with a test for each cause.*
