---
title: "Feature Engineering, Selection and Importance"
book: "Machine Learning for Markets"
subject: quant
language: en
chapter: 6
exercises: 8
source: https://one-course.com/books/quant/12/en/chapter/6-feature-engineering-selection-and-importance
---

# Chapter 6 — Feature Engineering, Selection and Importance

A model uses two versions of the same signal, computed over slightly different windows, whose correlation is 0.95. Its [permutation importance](#def-ml-feature-engineering-selection-and-importance-mdi) ranks them fourth and fifth of fifteen. Refitted without the first, the model loses 0.09 points of R-squared; without the second, it gains 0.02. Refitted without both, it loses 1.05 of its 2.67 points. Each copy looks useless because the other stands in for it, and a researcher who pruned “useless” features one at a time would have thrown away two-fifths of the model. This chapter builds features that behave, selects among them, and measures importance with methods whose failures are known in advance.

## 6.1 Transformations that respect the data

**Definition 6.1 (Feature engineering).**

*Feature engineering* is the construction of model inputs from raw data: choosing what to measure (a return, a volume, a book imbalance), over which window, and how to transform it (scale, rank, clip, difference) so that its relation to the target is stable and its values are known at decision time.

Market features are heavy-tailed and non-stationary. Book 7 (chapter 7) built the return, volatility and volume features; Book 4 gave winsorisation (chapter 15) and fractional differencing (chapter 17). The transformation matters more than the model for a feature with fat tails. A feature drawn from a Student $t$ with 1.5 degrees of freedom, acting on the target through $0.2\tanh(z)$ (its effect saturates), explains 2.17% of the variance at best. A linear regression on the raw feature reaches 0.13% out of sample: a few extreme values set the slope. Winsorised at the 1st and 99th percentiles it reaches 1.27%, ranked and mapped to normal scores 1.96%. Ranking by date (the cross-sectional rank transform of Book 7, chapter 6) is the default of this book for the same reason.

## 6.2 Feature selection and stability selection

**Definition 6.2 (Feature selection, stability selection).**

*Feature selection* chooses a subset of candidate features to use. *Stability selection* runs a selection procedure (the lasso’s nonzero coefficients, a model’s top features) on many random half-samples and keeps the features selected in a share of them above a threshold $\pi_{\mathrm{thr}}\in(\frac12, 1]$ (Meinshausen and Bühlmann, 2010).

**Theorem 6.3 (A bound on false selections).**

If each half-sample selects on average $q$ of $p$ features, the selections of the noise features are exchangeable, and the procedure is no worse than random guessing, the expected number of noise features selected by [stability selection](#def-ml-feature-engineering-selection-and-importance-selection) at threshold $\pi_{\mathrm{thr}}$ is at most $q^2/((2\pi_{\mathrm{thr}} - 1)p)$.

**Proof.** *Admitted here.* ∎

The proof, by a Markov-type inequality on the selection frequencies, is in Meinshausen and Bühlmann (2010).

The chapter’s task plants the roles: out of fifteen standard normal features, $x_1$ acts linearly, $x_2$ through its square, $x_3$ and $x_4$ only through their product, $x_{1b}$ is $x_1$’s near-copy (correlation 0.95, no effect of its own), and ten are noise; the signal is 4% of the target’s variance. For [stability selection](#def-ml-feature-engineering-selection-and-importance-selection) 85 more noise columns are added (100 candidates) and 30 half-samples of 5 000 rows are drawn. The lasso keeps $x_1$ alone at a threshold of 0.6 (it selects 6.7 features per half-sample; the bound allows 2.3 false selections, and none occurs); boosting’s six largest-gain features keep $x_1$ and $x_2$ (bound 1.8, none false). Neither finds $x_3$ and $x_4$: an interaction with no main effects is invisible to a linear selector, and to a greedy tree whose first split on either feature gains nothing. [Stability selection](#def-ml-feature-engineering-selection-and-importance-selection) controls what is selected, not what is missed.

## 6.3 Importance: impurity, permutation, Shapley

**Definition 6.4 (Mean decrease in impurity, permutation importance).**

The *mean decrease in impurity* (MDI, or gain importance) of a feature is the total loss reduction of the splits on it across a tree ensemble, computed on the training data. The *permutation importance* is the drop in a held-out score when the feature’s column is randomly permuted, breaking its link with the target and the other features.

**Definition 6.5 (Shapley value, SHAP value).**

For a game $v$ on the set $P$ of players, the *Shapley value* of player $j$ is

$$
\phi_j = \sum_{S\subseteq P\setminus\{j\}}\frac{|S|!\,(|P| - |S| - 1)!}{|P|!}\bigl(v(S\cup\{j\}) - v(S)\bigr),
$$

the only allocation that is efficient ($\sum_j\phi_j = v(P) - v(\emptyset)$), symmetric, zero for a player that adds nothing, and additive across games. A *SHAP value* is the Shapley value of feature $j$ for one prediction, with $v(S)$ the model’s expected output when the features in $S$ are fixed at their values (Lundberg and Lee, 2017); for tree ensembles it is computed exactly and fast (TreeSHAP).

The four measures answer different questions. MDI says where the trees spent their splits on the training data; [permutation importance](#def-ml-feature-engineering-selection-and-importance-mdi) says how much the fitted model needs a column on new data; drop-column importance (refit without the feature) says how much the *data* need it; the mean absolute [SHAP value](#def-ml-feature-engineering-selection-and-importance-shapley) says how much each feature moves the model’s individual predictions. On the planted task ([Figure 6.1](#fig-ml-importance-bars)), boosted trees fitted on 20 000 rows score 2.67% on 20 000 others (the truth scores 4.01%). Permutation and drop-column importance put $x_2$, $x_3$ and $x_4$ first and give the noise almost nothing. MDI ranks the same four features first but hands every noise column a share of the splits (the largest 5.4% of the total, against 9.5% for $x_1$): gain is earned on training data, noise included. SHAP credits $x_{1b}$ more than $x_3$ or $x_4$: it explains what the model does, and the model uses the copy.

![Four importances of the same boosted trees on the planted task, each as a share of its positive total: the four signal features, the copy x_1b, and the largest of the ten noise features. Permutation, drop-column and SHAP are computed on the 20 000 held-out rows, gain on the training rows. Data: ml_importance.importances.](https://one-course.com/images/onecourse/chapters/quant-12/ml-feature-engineering-selection-and-importance/fig-8f0ae9629ec4.svg)

***Figure 6.1.** Four importances of the same boosted trees on the planted task, each as a share of its positive total: the four signal features, the copy $x_{1b}$, and the largest of the ten noise features. Permutation, drop-column and SHAP are computed on the 20 000 held-out rows, gain on the training rows. Data: `ml_importance.importances`.*

## 6.4 Substitution and clustered importance

**Definition 6.6 (Substitution effect, clustered feature importance).**

The *substitution effect* is the understatement of a feature’s importance when a correlated feature can stand in for it: permuting or dropping one leaves the other. *Clustered feature importance* measures importance for groups of correlated features, permuted or dropped together, the groups found by hierarchical clustering on $1 - |\text{correlation}|$ (López de Prado, 2020).

$x_1$ and its copy are the case. Permuted singly they cost 0.76 and 0.41 points of R-squared; permuted together, 2.72 points, more than the model’s whole score, because the model has learned to use both and is now fed noise by both. Dropped singly and refitted they cost 0.09 and $-0.02$ points; dropped together, 1.05. Clustering the fifteen features at a distance of 0.5 finds exactly one group of two, $\{x_1, x_{1b}\}$, and every other feature alone: clustered importance restores $x_1$’s signal to the pair and names the pair, which is the honest answer when the data cannot say which copy matters. In MDI terms the pair also splits its credit (9.5% and 8.1%), so a pruning rule based on gain would have hesitated over both.

## 6.5 When importance lies

**Method 6.7 (Reading importances).**

1. Compute importances on held-out data (permutation, drop-column); read training-data MDI as a description of the fit, never as evidence.
2. Cluster the features by correlation first and report the clusters’ importances; prune whole clusters.
3. Compare with a permutation of the target (the canary of chapter 3): a feature is important only if it beats the importance noise features receive.
4. Remember what each measure answers: the model’s use (SHAP, permutation) or the data’s need (drop-column); a feature can be used and not needed.
5. Look for interactions separately: an effect through a product can be missed by every marginal screen.

Importance also lies with time. A feature’s importance in a model fitted on ten years is an average over regimes; a feature that mattered in the first five years and not the last five gets half its old importance and looks dependable. Chapter 12 measures importance on rolling windows, and chapter 21 asks whether an explanation is stable across retrains.

## 6.6 Tutorial: which feature mattered?

**Goal.** Compute the four importances and their clustered versions on the planted task, run [stability selection](#def-ml-feature-engineering-selection-and-importance-selection) over 100 candidates, and transform a heavy-tailed feature three ways. **End state:** [Figure 6.1](#fig-ml-importance-bars), the chapter’s numbers.

1. **Permutation and drop-column importance**, singly or by group. `def permutation_importance (predict, X, y, score, reps: int = 5 , seed: int = 0 , groups=None ): rng = np.random.default_rng(seed) base = score(y, predict(X)) out = [] for g in _groups(X.shape[1 ], groups): drops = [] for _ in range (reps): Z = X.copy() perm = rng.permutation(len (X)) Z[:, g] = X[perm][:, g] # a group is shuffled jointly drops.append(base - score(y, predict(Z))) out.append(np.mean(drops)) return np.array(out) def drop_column_importance (fit_predict, X, y, Xt, yt, score, groups=None ): """fit_predict(X, y, Xt) -> predictions on Xt of a model fitted on (X, y).""" base = score(yt, fit_predict(X, y, Xt)) out = [] for g in _groups(X.shape[1 ], groups): keep = [j for j in range (X.shape[1 ]) if j not in g] out.append(base - score(yt, fit_predict(X[:, keep], y, Xt[:, keep]))) return np.array(out)` **Listing 6.1.** Held-out importances, for single features or groups. code/firm/featimp/firm_featimp.py
2. **Clusters and [stability selection](#def-ml-feature-engineering-selection-and-importance-selection).** `def cluster_features (X, threshold: float = 0.5 ): from scipy.cluster.hierarchy import fcluster, linkage from scipy.spatial.distance import squareform C = np.corrcoef(X, rowvar=False ) D = 1.0 - np.abs(C) np.fill_diagonal(D, 0.0 ) Z = linkage(squareform(D, checks=False ), method=" average " ) lab = fcluster(Z, t=threshold, criterion=" distance " ) return [list (np.flatnonzero(lab == k)) for k in np.unique(lab)] def stability_selection (select, X, y, n_sub: int = 50 , frac: float = 0.5 , seed: int = 0 ): """select(X, y) -> boolean (p,) mask of selected features; run on n_sub random subsamples of size frac * n.""" rng = np.random.default_rng(seed) n = len (y) freq = np.zeros(X.shape[1 ]) for _ in range (n_sub): idx = rng.choice(n, int (frac * n), replace=False ) freq += np.asarray(select(X[idx], y[idx]), float ) return freq / n_sub def false_selection_bound (q: float , p: int , threshold: float ) -> float : """Expected number of falsely selected variables when each subsample selects q of p variables and the selection frequency threshold is in (0.5, 1] (Meinshausen and Buhlmann, 2010, under exchangeability).""" return q * q / ((2.0 * threshold - 1.0 ) * p)` **Listing 6.2.** Correlation clusters, stability selection and its bound. code/firm/featimp/firm_featimp.py
3. **Run** `ml_importance.importances()` , `twins()` , `stability()` , `heavy_tail()` and `fig_importance.py` .

**What to change next.** Give $x_3$ a small main effect and see whether [stability selection](#def-ml-feature-engineering-selection-and-importance-selection) now finds the interaction; raise the copy’s correlation to 0.999 and watch its [SHAP value](#def-ml-feature-engineering-selection-and-importance-shapley) overtake $x_1$’s.

## 6.7 Build: feature importance and selection

**Purpose.** Every model report carries held-out, clustered importances and a record of how its features were selected.

**Interface.** `winsorise`, `rank_gauss`, `mdi(model)`, `permutation_importance(predict, X, y, score, reps, seed, groups)`, `drop_column_importance(fit_predict, X, y, Xt, yt, score, groups)`, `shap_values(model, X)`, `cluster_features(X, threshold)`, `stability_selection(select, X, y, n_sub, frac, seed)`, `false_selection_bound(q, p, threshold)`.

**Rules.** Importances on held-out rows; groups permuted jointly with one permutation of the rows; the selection procedure is a callable, so any model can be stability-selected.

**Acceptance tests.** `code/firm/featimp/tests/`: a pure-noise feature has near-zero [permutation importance](#def-ml-feature-engineering-selection-and-importance-mdi) and a duplicated signal is found only as a group; [SHAP values](#def-ml-feature-engineering-selection-and-importance-shapley) sum to the prediction minus the bias; clustering puts a near-copy with its original; [stability selection](#def-ml-feature-engineering-selection-and-importance-selection) keeps a strong feature and drops noise; the bound’s formula.

**Stretch.** Conditional permutation (permute within bins of the correlated features); importance on rolling windows; SHAP interaction values for pairs.

Sources and further reading

- N. Meinshausen and P. Bühlmann, “Stability selection”, *Journal of the Royal Statistical Society B* 72(4), 2010.
- S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions”, *NeurIPS* , 2017; S. M. Lundberg et al., “From local explanations to global understanding with explainable AI for trees”, *Nature Machine Intelligence* 2, 2020.
- C. Strobl, A.-L. Boulesteix, A. Zeileis and T. Hothorn, “Bias in random forest variable importance measures”, *BMC Bioinformatics* 8, 2007.
- M. López de Prado, *Machine Learning for Asset Managers* , Cambridge University Press, 2020.
- L. S. Shapley, “A value for $n$ -person games”, in *Contributions to the Theory of Games II* , 1953.

## 6.8 Exercises

**Exercise 6.1 ★.**

[Stability selection](#def-ml-feature-engineering-selection-and-importance-selection) over 200 candidates selects 10 per half-sample on average. What does the bound give at thresholds 0.6 and 0.9?

**Solution of Exercise 6.1.**

$10^2/((2\times0.6 - 1)\times200) = 2.5$ expected false selections at 0.6; $100/(0.8\times200) = 0.625$ at 0.9.

**Exercise 6.2 ★.**

Three players: $v(\emptyset) = 0$, singletons 1, 1, 0; pairs $\{1,2\} = 3$, $\{1,3\} = 1$, $\{2,3\} = 1$; all three 3. Compute the [Shapley values](#def-ml-feature-engineering-selection-and-importance-shapley).

**Solution of Exercise 6.2.**

Averaging marginal contributions over the six orderings: player 1 contributes $1, 1, 2, 2, 1, 2$, so $\phi_1 = 1.5$; player 2 likewise $\phi_2 = 1.5$; player 3 adds nothing in any ordering, $\phi_3 = 0$. They sum to $v(P) = 3$.

**Exercise 6.3 ★.**

A held-out R-squared falls from 2.67% to 2.58% when a feature is dropped and the model refitted. What is its drop-column importance, and what can you not conclude from it?

**Solution of Exercise 6.3.**

0.09 points. It does not show that the feature is unimportant: another feature may stand in for it (the [substitution effect](#def-ml-feature-engineering-selection-and-importance-substitution)); only the importance of its cluster answers that.

**Exercise 6.4 ★★.**

Why does permuting $x_1$ and $x_{1b}$ together cost more (2.72 points) than the model’s whole R-squared (2.67%)?

**Solution of Exercise 6.4.**

The model uses both copies. Permuted together, both feed it noise with the variance of the signal they carried, so its predictions become noise of the same size, and a noisy prediction scores worse than predicting the mean: R-squared falls to $2.67 - 2.72 = -0.05\%$.

**Exercise 6.5 ★★.**

Why does MDI give the noise features up to 5.4% of the total while [permutation importance](#def-ml-feature-engineering-selection-and-importance-mdi) gives them 0.2%?

**Solution of Exercise 6.5.**

MDI adds up the training-loss reduction of every split, and on training data a split on noise does reduce the loss (it fits noise). [Permutation importance](#def-ml-feature-engineering-selection-and-importance-mdi) is measured on held-out rows, where those splits help nothing.

**Exercise 6.6 ★★.**

*Find the flaw.* “We dropped every feature whose [permutation importance](#def-ml-feature-engineering-selection-and-importance-mdi) was below the median, one at a time, refitting after each drop, and kept what was left.”

**Solution of Exercise 6.6.**

One-at-a-time pruning removes correlated copies one after the other: each looks unimportant while the other remains, and the second removal then costs what neither looked worth (1.05 points here). Prune clusters, not features. If the importances were computed on the test data, the selection is also a leak (chapter 3).

**Exercise 6.7 ★★★.**

*Coding.* In `ml_importance.heavy_tail`, replace the $t$ with 1.5 degrees of freedom by one with 5. Report the three R-squared and explain why the gap between raw and ranked closes.

**Solution of Exercise 6.7.**

`heavy_tail(df=5.0)`: raw 1.36%, winsorised 1.49%, ranked 1.57%, truth 1.79%. With five degrees of freedom extreme values are rare and moderate, so they no longer set the slope; the transforms matter much less than with 1.5.

**Exercise 6.8 ★★★.**

Show that if $x_{1b} = x_1$ exactly and a model’s prediction depends on $x_1 + x_{1b}$ only, the [SHAP values](#def-ml-feature-engineering-selection-and-importance-shapley) of the two are equal, and that the [permutation importance](#def-ml-feature-engineering-selection-and-importance-mdi) of the pair is not the sum of their single importances.

**Solution of Exercise 6.8.**

If the prediction depends on $x_1 + x_{1b}$ only and the two are identical, the value function is symmetric in them, and the symmetry axiom gives equal [Shapley values](#def-ml-feature-engineering-selection-and-importance-shapley). Take $p = \gamma(x_1 + x_{1b}) = 2\gamma x_1$ with $\Var(x_1) = \sigma^2$ and residuals uncorrelated with the features. Permuting $x_1$ alone replaces $p$ by $\gamma(x_{1,\pi} + x_1)$ and adds $\E[\gamma^2(x_1 - x_{1,\pi})^2] = 2\gamma^2\sigma^2$ to the mean squared error; the same for $x_{1b}$; permuting both with one permutation adds $\E[4\gamma^2(x_1 - x_{1,\pi})^2] = 8\gamma^2\sigma^2$, twice the sum of the singles.

## 6.9 Problem: Which Feature Mattered?

**Problem 6.1.**

Weekend problem — planted roles, four importances

The chapter’s task: $x_1$ linear, $x_2$ squared, $x_3x_4$, a copy $x_{1b}$, ten noise features; boosted trees on 20 000 rows, scored on 20 000 more.

**Part I — Transformations.**

1. What do the raw, winsorised and ranked versions of the heavy-tailed feature score, against the truth’s 2.17%?
2. Why does ranking beat winsorising here?
3. What does the book use by default, and why?
4. What does the model score, and the truth?

**Part II — Four importances.**

5. How does each measure rank $x_1$ to $x_4$ ?
6. Why does SHAP credit the copy more than $x_3$ or $x_4$ ?
7. What does MDI give the largest noise feature, and why?
8. Which measures answer the question “does the data need this feature”?

**Part III — The copy.**

9. What do $x_1$ and $x_{1b}$ cost permuted singly and together?
10. What do they cost dropped singly and together?
11. Which group does the clustering find, at what distance?
12. What would one-at-a-time pruning have done?

**Part IV — Selection and the verdict.**

13. What do the lasso and boosting keep at a threshold of 0.6 over 100 candidates, and what does the bound allow?
14. Why do both miss $x_3$ and $x_4$ ?
15. State the *named result* : the ranks of the signal features under each method, and the false-selection rate of [stability selection](#def-ml-feature-engineering-selection-and-importance-selection) at 0.6.
16. Which importance would you show a portfolio manager, and which a risk committee?
17. How would you make [stability selection](#def-ml-feature-engineering-selection-and-importance-selection) see the interaction?
18. What would change with features that drift over time?
19. Write the importance section of a model report in four lines.
20. In one sentence: what does an importance measure?

**Solution of Problem 6.1.**

1. 0.13%, 1.27% and 1.96%.
2. Winsorising at 1% still leaves large values in a distribution this heavy; ranking removes the scale entirely and matches the saturating effect.
3. Cross-sectional ranks by date: robust to tails and outliers, comparable across dates.
4. 2.67% held out; the truth 4.01%.
5. Gain, permutation and drop-column rank $x_2$ first and $x_1$ fourth; SHAP ranks $x_2$ , $x_1$ , the copy, then $x_3$ and $x_4$ .
6. SHAP explains the fitted model, and the model uses the copy.
7. 5.4% of the total (against 9.5% for $x_1$ ): gain is measured on training data, where noise splits reduce the loss.
8. Drop-column importance (refit without the feature), and [permutation importance](#def-ml-feature-engineering-selection-and-importance-mdi) on held-out data only in part.
9. 0.76 and 0.41 points singly, 2.72 together.
10. 0.09 and $-0.02$ singly, 1.05 together.
11. $\{x_1, x_{1b}\}$ at a distance of 0.5 (correlation 0.95), all other features alone.
12. Dropped one copy, then the other, losing 1.05 of 2.67 points.
13. The lasso keeps $x_1$ (6.7 selected per half-sample; bound 2.3); boosting keeps $x_1$ and $x_2$ (bound 1.8); no noise feature in either.
14. Their effect is only through the product: no linear screen sees it, and a greedy tree’s first split on either gains nothing.
15. *Named result* : permutation and drop-column importance rank the signals $x_2, x_3, x_4, x_1$ (drop-column $x_2, x_4, x_3, x_1$ ) and the copy far below; SHAP puts the copy third; the false-selection rate of [stability selection](#def-ml-feature-engineering-selection-and-importance-selection) at 0.6 is zero for both selectors (bounds 2.3 and 1.8 false features), and both miss the interaction.
16. A portfolio manager: clustered drop-column importance (what the data need); a risk committee: SHAP for individual decisions (what the model did).
17. Add candidate products or use a selector that can split on pairs (boosting with more trees, or interaction constraints lifted), and give the pair a small main effect to find.
18. Importances should be computed on rolling windows; a feature’s full-sample importance averages regimes.
19. Held-out clustered importances with the method named; the copies and their cluster; features that no method ranks above noise; the selection procedure and its false-selection bound.
20. How much a fitted model, or the data, depend on a feature, in one sense that must be named.

## 6.10 Interview questions

**Interview question 6.1 ★ researcher, mle.**

What is the difference between impurity-based and [permutation importance](#def-ml-feature-engineering-selection-and-importance-mdi), and which do you trust?

**Solution of Interview question 6.1.**

Impurity importance sums the training-loss reductions of splits on a feature: biased toward features with many split points and toward noise that was fitted. [Permutation importance](#def-ml-feature-engineering-selection-and-importance-mdi) measures the held-out score drop when the feature is shuffled. Trust the held-out one, clustered by correlation.

*What the interviewer is looking for: training versus held-out, and the known biases of MDI.*

**Interview question 6.2 ★★ researcher.**

Two of your features are 95% correlated and both show low importance. What do you do?

**Solution of Interview question 6.2.**

Measure them as a group (permute or drop both): substitution hides each behind the other. If the group matters, keep one or a combination (their mean or first principal component), chosen for stability and cost, not for its single importance.

*What the interviewer is looking for: the [substitution effect](#def-ml-feature-engineering-selection-and-importance-substitution) and clustered importance.*

**Interview question 6.3 ★★ mle.**

Explain [SHAP values](#def-ml-feature-engineering-selection-and-importance-shapley) and one situation where they mislead.

**Solution of Interview question 6.3.**

The [Shapley values](#def-ml-feature-engineering-selection-and-importance-shapley) of the features for one prediction, with the model’s expected output given subsets of features as the game; they add up to the prediction minus the average. They mislead with correlated features (credit is shared by a copy the model happens to use) and when read as causal effects rather than descriptions of the model.

*What the interviewer is looking for: the additivity property and the correlated-feature caveat.*

**Interview question 6.4 ★★ researcher.**

How does [stability selection](#def-ml-feature-engineering-selection-and-importance-selection) control false discoveries, and what does it not control?

**Solution of Interview question 6.4.**

By keeping only features selected in most half-samples; under exchangeability the expected number of noise features kept is at most $q^2/((2\pi_{\mathrm{thr}} - 1)p)$. It does not control misses: features whose effect the base selector cannot see (interactions, nonlinearities for a lasso) are never selected.

*What the interviewer is looking for: the bound and its blind spot.*

**Interview question 6.5 ★★ researcher, trader.**

A volume feature has occasional values a hundred times its median. How do you feed it to a model?

**Solution of Interview question 6.5.**

Take logs, rank it by date or winsorise it at quantiles estimated on training data only, and check the transformed feature’s relation to the target is stable; consider a volume surprise (relative to its own history) rather than the level.

*What the interviewer is looking for: robust transforms fitted without look-ahead.*

**Interview question 6.6 ★★★ researcher.**

Construct a target and two features such that each feature alone has zero correlation with the target but together they predict it perfectly. What does this imply for feature screening?

**Solution of Interview question 6.6.**

Let $x_1, x_2$ be independent signs $\pm1$ and $y = x_1x_2$: each is uncorrelated with $y$, together they determine it. Marginal screens (correlations, univariate tests, a lasso) discard both; screening must allow pairs, or a model that can represent interactions must be the selector.

*What the interviewer is looking for: the XOR example and its consequence.*
