---
title: "Signal Combination"
book: "Research Craft: Predictors, Backtests, Measurement, Portfolios"
subject: quant
language: en
chapter: 14
exercises: 8
source: https://one-course.com/books/quant/7/en/chapter/14-signal-combination
---

# Chapter 14 — Signal Combination

A research team has forty signals, each with an [information coefficient](https://one-course.com/books/quant/7/en/chapter/6-anatomy-of-a-predictor#def-rs-anatomy-of-a-predictor-ic) near 0.015. Blended with equal weights they reach 0.064. Weighted by their past ICs they reach 0.099 in one simulated market and 0.075 in another; with the textbook weights that maximise the blend’s information ratio, 0.064 and 0.041. In a universe of a hundred names with a year of history, regression weights earn almost nothing in one of the two markets, and there nothing beats equal weights. Forecasters know this pattern as a puzzle: Stock and Watson (2004) found that the most successful combinations of output-growth forecasts, like the mean, were the least sensitive to the recent performance of the individual forecasts. This chapter builds the blends (`firm.combine`), explains what each assumes, and measures on planted data when the estimated weights are worth their estimation error.

## 14.1 The geometry of a linear blend

**Definition 14.1 (Signal combination, composite signal, equal-weight blend).**

*Signal combination* is the construction of one [predictor](https://one-course.com/books/quant/7/en/chapter/6-anatomy-of-a-predictor#def-rs-anatomy-of-a-predictor-predictor) from several. A *composite signal* is a weighted sum $F = \sum_k w_k z_k$ of the signals $z_k$, each standardised across names on every date; the *equal-weight blend* sets every $w_k = 1/K$.

**Proposition 14.2 (The IC of an equal-weight blend).**

If $K$ standardised signals each have correlation $\rho$ with the target and pairwise correlation $c$ with each other, their [equal-weight blend](#def-rs-signal-combination-blend) has correlation $K\rho/\sqrt{K + K(K - 1)c}$ with the target, which rises with $K$ toward $\rho/\sqrt c$.

**Proof.** The covariance of $\sum_k z_k$ with the (standardised) target is $K\rho$; its variance is $K + K(K-1)c$. Divide, and let $K \to \infty$. ∎

The ceiling is set by the correlation between signals, not by their number: forty signals that repeat each other are one signal. The chapter’s first market (`rs_combine`, the *stable* world) plants the structure: each month’s target is a name’s next-month return on `firm.synthmkt` (market-adjusted, standardised across the 704 names listed throughout) plus an expected return from eight latent themes, five of which carry information (from 0.10 down to 0.03) and three none; each of the forty signals measures one theme, five signals a theme, through a noise its theme shares and a noise of its own whose size grows from 1 to 4 within the theme. Over the last five years the signals’ mean IC is 0.015 and their mean pairwise correlation 0.027. The measured blend of $K$ signals taken in random order follows the proposition closely ([Figure 14.1](#fig-rs-signal-combination-geometry)): 0.021 for two, 0.042 for ten, 0.064 for all forty against 0.065 from the formula, on the way to a ceiling of 0.090.

![Out-of-sample IC of the equal-weight blend of K of the forty planted signals (stable world, last five years), against the proposition with their mean IC 0.015 and mean pairwise correlation 0.027. Data: rs_combine on firm.synthmkt.](https://one-course.com/images/onecourse/chapters/quant-7/rs-signal-combination/fig-903b74c4a067.svg)

***Figure 14.1.** Out-of-sample IC of the [equal-weight blend](#def-rs-signal-combination-blend) of $K$ of the forty planted signals (stable world, last five years), against the proposition with their mean IC 0.015 and mean pairwise correlation 0.027. Data: `rs_combine` on `firm.synthmkt`.*

## 14.2 Weights from estimates

**Definition 14.3 (IC-weighted blend, maximum-ICIR blend).**

An *IC-weighted blend* sets each weight proportional to the signal’s past mean IC (zero if negative). A *maximum-ICIR blend* sets $w \propto \Sigma^{-1}\mu$, with $\mu$ the vector of the signals’ past mean ICs and $\Sigma$ the covariance of their IC time series: the weights that maximise the blended IC’s mean over its standard deviation.

Equal weights assume nothing about the signals; IC weights assume their quality differs and persists; the maximum-ICIR rule assumes, in addition, that the covariance of forty IC series estimated from 59 months is accurate enough to invert. The second market (the *drifting* world) removes the first two premises: every signal has the same noise, and each theme’s information wanders around a common mean of 0.04 as an autoregression with a half-life of about two years, so equal weights are right on average and any estimated weights chase noise and stale information. Fitted on the first five years and judged on the last five, with all 704 names:

|  | stable world | drifting world |
| --- | --- | --- |
| blend | in sample | out of sample | in sample | out of sample |
| equal weights | 0.065 | 0.064 | 0.055 | 0.060 |
| IC weights | 0.091 | 0.099 | 0.073 | 0.075 |
| maximum-ICIR, no shrinkage | 0.050 | 0.064 | 0.052 | 0.041 |
| maximum-ICIR, diagonal $\Sigma$ | 0.090 | 0.098 | 0.074 | 0.074 |
| pooled regression (OLS) | 0.093 | 0.100 | 0.076 | 0.071 |
| lasso | 0.093 | 0.102 | 0.075 | 0.071 |
| the truth (oracle weights) | — | 0.102 | — | 0.087 |

With 704 names a month, IC weights capture most of what the truth offers in the stable world, and still beat equal weights in the drifting one (information that drifts over two years is still there a year later). The unshrunk maximum-ICIR rule is the worst estimated blend in both: it inverts a covariance matrix of forty noisy series and loads on its smallest eigenvalues. Shrinking that covariance to its diagonal, the shrinkage of Book 4 (chapters 14 and 22) taken to its limit, comes within about a thousandth of the IC-weighted result in both worlds (0.098 against 0.099, 0.074 against 0.075).

## 14.3 Regularised combination

The weights of a pooled regression of the target on the forty signals are estimated from $59 \times 704$ observations, and there the regression is as good as IC weights. The puzzle appears when the sample shrinks. In a universe of 100 names with twelve months of history ([Figure 14.2](#fig-rs-signal-combination-shrink)), unpenalised regression weights earn an out-of-sample IC of 0.073 in the stable world and 0.005 in the drifting one, against 0.069 for equal weights in both. Shrinking the regression weights toward equal weights, $w_\lambda = (1 - \lambda)\hat w + \lambda/K$, is the simplest regularisation, and $\lambda$ has an optimum: 0.4 in the stable world (IC 0.081), 1 in the drifting world (equal weights). Smith and Wallis (2009) located the puzzle exactly there: the finite-sample error in estimating the combining weights, which can exceed the gain from estimating them at all.

![A universe of 100 names, blends fitted on twelve months and judged on the next sixty: out-of-sample IC of the pooled regression weights shrunk toward equal weights (= 1), and of IC weights (dashed). The stable world’s optimum is interior (= 0.4); the drifting world’s is equal weights. Data: rs_combine.small_universe.](https://one-course.com/images/onecourse/chapters/quant-7/rs-signal-combination/fig-b567f7ae5086.svg)

***Figure 14.2.** A universe of 100 names, blends fitted on twelve months and judged on the next sixty: out-of-sample IC of the pooled regression weights shrunk toward equal weights ($\lambda = 1$), and of IC weights (dashed). The stable world’s optimum is interior ($\lambda = 0.4$); the drifting world’s is equal weights. Data: `rs_combine.small_universe`.*

**Definition 14.4 (Forecast combination puzzle).**

The *forecast combination puzzle* is the repeated empirical finding that simple combinations of forecasts (the mean) outperform combinations with estimated weights.

The window matters less than the world: at 6, 12, 24 and 48 months of history the stable world’s IC weights beat equal weights (0.085, 0.086, 0.082 and 0.092 against 0.069), and the drifting world’s lose until four years of history (0.043, 0.049, 0.067, then 0.074). Ridge and lasso penalise the regression toward zero rather than toward equal weights; on a standardised panel that shrinks the composite’s scale more than its direction, and in the full panel they moved the out-of-sample IC by at most 0.003 (ridge with a penalty of 1 in the drifting world). The penalty that matters is toward the right prior, and for signals of similar quality and unstable information the right prior is equal weights.

## 14.4 Stacking

**Definition 14.5 (Stacking).**

*Stacking* combines several [predictors](https://one-course.com/books/quant/7/en/chapter/6-anatomy-of-a-predictor#def-rs-anatomy-of-a-predictor-predictor) (here, several blends) with weights fitted on their predictions for data not used to build them: Breiman’s (1996) stacked regressions use cross-validated predictions and least squares under non-negativity constraints, an idea he credits to Wolpert (1992).

Time-ordered folds are compulsory: a blend fitted on months that follow the months it predicts is a leak. `firm.stack` fits each base blend (equal weights, IC weights, regression weights) on an expanding window, collects its predictions for the next block of months, and regresses the target on them without negative weights. On the full panel the stack put all its weight on IC weights in both worlds, matching their out-of-sample IC (0.099 and 0.075): [stacking](#def-rs-signal-combination-stacking) chooses among blends, and it is only as good as the best of them and the length of its folds.

## 14.5 Orthogonalisation

**Definition 14.6 (Orthogonalisation, symmetric orthogonalisation).**

*Orthogonalisation* transforms a set of signals, on each date, into uncorrelated ones spanning the same space: *sequentially* (Gram–Schmidt), each signal residualised on those before it, so that the order decides who keeps the shared part; or by *symmetric orthogonalisation*, $Z S^{-1/2}$ with $S$ the signals’ correlation matrix, the orthonormal set closest to the original signals, a construction that goes back to Löwdin’s orthonormalisation of overlapping atomic orbitals (1950).

Orthogonalised signals make weights readable (each weight is the marginal contribution of information no other signal carries) and make risk-model exposures separable (residualising on a book’s factor exposures is the same operation). They do not create information. Equal weights on the symmetrically orthogonalised signals, recomputed every month, give an out-of-sample IC of 0.064 in the stable world, the same as the raw signals (0.064), with each orthogonalised signal correlated 0.95 with its original; sequential [orthogonalisation](#def-rs-signal-combination-orth) in the order of the in-sample ICs gives 0.050, because the best signal keeps its theme’s shared part and the others are reduced to their noise before being weighted equally with it.

## 14.6 Predictor cards

**Predictor card 14.1 — Composite of forty signals.**

**Definition.** $\sum_k w_k z_k$ with $z_k$ standardised by date and $w$ the regression weights shrunk toward equal weights by $\lambda$ chosen on time-ordered folds.

**Inputs and timestamps.** Each signal as known at the decision date; the weights refitted only on dates before it.

**Rationale.** The equal-weight ceiling $\bar\rho/\sqrt{\bar c}$ ([Proposition 14.2](#prop-rs-signal-combination-geometry)); weights earn their keep only when quality differs persistently.

**Horizon and half-life.** Those of the constituents, weighted.

**Normalisation.** Standardise the composite by date; neutralise to the book (chapter 6).

**Failure modes.** Unshrunk inverse-covariance weights; refitting with look-ahead; a composite dominated by one theme’s repeated signals.

**Sources.** Stock and Watson (2004); Smith and Wallis (2009); `rs_combine`.

## 14.7 Tutorial: forty weak signals

**Goal.** Blend forty planted signals in two worlds with every rule of the chapter, in sample and out, on the full panel and on a small universe, and find the shrinkage toward equal weights that wins. **End state:** Figures [14.1](#fig-rs-signal-combination-geometry) and [14.2](#fig-rs-signal-combination-shrink); the table.

1. **The maximum-ICIR weights**, with the covariance shrunk toward its diagonal. `def max_icir (ic, shrink: float = 0.0 ) -> np.ndarray: """The weights that maximise the mean over the standard deviation of the blended IC, with the IC covariance shrunk toward its diagonal (shrink = 1: each signal weighted by its own mean IC over its IC variance).""" ic = np.asarray(ic, float ) S = np.cov(ic, rowvar=False ) S = (1.0 - shrink) * S + shrink * np.diag(np.diag(S)) w = np.linalg.solve(S, ic.mean(axis=0 )) return w / np.abs(w).sum()` **Listing 14.1.** Maximum-ICIR weights with a shrunk IC covariance. code/firm/combine/firm_combine.py
2. **[Symmetric orthogonalisation](#def-rs-signal-combination-orth)** of one date’s cross-section. `def orth_symmetric (X) -> np.ndarray: X = np.asarray(X, float ) - np.asarray(X, float ).mean(axis=0 ) S = X.T @ X / len (X) val, vec = np.linalg.eigh(S) return X @ (vec @ np.diag(val ** -0.5 ) @ vec.T)` **Listing 14.2.** The orthonormal signals closest to the originals. code/firm/combine/firm_combine.py
3. **Run** `results()` for both worlds, `small_universe()` for windows of 6 to 48 months, `orthogonalised()` , `geometry()` and `fig_combine.py` .

**What to change next.** Make the drifting world’s half-life six months and watch IC weights lose even on the full panel; give three signals a negative IC and compare IC weights (which drop them) with equal weights (which keep them).

## 14.8 Build: the blender

**Purpose.** One place where the firm’s signals become a composite, with every rule of the chapter behind the same interface, and the out-of-sample comparison that chooses among them.

**Interface.** `zscore`, `ic_series`, `ic_matrix`, `equal`, `blend(X, w)`, `ic_weights`, `max_icir(ic, shrink)`, `toward_equal(w, lam)`, `ridge`, `lasso`, `nnls`, `stack(preds, y, folds)`, `time_folds`, `orth_sequential`, `orth_symmetric`, `residualise_on`.

**Rules.** Signals are standardised by date before any weight is applied; weights and hyperparameters are fitted on dates before the ones they are judged on; the [equal-weight blend](#def-rs-signal-combination-blend) is always reported alongside.

**Acceptance tests.** `code/firm/combine/tests/`: z-scores and ICs by hand; a planted panel where the good signal earns the largest IC weight, the diagonal maximum-ICIR equals mean over variance, ridge shrinks, lasso zeroes the noise signal; NNLS by hand; a stack that picks the informative base; both orthogonalisers orthonormal, the symmetric one closest to the originals, the sequential one keeping its first column.

**Stretch.** Weights by regime; hierarchical blends (within themes, then across); Bayesian weights with a prior at equal weights.

Sources and further reading

- J. H. Stock and M. W. Watson, “Combination forecasts of output growth in a seven-country data set”, *Journal of Forecasting* 23(6), 2004.
- J. Smith and K. F. Wallis, “A simple explanation of the forecast combination puzzle”, *Oxford Bulletin of Economics and Statistics* 71(3), 2009.
- L. Breiman, “Stacked regressions”, *Machine Learning* 24, 1996; D. H. Wolpert, “Stacked generalization”, *Neural Networks* 5(2), 1992.
- P.-O. Löwdin, “On the non-orthogonality problem connected with the use of atomic wave functions in the theory of molecules and crystals”, *Journal of Chemical Physics* 18(3), 1950.

## 14.9 Exercises

**Exercise 14.1 ★.**

Ten signals each have an IC of 0.02 and a pairwise correlation of 0.25. What is the IC of their [equal-weight blend](#def-rs-signal-combination-blend), and its limit as the number of such signals grows?

**Solution of Exercise 14.1.**

$10 \times 0.02/\sqrt{10 + 90 \times 0.25} = 0.2/\sqrt{32.5} = 0.035$; the limit is $0.02/\sqrt{0.25} = 0.04$.

**Exercise 14.2 ★.**

Three signals have past mean ICs of 0.03, 0.01 and $-0.01$. Give their IC weights, and their weights shrunk halfway toward equal weights.

**Solution of Exercise 14.2.**

IC weights use the positive parts 0.03, 0.01 and 0: weights 0.75, 0.25 and 0. Halfway toward equal weights: $\frac12(0.75, 0.25, 0) + \frac12(\frac13, \frac13, \frac13) = (0.542, 0.292, 0.167)$.

**Exercise 14.3 ★.**

Why must [stacking](#def-rs-signal-combination-stacking) folds be ordered in time for financial data, when Breiman’s [stacking](#def-rs-signal-combination-stacking) used ordinary cross-validation?

**Solution of Exercise 14.3.**

Returns and signals are ordered in time, correlated across neighbouring dates and exposed to regimes; a fold that trains on later dates and tests on earlier ones lets the [stacking](#def-rs-signal-combination-stacking) weights learn from the future (which blend worked in the test period’s own regime). Expanding windows that always predict forward reproduce what the firm can do in production.

**Exercise 14.4 ★★.**

Two signals have ICs $\rho_1 = 0.03$, $\rho_2 = 0.02$ and correlation $c = 0.5$. Find the weights of the best linear blend (for correlation with the target) and its IC, and compare with equal weights.

**Solution of Exercise 14.4.**

$w \propto C^{-1}\rho$ with $C = \begin{pmatrix} 1 & 0.5\\ 0.5 & 1\end{pmatrix}$: $C^{-1}\rho = \frac{1}{0.75}(0.03 - 0.01, -0.015 + 0.02) = (0.0267, 0.0067)$, weights 0.8 and 0.2. The best IC is $\sqrt{\rho^\top C^{-1}\rho} = \sqrt{0.03 \times 0.0267 + 0.02 \times 0.0067} = 0.0306$; equal weights give $0.05/\sqrt{2 + 2 \times 0.5} = 0.0289$.

**Exercise 14.5 ★★.**

Show that sequential [orthogonalisation](#def-rs-signal-combination-orth) of two standardised signals with correlation $c$ leaves the first unchanged and replaces the second by $(z_2 - cz_1)/\sqrt{1 - c^2}$. What is the IC of the new second signal in terms of $\rho_1$, $\rho_2$ and $c$?

**Solution of Exercise 14.5.**

The first signal is its own first Gram–Schmidt vector. The second, residualised on the first, is $z_2 - cz_1$, with variance $1 - c^2$; standardised, $(z_2 - cz_1)/\sqrt{1 - c^2}$. Its correlation with the target is $(\rho_2 - c\rho_1)/\sqrt{1 - c^2}$: what the second signal knows that the first does not.

**Exercise 14.6 ★★.**

Why does inverting the covariance of forty IC series estimated from sixty months produce extreme weights? Relate it to the eigenvalues of a sample covariance matrix (Book 4, chapter 22).

**Solution of Exercise 14.6.**

Sixty observations of forty series give a sample covariance whose smallest eigenvalues are far below the true ones (the spread of sample eigenvalues of Book 4, chapter 22), and $\Sigma^{-1}\mu$ divides by them: directions that look almost riskless because of sampling error receive the largest weights. Shrinking toward the diagonal lifts those eigenvalues; at the diagonal the rule becomes mean over variance per signal, close to IC weights.

**Exercise 14.7 ★★★.**

*Coding.* With `small_universe`, find the optimal $\lambda$ in the stable world for windows of 6, 24 and 48 months. Does the optimum move toward the regression weights as the window lengthens?

**Solution of Exercise 14.7.**

The optimal $\lambda$ (on a grid of tenths) is 0.5 at 6 months (IC 0.074), 0.4 at 12, 0.6 at 24 and 0.4 at 48 (0.083). It does not move steadily toward the regression weights: with 100 names even four years leave enough error in forty regression weights that a large pull toward equal weights pays, and the optimum is noisy from one window to the next.

**Exercise 14.8 ★★★.**

*Find the flaw.* “We fitted the blend weights by regression on the full ten years and they beat equal weights by 40% over the same period.”

**Solution of Exercise 14.8.**

The weights were fitted on the period they are judged on: the comparison measures the fit, not the forecast. Refit on dates before each judged date (expanding windows or time-ordered folds), compare out of sample, and report equal weights alongside; on the chapter’s full panel the in-sample advantage of the unshrunk maximum-ICIR rule reversed out of sample in the drifting world.

## 14.10 Problem: Forty Weak Signals

**Problem 14.1.**

Weekend problem — when weights are worth estimating

The chapter’s two planted worlds on `firm.synthmkt`: forty signals in eight themes, 704 names, ten years of months split in two; then a universe of 100 names.

**Part I — The geometry.**

1. What are the signals’ mean out-of-sample IC and mean pairwise correlation in the stable world?
2. What does [Proposition 14.2](#prop-rs-signal-combination-geometry) give for forty signals, and what was measured?
3. What is the ceiling, and why is it not reached by adding signals?
4. Which planted feature makes the signals correlated?

**Part II — Full panel.**

5. Give the out-of-sample ICs of equal weights, IC weights and the truth in both worlds.
6. Why does the unshrunk maximum-ICIR rule do worst?
7. What does shrinking its covariance to the diagonal give?
8. Why do IC weights still beat equal weights in the drifting world?

**Part III — Small universe.**

9. With 100 names and twelve months, give the out-of-sample IC of unpenalised regression weights in each world.
10. State the *named result* : the out-of-sample IC of the best blend against equal weights, and the shrinkage toward equal weights that maximises it, in each world.
11. How do IC weights fare against equal weights at 6, 12, 24 and 48 months in each world?
12. What did Smith and Wallis identify as the explanation of the puzzle?

**Part IV — Beyond.**

13. What weights did [stacking](#def-rs-signal-combination-stacking) choose, and why could it do no better?
14. What do equal weights on symmetrically and on sequentially orthogonalised signals give, and why do they differ?
15. When should a team orthogonalise its signals?
16. How would you choose $\lambda$ in production?
17. What would make equal weights the wrong prior?
18. Why do ridge and lasso change the full-panel result so little here?
19. What should every blend report beside its own IC?
20. In one sentence: when are estimated weights worth estimating?

**Solution of Problem 14.1.**

1. 0.015 and 0.027.
2. 0.065 from the proposition; 0.064 measured.
3. $\bar\rho/\sqrt{\bar c} = 0.090$ . Each new signal repeats part of its theme’s noise, shared with the other signals of the theme; that shared part does not average away, and the blend’s IC approaches the ceiling only as the number of signals grows without bound.
4. Five signals per theme share the theme’s noise.
5. Stable world: 0.064, 0.099 and 0.102. Drifting world: 0.060, 0.075 and 0.087.
6. It inverts the covariance of forty IC series estimated from 59 months and loads on its smallest, most underestimated eigenvalues.
7. 0.098 in the stable world and 0.074 in the drifting one: close to the IC-weighted result.
8. The themes’ information drifts with a half-life of about two years, so the first five years still say something about the next ones, and 704 names estimate the weights precisely.
9. 0.073 in the stable world, 0.005 in the drifting one.
10. **Named result.** With 100 names and a year of history, the best out-of-sample blend in the stable world is IC weights (0.086 against 0.069 for equal weights), and regression weights are best shrunk 40% toward equal weights (0.081); in the drifting world the best blend is equal weights ( $\lambda = 1$ , 0.069), and unshrunk regression weights earn 0.005.
11. Stable world: 0.085, 0.086, 0.082 and 0.092 against 0.069, better at every window. Drifting world: 0.043, 0.049, 0.067 and 0.074, worse until four years of history.
12. The finite-sample error in estimating the combining weights.
13. All its weight on IC weights, in both worlds. It chooses among its base blends on held-out dates; it cannot be better than the best of them by more than their errors’ diversification, and here one base dominated.
14. 0.064 for the symmetric set, as for the raw signals; 0.050 for the sequential set. The sequential order gives the best signal its theme’s shared part and leaves the others with residual noise, which equal weights then overweight; the symmetric set changes each signal as little as possible.
15. When weights must be interpreted (marginal information), when the book’s exposures must be removed (residualise on the risk model), or when a new signal must be judged on what it adds; not to create information.
16. On time-ordered folds within the training period, choosing the $\lambda$ that maximises the average out-of-fold IC, refitted on a schedule, with equal weights as the default when the choice is not clearly better.
17. Signals of very different and persistent quality, or signals with negative ICs: equal weights keep them all.
18. The panel is large (704 names a month), so the unpenalised regression is already precise, and penalising toward zero mostly rescales a composite that is standardised anyway.
19. The [equal-weight blend](#def-rs-signal-combination-blend) ’s IC on the same dates, in sample and out of sample.
20. When signal quality differs persistently and the sample is large enough that the weights’ estimation error is smaller than the differences they estimate.

## 14.11 Interview questions

**Interview question 14.1 ★ researcher.**

You have fifty alpha signals. How do you combine them?

**Solution of Interview question 14.1.**

Standardise each by date, start from equal weights, and measure the blend’s IC out of sample; then try IC weights and regression weights shrunk toward equal weights, with the shrinkage chosen on time-ordered folds; keep the estimated weights only if they beat equal weights out of sample by more than their noise. Check the signals’ correlation structure first: the blend’s ceiling is set by it.

**Interview question 14.2 ★★ researcher.**

What is the [forecast combination puzzle](#def-rs-signal-combination-puzzle), and what explains it?

**Solution of Interview question 14.2.**

Simple averages of forecasts repeatedly beat combinations with estimated weights (Stock and Watson; Smith and Wallis). The explanation: the error in estimating the weights from finite samples, and the instability of the forecasts’ relative quality, can exceed the gain from weighting.

**Interview question 14.3 ★★ researcher, mle.**

How would you use [stacking](#def-rs-signal-combination-stacking) for alpha signals without leaking?

**Solution of Interview question 14.3.**

Build base blends on expanding windows, predict the following block of dates with each, and fit non-negative [stacking](#def-rs-signal-combination-stacking) weights on those out-of-sample predictions only; never let a base or the stacker see dates it is later judged on, and compare the stack with its best base and with equal weights.

**Interview question 14.4 ★★ researcher.**

Adding a new signal with IC 0.02 did not improve your composite. Why might that be?

**Solution of Interview question 14.4.**

It may repeat information already in the composite (high correlation with existing signals: its residual IC, $(\rho - c\rho_1)/\sqrt{1 - c^2}$ for one correlated signal, can be near zero); it may be noisier than its IC suggests; or the weighting may have given it too little weight to matter. Measure its incremental IC after residualising on the composite (chapter 12).

**Interview question 14.5 ★★ researcher, risk.**

What is the difference between sequential and [symmetric orthogonalisation](#def-rs-signal-combination-orth), and when would you use each?

**Solution of Interview question 14.5.**

Sequential (Gram–Schmidt) [orthogonalisation](#def-rs-signal-combination-orth) keeps the first signal and residualises each next one on those before: the order decides who keeps the shared information, useful when there is a natural priority (a book’s factors first, new signals after). [Symmetric orthogonalisation](#def-rs-signal-combination-orth) ($ZS^{-1/2}$) treats all signals alike and gives the uncorrelated set closest to the originals: useful for interpretation without an order.

**Interview question 14.6 ★★★ researcher.**

Derive the weights that maximise the correlation of a linear blend with the target, given the signals’ ICs and correlation matrix, and explain why they are dangerous to estimate.

**Solution of Interview question 14.6.**

Maximise $w^\top\rho/\sqrt{w^\top Cw}$: the gradient condition gives $\rho(w^\top Cw) = Cw(w^\top\rho)$, so $w \propto C^{-1}\rho$ and the maximum is $\sqrt{\rho^\top C^{-1}\rho}$. Estimated $C^{-1}$ amplifies the sampling error of $C$’s smallest eigenvalues and $\hat\rho$ is itself noisy, so the weights are extreme and unstable; shrink $C$ toward its diagonal and $w$ toward equal weights.
