---
title: "Residual and Principal-Component Stat Arb"
book: "Strategies I: Equities and Futures"
subject: quant
language: en
chapter: 3
exercises: 8
source: https://one-course.com/books/quant/8/en/chapter/3-residual-and-principal-component-stat-arb
---

# Chapter 3 — Residual and Principal-Component Stat Arb

Regress each stock’s last sixty daily returns on its sector fund, add up the residuals, and trade the sum when it has drifted far from where it usually sits: buy when it is 1.25 of its standard deviations below, sell when it is 1.25 above, close near the middle. Avellaneda and Lee published this rule, with every parameter, and reported a Sharpe ratio after costs of 1.1 from 1997 to 2007 with sector funds, and 1.44 with principal components, and they reported the degradation too: 0.9 from 2003 to 2007. The method is public; the parameters are printed; what decides the result is whether the residuals actually revert. On the synthetic market, the same rule nets a Sharpe ratio of 3.0 when the stocks carry a mean-reverting component, and nothing when they do not. The build is `firm.residarb`.

## 3.1 Residuals against factors or principal components

The strategy starts where chapter 2 ended: with residual returns, but over weeks rather than a day. For each stock $i$ and each day, the last sixty daily returns are regressed on factor returns,

$$
R_{i,t} = \beta_{i,0} + \sum_j \beta_{ij} F_{j,t} + \epsilon_{i,t},
$$

and the residuals $\epsilon$ are the part of the stock’s moves its factors do not explain. Avellaneda and Lee used two kinds of factors. In the “ETF” version, the stock’s sector fund, one regressor per stock. In the “PCA” version, principal components of the whole market.

**Definition 3.1 (Eigenportfolio).**

An *eigenportfolio* is the portfolio that invests in each stock the corresponding entry of an eigenvector of the stocks’ return correlation matrix divided by the stock’s volatility, $Q^{(j)}_i = v^{(j)}_i/\sigma_i$; the returns of the leading eigenportfolios serve as the factors of a statistical factor model.

The correlation matrix is estimated on the last 252 days; Avellaneda and Lee kept fifteen [eigenportfolios](#def-s1-residual-and-principal-component-stat-arb-eigenportfolio), or as many as explain 55% of the variance. On the synthetic market, whose true structure has fourteen factors (the market, ten industries, three styles), fifteen [eigenportfolios](#def-s1-residual-and-principal-component-stat-arb-eigenportfolio) explain 45.7% of the variance, and the first alone 26.1%.

## 3.2 The residual as an Ornstein–Uhlenbeck process: the s-score

Add up the sixty residuals to get a path $X_1, \dots, X_{60}$, the stock’s cumulative idiosyncratic move. If the stock has a fair value relative to its factors, the path should wander around it and come back: an Ornstein–Uhlenbeck process (Book 4, chapter 4), which sampled daily is an AR(1),

$$
X_{n+1} = a + b\,X_n + \zeta_{n+1},\qquad \kappa = -252\ln b,\qquad m = \frac{a}{1-b},\qquad \sigma_{\mathrm{eq}} = \sqrt{\frac{\operatorname{Var}\zeta}{1-b^2}}.
$$

Here $\kappa$ is the speed of mean reversion (a year divided by it is the mean-reversion time), $m$ the level the path reverts to and $\sigma_{\mathrm{eq}}$ its standard deviation around that level.

**Definition 3.2 (s-score).**

The *s-score* of a stock is the distance of its cumulative residual from the level it reverts to, in units of its equilibrium standard deviation, $s = (X - m)/\sigma_{\mathrm{eq}}$, from an AR(1) fit of the cumulative residual over a trailing window. Because the regression’s residuals sum to zero, the last point is $X_{60} = 0$ and $s = -m/\sigma_{\mathrm{eq}}$; Avellaneda and Lee centred $m$ across stocks.

A large positive [s-score](#def-s1-residual-and-principal-component-stat-arb-sscore) says the stock is rich relative to its factors, and a large negative one that it is cheap. The trading rule is a band. Buy (the stock, and sell its betas in the factors) when $s < -1.25$; sell when $s > 1.25$; close a long when $s > -0.50$ and a short when $s < 0.75$. The cutoffs were chosen on 2000–2004. Every position is the same fraction of capital (the paper’s “bang-bang” rule), and each trade costs five basis points.

## 3.3 Entry, exit and the mean-reversion speed filter

**Definition 3.3 (Mean-reversion speed filter).**

A *mean-reversion speed filter* trades a stock’s [s-score](#def-s1-residual-and-principal-component-stat-arb-sscore) only when its fitted mean-reversion speed is high enough that the reversion is expected well within the estimation window; Avellaneda and Lee required $\kappa > 252/30 = 8.4$, a mean-reversion time under half the sixty-day window ($0 < b < 0.9672$).

To test the method, the synthetic market needs something for it to find. `MarketConfig.ou_share` (new in this chapter, off by default so every earlier panel is unchanged) gives each stock’s specific return a mean-reverting level that takes a share of its specific variance, with mean-reversion times spread lognormally around twenty days. The planted level’s expected pull is recorded, like every planted effect, so the [s-score](#def-s1-residual-and-principal-component-stat-arb-sscore) can be scored against it: its correlation with the pull is $-0.128$ at a share of 0.3 (the sign is right: a high score means a fall to come).

| mean-reverting share | 0 | 0.15 | 0.3 |
| --- | --- | --- | --- |
| factors | sector | PCA | sector | PCA | sector | PCA |
| Sharpe ratio before costs | 1.04 | 1.30 | 1.92 | 2.56 | 3.15 | 4.34 |
| Sharpe ratio after costs | 0.07 | $-0.06$ | 0.93 | 1.22 | 2.15 | 3.00 |
| net return a year | 0.9% | $-0.5\%$ | 11.5% | 9.0% | 26.0% | 22.4% |
| costs a year | 12.0% | 9.8% | 12.3% | 9.8% | 12.2% | 10.0% |
| volatility | 12.3% | 7.2% | 12.4% | 7.3% | 12.1% | 7.5% |
| open positions | 347 | 343 | 344 | 340 | 336 | 338 |
| turnover a day | 37% | 39% | 37% | 39% | 37% | 40% |

Three things stand out. Without a mean-reverting component the method earns its costs and nothing more; the gross Sharpe ratio of about 1 comes from the one-day reversal the market plants anyway, and five basis points a trade eat it. With one, the method finds it, and the principal components beat the sector indices because they remove the style factors too: the PCA book’s volatility is 7.5% against 12.1%. And the book is large: about 340 positions of 1% of capital each, a gross exposure of 3.4, close to the paper’s “2+2”.

The speed filter barely filters. At Avellaneda and Lee’s threshold it passes 99.3% of the fits, and it passes 99.1% when there is no mean reversion at all: an AR(1) fitted on sixty points of a path pinned to zero at its end is strongly biased toward fast reversion, so even a random walk looks fast. Stricter thresholds reject more names but do not help ([Figure 3.1](#fig-s1-residual-and-principal-component-stat-arb-sens)): requiring a mean-reversion time under fifteen days keeps 87.9% of the fits and nets 2.03; under ten days, 69.2% and 1.81; under five, 26.8% and 0.74. Without the filter the sector book nets 2.16 and the PCA book 3.04, against 2.15 and 3.00 with it. The entry threshold matters as little, which is good news for a published rule: from 0.75 to 1.5 the net Sharpe ratio stays between 2.20 and 2.30, and it falls to 1.87 at 2.0, where the book holds 101 positions instead of 336.

![The sector-index s-score book with a planted mean-reverting share of 0.3. Left: Sharpe ratio against the entry threshold. Right: net Sharpe ratio against the speed filter’s largest allowed mean-reversion time (30 days is Avellaneda and Lee’s); labels give the percentage of fits kept. Data: s1_residarb.run.](https://one-course.com/images/onecourse/chapters/quant-8/s1-residual-and-principal-component-stat-arb/fig-9c87b2aed72c.svg)

***Figure 3.1.** The sector-index [s-score](#def-s1-residual-and-principal-component-stat-arb-sscore) book with a planted mean-reverting share of 0.3. Left: Sharpe ratio against the entry threshold. Right: net Sharpe ratio against the speed filter’s largest allowed mean-reversion time (30 days is Avellaneda and Lee’s); labels give the percentage of fits kept. Data: `s1_residarb.run`.*

![One synthetic stock’s s-score (against its sector index, not centred) over a year, with the entry thresholds dashed; the step line is the position (up long, down short, drawn at ± 2.6). Data: s1_residarb.path.](https://one-course.com/images/onecourse/chapters/quant-8/s1-residual-and-principal-component-stat-arb/fig-12c6fef178dc.svg)

***Figure 3.2.** One synthetic stock’s [s-score](#def-s1-residual-and-principal-component-stat-arb-sscore) (against its sector index, not centred) over a year, with the entry thresholds dashed; the step line is the position (up long, down short, drawn at $\pm 2.6$). Data: `s1_residarb.path`.*

## 3.4 Eigenportfolios and their stability

A factor that changes every month is a factor the book cannot hedge cheaply. Re-estimated every twenty-one days on a sliding year, the first eigenvector is almost the same each time (the absolute cosine between successive estimates averages 0.999: it is the market). The fifteen-dimensional subspace moves more: its average overlap with the previous month’s (the mean squared cosine of the principal angles) is 0.89, and the least stable direction’s cosine averages 0.63 ([Figure 3.3](#fig-s1-residual-and-principal-component-stat-arb-stab)). The later components mix industries and styles differently from one estimate to the next, and a hedge in them is rebuilt each month. The sector-index version pays for its stability with a worse fit; the PCA version pays for its fit with hedges that turn over, which this chapter’s cost model charges at five basis points per unit of each [eigenportfolio](#def-s1-residual-and-principal-component-stat-arb-eigenportfolio), a generous simplification for a basket of a thousand names.

![Stability of the fifteen leading eigenportfolios, re-estimated every 21 days on the last 252: the overlap with the previous estimate of the subspace they span (mean squared cosine of the principal angles) and the cosine of the largest principal angle. Data: s1_residarb.stability.](https://one-course.com/images/onecourse/chapters/quant-8/s1-residual-and-principal-component-stat-arb/fig-5896755b8402.svg)

***Figure 3.3.** Stability of the fifteen leading [eigenportfolios](#def-s1-residual-and-principal-component-stat-arb-eigenportfolio), re-estimated every 21 days on the last 252: the overlap with the previous estimate of the subspace they span (mean squared cosine of the principal angles) and the cosine of the largest principal angle. Data: `s1_residarb.stability`.*

## 3.5 Where the returns went

Avellaneda and Lee’s own numbers tell the story of the strategy’s decline: the PCA version’s Sharpe ratio after costs averaged 1.44 over 1997–2007, much stronger before 2003 and 0.9 from 2003 to 2007; the ETF version 1.1 over the whole period, with a similar degradation after 2002. Two further findings in the same paper point to what kept working. Using trading time, dividing each day’s return by that day’s volume relative to its ten-day average, lifted the ETF version to 1.51 in 2003–2007. And the strategies’ behaviour in the summer of 2007 fitted Khandani and Lo’s account of books like these being unwound together (Book 7, chapter 28).

The synthetic market cannot say why real residuals reverted less after 2002, only what the method needs: a mean-reverting component large enough to pay five basis points a trade on a third of the book a day. At a share of 0.15 of specific variance the sector book nets 0.93; at 0.3, 2.15; at zero, nothing. Its own volume effect goes the other way: the synthetic market’s volume rises with the size of the day’s specific shock, so trading time shrinks exactly the moves that reverse, and the volume-adjusted sector book nets 1.78 against 2.15. Whether volume helps depends on what volume measures in a given market; the paper found it helped in the United States in 2003–2007.

## 3.6 Strategy files

**Strategy file 3.1 — ETF-residual s-score strategy.**

**Who pays you, and why.** Flows that push a stock away from its sector without news: index and fund trades, hurried sellers; the book absorbs them and waits for the stock to rejoin its sector.

**Instruments and venues.** Liquid stocks and their sector funds, long and short; the hedges in the funds.

**Signal.** The [s-score](#def-s1-residual-and-principal-component-stat-arb-sscore) of the 60-day cumulative residual against the sector fund; open at $|s| > 1.25$, close at $-0.50$ (longs) and $0.75$ (shorts).

**Sizing and execution.** Equal positions of a fixed fraction of capital; each hedged with its beta in the fund; hedges netted across the book.

**Costs.** Five basis points a trade in the paper; a third of the book traded a day here.

**How it dies.** Residuals stop reverting (the published record after 2002); crowding and simultaneous unwinds (August 2007); costs above the edge.

**Horizon, capacity, infrastructure.** Days to weeks; hundreds of positions; a daily regression and fit per stock, sector-fund hedging.

**Backtest honestly.** Point-in-time universes and sector membership; costs on both legs; the cutoffs chosen on an earlier period than the one reported.

**Sources.** Avellaneda and Lee (2010): Sharpe ratio 1.1 after costs over 1997–2007, degrading after 2002; this chapter’s simulation.

**Strategy file 3.2 — Principal-component residual strategy.**

**Who pays you, and why.** As the ETF version; removing more common structure leaves a cleaner residual.

**Instruments and venues.** Liquid stocks; hedges in the [eigenportfolios](#def-s1-residual-and-principal-component-stat-arb-eigenportfolio), which are baskets of the same stocks.

**Signal.** The [s-score](#def-s1-residual-and-principal-component-stat-arb-sscore) against 15 [eigenportfolios](#def-s1-residual-and-principal-component-stat-arb-eigenportfolio) of the last year’s correlation matrix (or enough to explain 55% of the variance).

**Sizing and execution.** As the ETF version; hedges netted into the stock book, since the [eigenportfolios](#def-s1-residual-and-principal-component-stat-arb-eigenportfolio) are stocks.

**Costs.** Stock trades plus the drift of the [eigenportfolios](#def-s1-residual-and-principal-component-stat-arb-eigenportfolio) between re-estimations.

**How it dies.** As the ETF version; [eigenportfolios](#def-s1-residual-and-principal-component-stat-arb-eigenportfolio) that rotate faster than the book can hedge.

**Horizon, capacity, infrastructure.** Days to weeks; a correlation-matrix estimate and its eigenvectors, refreshed regularly.

**Backtest honestly.** Eigenvectors estimated only on data before each trade; the number of components fixed in advance.

**Sources.** Avellaneda and Lee (2010): Sharpe ratio 1.44 after costs over 1997–2007, 0.9 in 2003–2007.

**Strategy file 3.3 — Industry-neutral residual book.**

**Who pays you, and why.** As above, with the [s-scores](#def-s1-residual-and-principal-component-stat-arb-sscore) turned into a portfolio rather than a set of trades.

**Instruments and venues.** Liquid stocks only; no fund hedges.

**Signal.** Minus the [s-score](#def-s1-residual-and-principal-component-stat-arb-sscore), as a forecast in return units (chapter 2).

**Sizing and execution.** An optimiser neutral to industries and styles with costs inside (chapter 2), instead of equal bang-bang positions.

**Costs.** Lower than the bang-bang rule: the optimiser trades only where the forecast pays.

**How it dies.** As above.

**Horizon, capacity, infrastructure.** Days to weeks; a risk model and an optimiser.

**Backtest honestly.** The same fits and thresholds as the bang-bang version, so that the comparison isolates the portfolio construction.

**Sources.** Avellaneda and Lee (2010) note that the all-or-nothing rule outperformed continuous adjustments in their tests, which they attribute to model misspecification; the comparison is the reader’s to run (tutorial).

**Strategy file 3.4 — Volume-conditioned residual strategy.**

**Who pays you, and why.** As the ETF version, weighting moves by how much trading they took.

**Instruments and venues.** As the ETF version.

**Signal.** The [s-score](#def-s1-residual-and-principal-component-stat-arb-sscore) of residuals of returns put in trading time: each day’s return multiplied by its ten-day average volume over its own volume.

**Sizing and execution.** As the ETF version.

**Costs.** As the ETF version.

**How it dies.** When volume stops distinguishing liquidity moves from news moves; on the synthetic market, where volume rises with every large shock, it hurts.

**Horizon, capacity, infrastructure.** Days to weeks; daily volume by stock.

**Backtest honestly.** Volume known at the close used; the adjustment compared with the plain version on the same period.

**Sources.** Avellaneda and Lee (2010): 1.51 in 2003–2007 for the ETF version with volume information.

## 3.7 Tutorial: the public method

**Goal.** Run the published rule with sector indices and with principal components, with and without a mean-reverting component, and measure what its parameters do. **End state:** the table and [Figure 3.1](#fig-s1-residual-and-principal-component-stat-arb-sens).

1. **The fit and the rule**: the cumulative residual’s AR(1), the [s-score](#def-s1-residual-and-principal-component-stat-arb-sscore), and the band. `X = np.cumsum(np.asarray(resid, float ), axis=0 ) x, y = X[:-1 ], X[1 :] xm, ym = x.mean(axis=0 ), y.mean(axis=0 ) b = ((x - xm) * (y - ym)).sum(axis=0 ) / ((x - xm) ** 2 ).sum(axis=0 ) a = ym - b * xm var = ((y - a - b * x) ** 2 ).sum(axis=0 ) / (len (x) - 2 ) ok = (b > 0 ) & (b < 1 ) bb = np.where(ok, b, np.nan) return {" a " : a, " b " : b, " var " : var, " kappa " : -np.log(bb) * YEAR, " m " : a / (1 - bb), " sigma_eq " : np.sqrt(var / (1 - bb * bb))} def s_score (fit, center: bool = True ): m = fit[" m " ] if center: m = m - np.nanmean(m) return -m / fit[" sigma_eq " ] def step (pos, s, ok, s_open: float = 1.25 , s_close_long: float = 0.5 , s_close_short: float = 0.75 ): """Buy to open below -s_open, sell to open above s_open; close longs above -s_close_long and shorts below s_close_short; names that are not ok (no fit, or too slow) are closed.""" pos, s, ok = np.asarray(pos, int ).copy(), np.asarray(s, float ), np.asarray(ok, bool ) good = ok & np.isfinite(s) pos[(pos == 1 ) & (~good | (s > -s_close_long))] = 0 pos[(pos == -1 ) & (~good | (s < s_close_short))] = 0 flat = pos == 0 pos[flat & good & (s < -s_open)] = 1 pos[flat & good & (s > s_open)] = -1 return pos` **Listing 3.1.** The OU fit by AR(1), the s-score and the trading rule. code/firm/residarb/firm_residarb.py
2. **Every day**: regress, fit, score, trade, hedge. `if method == " etf " : beta, e = regress(Rw, ind[t - W + 1 :t + 1 ][:, P.industry[idx]], per_name=True ) else : Fw = np.nan_to_num(R[t - W + 1 :t + 1 ][:, uni]) @ Q.T beta, e = regress(Rw, Fw) fit = ou_fit(e) s = s_score(fit) ok = speed_ok(fit[" kappa " ], min_kappa) new = np.zeros(M, int ) new[idx] = step(pos[idx], s, ok, s_open) pos = new w = LAM * pos r1 = np.nan_to_num(R[t + 1 ]) if method == " etf " : h = np.bincount(P.industry[idx], weights=w[idx] * beta, minlength=ind.shape[1 ]) f1 = ind[t + 1 ] else : h = beta @ w[idx] f1 = np.nan_to_num(R[t + 1 ][uni]) @ Q.T gross.append(float (w @ r1 - h @ f1))` **Listing 3.2.** The daily step of the s-score book. code/strategies-1/03-residual-and-principal-component-stat-arb/python/s1_residarb.py
3. **Run** `run(share, method, s_open, min_kappa, volume)` over the grid, `stability()` and `fig_residarb.py` .

**What to change next.** Replace the bang-bang rule with chapter 2’s cost-aware optimiser (the industry-neutral strategy file); estimate the regression on 90 days and the process on 60, as the paper suggests; make the mean-reversion times shorter and see when the speed filter starts to matter.

## 3.8 Build: residual stat arb

**Purpose.** Residuals, OU fits, [s-scores](#def-s1-residual-and-principal-component-stat-arb-sscore) and the trading rule of residual stat arb, for any set of factors.

**Interface.** `regress(R, F, per_name)`, `eigenportfolios(R, k)`, `ou_fit(resid)`, `s_score(fit, center)`, `step(pos, s, ok, s_open, s_close_long, s_close_short)`, `speed_ok(kappa, min_kappa)`, `volume_adjust(R, V, avg)`.

**Rules.** Only data before each trade; fits rejected where $b \notin (0, 1)$; parameters fixed in advance and reported with the results.

**Acceptance tests.** `code/firm/residarb/tests/`: regressions recover known betas; the AR(1) fit recovers a known coefficient within its small-sample bias; the [s-score](#def-s1-residual-and-principal-component-stat-arb-sscore) and the band by hand; [eigenportfolios](#def-s1-residual-and-principal-component-stat-arb-eigenportfolio) of a one-factor market.

**Stretch.** A drift term in the [s-score](#def-s1-residual-and-principal-component-stat-arb-sscore) (the paper’s “modified [s-score](#def-s1-residual-and-principal-component-stat-arb-sscore)”); bias-corrected AR(1) estimates; hedges in real ETFs with their own costs.

Sources and further reading

- M. Avellaneda and J.-H. Lee, “Statistical arbitrage in the US equities market”, *Quantitative Finance* 10(7), 2010 (working paper version of June 2009 on the first author’s page).
- A. E. Khandani and A. W. Lo, “What happened to the quants in August 2007?”, *Journal of Financial Markets* 14(1), 2011.
- A. Pole, *Statistical Arbitrage* , Wiley.

## 3.9 Exercises

**Exercise 3.1 ★.**

A cumulative residual’s AR(1) coefficient is $b = 0.9$ per day. Compute $\kappa$ and the mean-reversion time in days.

**Solution of Exercise 3.1.**

$\kappa = -252 \ln 0.9 = 26.6$ a year; the mean-reversion time is $252/\kappa = 9.5$ days.

**Exercise 3.2 ★.**

With $b = 0.9$, $a = 0.002$ and $\operatorname{Var}\zeta = 10^{-4}$, compute $m$, $\sigma_{\mathrm{eq}}$ and the (uncentred) [s-score](#def-s1-residual-and-principal-component-stat-arb-sscore). Does the rule trade?

**Solution of Exercise 3.2.**

$m = 0.002/0.1 = 0.02$; $\sigma_{\mathrm{eq}} = \sqrt{10^{-4}/(1 - 0.81)} = 2.29\%$; $s = -0.02/0.0229 = -0.87$. No: $|s| < 1.25$.

**Exercise 3.3 ★.**

Show that $\kappa > 252/30$ is the same condition as $b < 0.9672$.

**Solution of Exercise 3.3.**

$\kappa = -252 \ln b > 252/30$ is $\ln b < -1/30$, that is $b < e^{-1/30} = 0.9672$.

**Exercise 3.4 ★★.**

The sector book turns over 37% of capital a day. What does five basis points per unit traded cost it a year in stocks alone, and why is the chapter’s cost figure higher?

**Solution of Exercise 3.4.**

The weight traded is twice the turnover: $2 \times 0.37 \times 0.0005 \times 252 = 9.3\%$ a year. The chapter’s 12.2% also charges the hedges, whose sector exposures change as positions open and close.

**Exercise 3.5 ★★.**

Why does the speed filter pass 99% of fits even on a market with no mean reversion?

**Solution of Exercise 3.5.**

The residuals of a regression with an intercept sum to zero, so the cumulative residual starts and ends near zero. An AR(1) fitted to such a pinned path of sixty points has a coefficient biased well below one, so nearly every path looks mean-reverting with a time far under thirty days, including a random walk. The filter tests the estimate, and the estimate is biased.

**Exercise 3.6 ★★.**

Why does the PCA book have a lower volatility than the sector book at the same number of positions?

**Solution of Exercise 3.6.**

The sector index removes the market and the stock’s industry but not the styles; fifteen [eigenportfolios](#def-s1-residual-and-principal-component-stat-arb-eigenportfolio) span the market, the industries and the styles, so the PCA residuals and the hedged positions carry less common risk: 7.5% against 12.1% of volatility with about the same number of positions.

**Exercise 3.7 ★★★.**

*Coding.* Run `run(0.15, ’pca’)` and `run(0.0, ’pca’)`. Report the net Sharpe ratios and explain what the second one measures.

**Solution of Exercise 3.7.**

1.22 at a share of 0.15 and $-0.06$ at zero. The second measures the method on a market with no mean-reverting level: its gross Sharpe ratio of 1.30 is the planted one-day reversal leaking into the [s-score](#def-s1-residual-and-principal-component-stat-arb-sscore), and costs of 9.8% a year remove it. It is the method’s floor.

**Exercise 3.8 ★★★.**

*Find the flaw.* “The method is published with its parameters and a Sharpe ratio of 1.44 after costs; we will run it as published.”

**Solution of Exercise 3.8.**

The Sharpe ratio of 1.44 is an average over 1997–2007 that was much stronger before 2003 and 0.9 afterwards, and the cutoffs were chosen on part of the same sample. The published parameters are not the risk; whether today’s residuals revert is. Run it first on recent data, with the firm’s costs, and measure the mean-reverting share of residual variance directly.

## 3.10 Problem: The Public Method

**Problem 3.1.**

Weekend problem — a published strategy, re-run

Avellaneda and Lee’s rule on `firm.synthmkt`, with and without a planted mean-reverting component.

**Part I — The method.**

1. Describe the two ways of computing residuals.
2. Define an [eigenportfolio](#def-s1-residual-and-principal-component-stat-arb-eigenportfolio) and say how many the paper used.
3. Fit the OU process to the cumulative residual: what are $\kappa$ , $m$ and $\sigma_{\mathrm{eq}}$ ?
4. Define the [s-score](#def-s1-residual-and-principal-component-stat-arb-sscore) and give the trading rule’s thresholds.

**Part II — The synthetic test.**

5. What does `ou_share` plant, and how is the [s-score](#def-s1-residual-and-principal-component-stat-arb-sscore) scored against it?
6. Give the net Sharpe ratios of both versions at shares of 0, 0.15 and 0.3.
7. Why does the PCA version do better here?
8. What does the method earn without mean reversion, and where does its gross return come from?

**Part III — The parameters.**

9. What share of fits passes the paper’s speed filter, with and without mean reversion, and why?
10. What do stricter filters do?
11. How sensitive is the result to the entry threshold?
12. What does the volume adjustment do here, and why?

**Part IV — The verdict.**

13. State the *named result* : the Sharpe ratio of the [s-score](#def-s1-residual-and-principal-component-stat-arb-sscore) strategy with and without the speed filter, and its sensitivity to the entry threshold.
14. How stable are the [eigenportfolios](#def-s1-residual-and-principal-component-stat-arb-eigenportfolio) ?
15. What did the paper report after 2002?
16. What does August 2007 add?
17. What does the method need from a market to work?
18. How would you tell, on real data, whether residuals revert?
19. Which parameter would you worry about most?
20. In one sentence: what is residual stat arb?

**Solution of Problem 3.1.**

1. Time-series regressions of each stock’s last 60 returns on its sector fund (“ETF”), or on the returns of leading [eigenportfolios](#def-s1-residual-and-principal-component-stat-arb-eigenportfolio) (“PCA”).
2. $Q^{(j)}_i = v^{(j)}_i/\sigma_i$ from the correlation matrix’s eigenvectors; 15, or as many as explain 55% of the variance.
3. $\kappa = -252 \ln b$ , $m = a/(1-b)$ , $\sigma_{\mathrm{eq}} = \sqrt{\operatorname{Var}\zeta/(1-b^2)}$ from $X_{n+1} = a + bX_n + \zeta$ .
4. $s = (X - m)/\sigma_{\mathrm{eq}} = -m/\sigma_{\mathrm{eq}}$ ; open at $\pm 1.25$ , close longs at $-0.50$ and shorts at $0.75$ .
5. A mean-reverting level in each stock’s specific return taking a share of its variance, mean-reversion times around twenty days; the [s-score](#def-s1-residual-and-principal-component-stat-arb-sscore) correlates $-0.128$ with the level’s expected pull at a share of 0.3.
6. Sector: 0.07, 0.93, 2.15. PCA: $-0.06$ , 1.22, 3.00.
7. The [eigenportfolios](#def-s1-residual-and-principal-component-stat-arb-eigenportfolio) remove the style factors as well; the PCA book’s volatility is 7.5% against 12.1%.
8. Nothing after costs (0.07 and $-0.06$ ); its gross Sharpe ratio of 1.0–1.3 comes from the planted one-day reversal.
9. 99.3% with, 99.1% without: the AR(1) on a pinned 60-point path is biased toward fast reversion.
10. They reject more names and lower the net Sharpe ratio: 2.03, 1.81 and 0.74 for times under 15, 10 and 5 days.
11. Hardly: 2.20 to 2.30 between thresholds of 0.75 and 1.5, 1.87 at 2.0.
12. It lowers the net Sharpe ratio to 1.78: synthetic volume rises with the day’s specific shock, so trading time shrinks the moves that reverse.
13. **Named result.** With a mean-reverting share of 0.3, the sector book nets 2.15 with the paper’s speed filter and 2.16 without (PCA: 3.00 and 3.04); between entry thresholds of 0.75 and 1.5 the net ratio stays within 2.20–2.30, and at 2.0 it is 1.87.
14. The first eigenvector is nearly fixed (cosine 0.999 between monthly estimates); the fifteen-dimensional subspace overlaps 0.89 on average, and its least stable direction only 0.63.
15. Sharpe ratios of 0.9 for PCA in 2003–2007, against 1.44 over 1997–2007; a similar degradation for the ETF version after 2002.
16. Books of this kind were unwound together, consistent with Khandani and Lo’s account.
17. Residuals with a mean-reverting component large enough to pay the costs of trading a third of the book a day.
18. Measure the autocorrelation of residual returns at horizons of days to weeks on data the fits did not use, and the out-of-sample correlation of [s-scores](#def-s1-residual-and-principal-component-stat-arb-sscore) with the next weeks’ residual returns.
19. The cost model, and then the estimation window: the entry threshold and the speed filter hardly matter.
20. Trading stocks back toward where their factors say they should be, when they have strayed far enough to pay for the round trip.

## 3.11 Interview questions

**Interview question 3.1 ★ researcher, trader.**

What is an [s-score](#def-s1-residual-and-principal-component-stat-arb-sscore), and how would you trade it?

**Solution of Interview question 3.1.**

The distance of a stock’s cumulative residual from its mean-reversion level, in equilibrium standard deviations, from an OU fit on a trailing window. Sell (and buy the hedge) when it is high, buy when it is low, close when it comes back toward zero; size equally or through an optimiser.

**Interview question 3.2 ★★ researcher.**

PCA factors or sector ETFs for residual stat arb: what are the trade-offs?

**Solution of Interview question 3.2.**

PCA fits the market’s actual common structure, including styles, and needs no fund data, but its factors rotate and are hard to interpret or hedge cheaply. ETFs are tradable, stable and cheap to hedge, but leave style exposures in the residuals and depend on fund composition.

**Interview question 3.3 ★★ researcher.**

You fit an AR(1) to 60 observations and get $b = 0.85$. How much do you trust the mean-reversion time?

**Solution of Interview question 3.3.**

Not much. The AR(1) slope is biased downward in small samples, more so on a path pinned by the regression’s intercept, and its standard error is large: a mean-reversion time of 6 days ($b = 0.85$) is compatible with much slower reversion or none. Check with a bias-corrected estimate, a longer window, or the fit’s out-of-sample record.

**Interview question 3.4 ★★ risk.**

A residual [stat-arb book](https://one-course.com/books/quant/8/en/chapter/1-anatomy-of-a-stat-arb-book#def-s1-anatomy-of-a-stat-arb-book-statarb) loses 5% in three days with no news. What do you check?

**Solution of Interview question 3.4.**

Factor exposures left in the book (styles not in the hedge), a crowded unwind (other books with the same positions, as in August 2007), news in a few large names, and the hedges: did the factors used for hedging change? Then decide whether to cut risk or provide liquidity.

**Interview question 3.5 ★★ developer.**

Compute [s-scores](#def-s1-residual-and-principal-component-stat-arb-sscore) for 3 000 stocks every day in under a second: how?

**Solution of Interview question 3.5.**

Vectorise: one least-squares solve with all stocks as right-hand sides (the same factor matrix), cumulative sums and the AR(1) moments as column operations; update the correlation matrix and eigenvectors weekly or monthly, not daily. That is a few matrix products per day.

**Interview question 3.6 ★★★ researcher.**

Derive the equilibrium standard deviation of an OU process sampled daily as an AR(1), and show why $X_{60} = 0$ in the paper’s [s-score](#def-s1-residual-and-principal-component-stat-arb-sscore).

**Solution of Interview question 3.6.**

The AR(1) $X_{n+1} = a + bX_n + \zeta$ has stationary variance $\operatorname{Var}\zeta/(1 - b^2)$, the square of $\sigma_{\mathrm{eq}}$. With an intercept in the regression the residuals sum to zero, so $X_{60} = \sum_{n=1}^{60} \epsilon_n = 0$ and $s = (0 - m)/\sigma_{\mathrm{eq}}$.
