---
title: "Linear Time Series"
book: "Quantitative Methods"
subject: quant
language: en
chapter: 17
exercises: 8
source: https://one-course.com/books/quant/4/en/chapter/17-linear-time-series
---

# Chapter 17 — Linear Time Series

The spread between the ten-year and the two-year Treasury yields, the 2s10s that a steepener trades, has a daily autocorrelation of 0.9987 over fifty years. A regression of tomorrow’s spread on today’s gives a slope of 0.99875 with a [standard error](https://one-course.com/books/quant/4/en/chapter/11-estimation#def-qm-estimation-estimator) of 0.00045: a $t$-statistic of $-2.79$ against a slope of one, a one-sided $p$-value of 0.003 by the normal table. The spread seems to mean-revert, with a [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) of 552 trading days. But if the spread had a [unit root](#def-qm-linear-time-series-unitroot), the normal table would be the wrong one: the right one puts the 5% critical value at $-2.86$, and the evidence falls short. This chapter is about the models that describe such series and the tests that tell them apart: stationarity and the [autocorrelation function](#def-qm-linear-time-series-acf), autoregressive and moving-average models, the spectral view, [unit roots](#def-qm-linear-time-series-unitroot) and the regressions they make spurious, and [long memory](#def-qm-linear-time-series-longmem).

## 17.1 Stationarity and autocorrelation

**Definition 17.1 (Stationary process, weak stationarity, white noise).**

A process $(X_t)$ is a *stationary process* (strictly) if the joint law of $(X_{t_1 + h}, \dots, X_{t_k + h})$ does not depend on $h$. It has *weak stationarity* if its mean is constant and $\Cov(X_t, X_{t+h})$ depends only on $h$. *White noise* is a weakly stationary sequence with mean zero and no autocorrelation at any nonzero lag.

**Definition 17.2 (Autocovariance, autocorrelation and partial autocorrelation functions).**

For a weakly [stationary process](#def-qm-linear-time-series-stationary), the *autocovariance function* is $\gamma(h) = \Cov(X_t, X_{t+h})$, and the *autocorrelation function* is $\rho(h) = \gamma(h)/\gamma(0)$. The *partial autocorrelation function* at lag $h$ is the last coefficient $\phi_{hh}$ of the best linear predictor of $X_{t+h}$ from $X_{t+h-1},
\dots, X_t$: the correlation at lag $h$ once the intermediate lags are accounted for.

Under [white noise](#def-qm-linear-time-series-stationary) the sample autocorrelations are approximately independent $\mathcal N(0, 1/n)$, whence the $\pm 2/\sqrt n$ bands of every correlogram. The decomposition that justifies the models of the next section is Wold’s (1938): every weakly stationary, purely nondeterministic process is an infinite moving average $X_t = \mu + \sum_{j \ge 0}\psi_j\varepsilon_{t-j}$ of its one-step prediction errors $\varepsilon_t$, a [white noise](#def-qm-linear-time-series-stationary), with $\psi_0 = 1$ and $\sum_j\psi_j^2 < \infty$.

The 2s10s spread’s correlogram ([Figure 17.1](#fig-qm-linear-time-series-acf)) is the signature of a nearly nonstationary level: autocorrelations that start at 0.9987 and fall almost linearly over a year of lags. Its daily changes are nearly [white noise](#def-qm-linear-time-series-stationary), with a standard deviation of 4.6 basis points and autocorrelations of 0.015, 0.015 and $-0.027$ at the first three lags, within the band $\pm 0.018$ for the first two.

![Sample autocorrelations of the 2s10s Treasury spread, 1 June 1976 to 22 September 2026 (12 574 days). Left: the level, lags 1 to 250, from 0.9987 down. Right: the daily change, lags 1 to 40, with the ± 2/√ n band. Data: FRED series DGS10 and DGS2 (Board of Governors of the Federal Reserve System, H.15).](https://one-course.com/images/onecourse/chapters/quant-4/qm-linear-time-series/fig-6e0cd8464288.svg)

***Figure 17.1.** Sample autocorrelations of the 2s10s Treasury spread, 1 June 1976 to 22 September 2026 (12 574 days). Left: the level, lags 1 to 250, from 0.9987 down. Right: the daily change, lags 1 to 40, with the $\pm 2/\sqrt n$ band. Data: FRED series DGS10 and DGS2 (Board of Governors of the Federal Reserve System, H.15).*

## 17.2 Autoregressive and moving-average models

**Definition 17.3 (Autoregressive, moving-average and ARMA processes).**

With $(\varepsilon_t)$ [white noise](#def-qm-linear-time-series-stationary) of variance $\sigma^2$: an *autoregressive process* of order $p$, AR($p$), satisfies $X_t = c + \sum_{i=1}^p\phi_iX_{t-i} + \varepsilon_t$; a *moving-average process* of order $q$, MA($q$), is $X_t = \mu +
\varepsilon_t + \sum_{j=1}^q\vartheta_j\varepsilon_{t-j}$; an *ARMA process* combines them, $\phi(L)X_t = c + \vartheta(L)\varepsilon_t$ with $\phi(z) = 1 - \sum_i\phi_iz^i$, $\vartheta(z) = 1 + \sum_j\vartheta_jz^j$ and $L$ the lag operator.

**Proposition 17.4 (Stationarity and invertibility of ARMA).**

The ARMA equation has a unique weakly stationary causal solution $X_t = \mu + \sum_{j \ge 0}\psi_j\varepsilon_{t-j}$ if and only if every root of $\phi(z)$ lies outside the unit circle; it is invertible, $\varepsilon_t = \sum_{j \ge 0}\pi_jX_{t-j}$ up to a constant, if and only if every root of $\vartheta(z)$ does, the polynomials having no common root.

**Proof.** Factor $\phi(z) = \prod_i(1 - z/z_i)$. For $|z_i| > 1$, $(1 - L/z_i)^{-1} = \sum_kz_i^{-k}L^k$ is an absolutely summable filter, so $\psi(z) = \vartheta(z)/\phi(z)$ has summable coefficients and defines the solution; if some $|z_i| \le 1$ no summable causal inverse exists (for $|z_i| = 1$, none at all). The MA side is the same argument applied to $\vartheta$. ∎

For the AR(1) $X_t = c + \phi X_{t-1} + \varepsilon_t$, the condition is $|\phi| < 1$, the autocorrelations are $\phi^h$, and a deviation from the mean decays by half in $\ln(1/2)/\ln\phi$ periods, the [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) of chapter 4’s [Ornstein–Uhlenbeck process](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) sampled daily. An AR($p$) has a [partial autocorrelation function](#def-qm-linear-time-series-acf) that cuts off after lag $p$; an MA($q$) has an [autocorrelation function](#def-qm-linear-time-series-acf) that cuts off after lag $q$. The orders are chosen by an [information criterion](#def-qm-linear-time-series-ic).

**Definition 17.5 (Information criterion).**

An *information criterion* ranks models by fit penalised for size: Akaike’s $\mathrm{AIC} = -2\ell + 2k$ and Schwarz’s $\mathrm{BIC} = -2\ell + k\ln n$, for maximised log-likelihood $\ell$ and $k$ parameters; the smallest wins.

AIC aims at prediction and tends to choose large models; BIC is consistent for the true order when there is one. On the 2s10s daily changes AIC chooses an AR(10) whose coefficients are all below 0.035 in absolute value: statistically detectable with 12 574 days, economically nothing.

## 17.3 The spectral view

**Definition 17.6 (Spectral density, periodogram).**

The *spectral density* of a weakly [stationary process](#def-qm-linear-time-series-stationary) with summable autocovariances is $f(\omega) = \frac1{2\pi}\sum_h\gamma(h)e^{-ih\omega}$, $\omega \in [-\pi, \pi]$: the decomposition of its variance by frequency, $\gamma(0) = \int_{-\pi}^\pi f$. The *periodogram* $I(\omega_j) =
\frac1{2\pi n}\bigl\lvert\sum_tX_te^{-it\omega_j}\bigr\rvert^2$ at the Fourier frequencies $\omega_j = 2\pi j/n$ is its sample counterpart.

[White noise](#def-qm-linear-time-series-stationary) has a flat spectrum, $\sigma^2/2\pi$; an AR(1) with $\phi$ near one concentrates its variance at low frequencies, $f(\omega) =
\sigma^2/(2\pi|1 - \phi e^{-i\omega}|^2)$. The [periodogram](#def-qm-linear-time-series-spectral) is asymptotically unbiased but not consistent: each ordinate is approximately $f(\omega_j)$ times an exponential variable, however long the sample, so it must be smoothed or averaged. Its low-frequency end is where persistence shows, and it is the basis of the long-memory estimate of the last section. $f(0) = \frac1{2\pi}\sum_h\gamma(h)$ is, up to $2\pi$, the long-run variance of chapter 11.

## 17.4 Unit roots

**Definition 17.7 (Unit root, spurious regression).**

A process has a *unit root* if its autoregressive polynomial has a root at $z = 1$: its first difference is stationary but the level is not, like a random walk. A *spurious regression* is a regression between independent processes with unit roots whose usual $t$-statistics and $R^2$ indicate a relation that does not exist.

Granger and Newbold (1974) showed the problem by simulation, and it is easy to repeat: regressing one random walk of 500 steps on another, independent one, the slope’s $t$-statistic exceeds 1.96 in absolute value in 89% of 5 000 trials, with a median $R^2$ of 0.17. The residual inherits the [unit root](#def-qm-linear-time-series-unitroot), the observations are not independent, and every formula that assumed they were is wrong. The same failure hits the test of $\phi = 1$ itself.

**Definition 17.8 (Dickey–Fuller test).**

The *Dickey–Fuller test* of a [unit root](#def-qm-linear-time-series-unitroot) regresses $\Delta X_t$ on $X_{t-1}$ (with a constant, and a trend if wanted) and compares the $t$-statistic $\tau$ of the coefficient on $X_{t-1}$ with the Dickey–Fuller distribution rather than the normal; the augmented version adds lagged differences $\Delta X_{t-k}$ to absorb short-run dependence (Said and Dickey, 1984).

Under the [unit root](#def-qm-linear-time-series-unitroot) and with a constant, $\tau$ converges to $\bigl(\frac12(W(1)^2 - 1) - W(1)\int_0^1W\bigr)/\bigl(\int_0^1W^2 - (\int_0^1W)^2\bigr)^{1/2}$ for a standard [Brownian motion](https://one-course.com/books/quant/4/en/chapter/2-brownian-motion#def-qm-brownian-motion-bm) $W$, a functional that is neither normal nor centred at zero ([Figure 17.2](#fig-qm-linear-time-series-df)). Its quantiles have no closed form; MacKinnon’s (2010) response surfaces give the 1%, 5% and 10% critical values $-3.43$, $-2.86$ and $-2.57$ for large samples. The chapter’s own simulation of 20 000 random walks of 500 steps gives $-3.46$, $-2.89$ and $-2.57$, against MacKinnon’s $-3.44$, $-2.87$ and $-2.57$ at that length.

![The law of the Dickey–Fuller t-statistic (regression with a constant) under a unit root, from 20 000 simulated random walks of 500 steps, against the standard normal. The 5% critical values are -2.86 (solid line) and -1.645 (dotted): a test with the normal table rejects a true unit root far too often. Data: the chapter’s tutorial, seeded.](https://one-course.com/images/onecourse/chapters/quant-4/qm-linear-time-series/fig-0b0f02137a59.svg)

***Figure 17.2.** The law of the Dickey–Fuller $t$-statistic (regression with a constant) under a [unit root](#def-qm-linear-time-series-unitroot), from 20 000 simulated random walks of 500 steps, against the standard normal. The 5% critical values are $-2.86$ (solid line) and $-1.645$ (dotted): a test with the normal table rejects a true [unit root](#def-qm-linear-time-series-unitroot) far too often. Data: the chapter’s tutorial, seeded.*

For the 2s10s spread, $\tau = -2.79$ with no lagged differences: not significant at 5%, significant at 10%. The augmented test depends on the lags: $-2.74$ with five, $-3.13$ with ten, $-3.55$ with twenty; AIC, searching up to forty, chooses 34 and $-3.14$. On the sample since 1990 the Dickey–Fuller statistic is $-2.01$. The honest reading is that the data cannot decide: the spread mean-reverts too slowly for fifty years of daily data to tell it from a random walk.

**Proposition 17.9 (Kendall’s bias of the AR(1) coefficient).**

For a stationary AR(1) with an estimated mean, the least-squares coefficient satisfies $\E[\hat\phi] - \phi = -(1 + 3\phi)/n + O(n^{-2})$ (Kendall, 1954).

The bias is small in the coefficient and large in the [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou), because the [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) divides by $\ln\phi$, close to zero. The spread’s estimate 0.99875, corrected by $(1 + 3 \times 0.99875)/12\,574 = 0.00032$, becomes 0.99906, and the [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) grows from 552 trading days (2.2 years) to 739 (2.9 years). The correction is an average: simulating 4 000 histories of 12 574 days from an AR(1) with $\phi = 0.99906$, the estimated [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) has median 576 days and a 10–90% range of 352 to 969 days ([Figure 17.3](#fig-qm-linear-time-series-halflife)), and in 57% of the histories the [Dickey–Fuller test](#def-qm-linear-time-series-df) does not reject the [unit root](#def-qm-linear-time-series-unitroot) at 5%, although there is none.

![Sampling distribution of the AR(1) half-life estimated from 12 574 daily observations of an AR(1) whose true half-life is 739 days, over 4 000 simulated histories (estimates above 2 000 days piled in the last bin): median 576, 10–90% range 352 to 969. The 2s10s estimate of 552 days is a typical draw. Data: the chapter’s tutorial, seeded.](https://one-course.com/images/onecourse/chapters/quant-4/qm-linear-time-series/fig-bddbc29a4088.svg)

***Figure 17.3.** Sampling distribution of the AR(1) [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) estimated from 12 574 daily observations of an AR(1) whose true [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) is 739 days, over 4 000 simulated histories (estimates above 2 000 days piled in the last bin): median 576, 10–90% range 352 to 969. The 2s10s estimate of 552 days is a typical draw. Data: the chapter’s tutorial, seeded.*

## 17.5 Long memory

**Definition 17.10 (Long memory, fractional differencing, Hurst exponent).**

A [stationary process](#def-qm-linear-time-series-stationary) has *long memory* if its autocorrelations decay like a power, $\rho(h) \sim ch^{2d-1}$ with $0 < d < \frac12$, so that they are not summable. *Fractional differencing* applies $(1 - L)^d = \sum_{j \ge 0}w_jL^j$ with $w_0 = 1$, $w_j = w_{j-1}(j - 1 -
d)/j$; the ARFIMA$(0, d, 0)$ process is $(1 - L)^dX_t = \varepsilon_t$ (Granger and Joyeux, 1980; Hosking, 1981). The *Hurst exponent* is $H = d + \frac12$, measured by the growth of ranges or variances of partial sums like $n^H$ (Hurst, 1951).

**Definition 17.11 (Fractional Brownian motion).**

*Fractional Brownian motion* with [Hurst exponent](#def-qm-linear-time-series-longmem) $H \in (0, 1)$ is the centred [Gaussian process](https://one-course.com/books/quant/4/en/chapter/2-brownian-motion#def-qm-brownian-motion-bm) with $B_H(0) = 0$ and $\E[B_H(t)B_H(s)] = \frac12(t^{2H} + s^{2H} - |t - s|^{2H})$ (Mandelbrot and Van Ness, 1968). For $H = \frac12$ it is [Brownian motion](https://one-course.com/books/quant/4/en/chapter/2-brownian-motion#def-qm-brownian-motion-bm); for $H > \frac12$ its increments are positively correlated with power-law decay, for $H < \frac12$ negatively.

For $H \ne \frac12$ [fractional Brownian motion](#def-qm-linear-time-series-fbm) is not a [semimartingale](https://one-course.com/books/quant/4/en/chapter/6-jump-processes#def-qm-jump-processes-jd), so the Itô calculus of chapter 3 does not apply to it, and a price that followed it would admit arbitrage; it is a model for volatility and for slowly mean-reverting quantities, not for prices. The difference between long and short memory is in the tail of the correlogram ([Figure 17.4](#fig-qm-linear-time-series-longmem)): an ARFIMA with $d = 0.3$ and an AR(1) with the same first autocorrelation, 0.41, differ by a factor of a thousand at lag 10. On 20 000 simulated observations the rescaled-range estimate of $H$ is 0.78 and the [log-periodogram](#def-qm-linear-time-series-spectral) (Geweke–Porter-Hudak) estimate 0.71, for a true 0.8; on [white noise](#def-qm-linear-time-series-stationary) they give 0.53 and 0.46, the rescaled range being biased upward in short blocks.

![Autocorrelations of a long-memory process (ARFIMA with d = 0.3: 20 000 simulated observations and the exact (h) = _i h(i - 1 + d)/(i - d), a straight line of slope 2d - 1 on these axes) and of an AR(1) with the same lag-one autocorrelation, which decays geometrically. Data: the chapter’s tutorial, seeded.](https://one-course.com/images/onecourse/chapters/quant-4/qm-linear-time-series/fig-ca337e612a86.svg)

***Figure 17.4.** Autocorrelations of a long-memory process (ARFIMA with $d = 0.3$: 20 000 simulated observations and the exact $\rho(h) = \prod_{i \le h}(i - 1 + d)/(i -
d)$, a straight line of slope $2d - 1$ on these axes) and of an AR(1) with the same lag-one autocorrelation, which decays geometrically. Data: the chapter’s tutorial, seeded.*

The spread’s daily changes give a [log-periodogram](#def-qm-linear-time-series-spectral) estimate $H = 0.34$ from the lowest $\sqrt n = 112$ frequencies, that is $d = -0.16$ with a [standard error](https://one-course.com/books/quant/4/en/chapter/11-estimation#def-qm-estimation-estimator) of $\pi/\sqrt{24 \times 112} = 0.06$: the changes are over-differenced, and the level behaves like a fractionally integrated process with $d \approx 0.84$, mean reverting in the long run but with hyperbolic rather than geometric decay. That is one reason AR(1) half-lives of the spread are unstable across samples: the AR(1) is the wrong shape of memory.

## 17.6 Tutorial: the half-life of a spread

**Goal.** Estimate the 2s10s spread’s persistence three ways and decide what the data can say about a [unit root](#def-qm-linear-time-series-unitroot). **End state:** Figures [17.1](#fig-qm-linear-time-series-acf), [17.2](#fig-qm-linear-time-series-df) and [17.3](#fig-qm-linear-time-series-halflife); the half-lives 552 and 739 days and the Dickey–Fuller statistics.

1. **The AR(1) with its [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) and Kendall’s correction.** `def ar1 (x) -> dict : """OLS of x_t on (1, x_{t-1}); half-life ln(1/2)/ln(rho); Kendall's correction rho + (1 + 3 rho)/n.""" x = np.asarray(x, dtype=float ) n = x.size X = np.column_stack([np.ones(n - 1 ), x[:-1 ]]) b, *_ = np.linalg.lstsq(X, x[1 :], rcond=None ) e = x[1 :] - X @ b s2 = float (e @ e) / (n - 3 ) xc = x[:-1 ] - x[:-1 ].mean() se = math.sqrt(s2 / float (xc @ xc)) rho = float (b[1 ]) rk = rho + (1 + 3 * rho) / n def hl (r): return math.log(0.5 ) / math.log(r) if 0 < r < 1 else math.inf return {" rho " : rho, " se " : se, " t_unit " : (rho - 1 ) / se, " c " : float (b[0 ]), " sigma " : math.sqrt(s2), " half_life " : hl(rho), " rho_kendall " : rk, " half_life_kendall " : hl(rk), " n " : n}` **Listing 17.1.** AR(1) by least squares, half-life and bias correction. code/firm/tsa/firm_tsa.py
2. **The augmented Dickey–Fuller regression and MacKinnon’s critical values.** `def _adf_reg (x, lags, trend): dx = np.diff(x) n = dx.size - lags cols = [x[lags:-1 ]] if trend in (" c " , " ct " ): cols.append(np.ones(n)) if trend == " ct " : cols.append(np.arange(n, dtype=float )) for k in range (1 , lags + 1 ): cols.append(dx[lags - k: dx.size - k]) X = np.column_stack(cols) y = dx[lags:] b, *_ = np.linalg.lstsq(X, y, rcond=None ) e = y - X @ b s2 = float (e @ e) / (n - X.shape[1 ]) se = math.sqrt(s2 * np.linalg.inv(X.T @ X)[0 , 0 ]) return float (b[0 ]) / se, float (e @ e), n, X.shape[1 ] def adf (x, lags: int | None = None , trend: str = " c " , max_lags: int | None = None ) -> dict : """Augmented Dickey-Fuller tau for Delta x_t = gamma x_{t-1} + deterministics + sum a_k Delta x_{t-k} + e_t. lags=None chooses them by AIC up to max_lags (default 12 (n/100)^(1/4)) on a common sample.""" x = np.asarray(x, dtype=float ) if lags is None : max_lags = int (12 * (x.size / 100 ) ** 0.25 ) if max_lags is None else max_lags best = None for k in range (max_lags + 1 ): _, ssr, n, m = _adf_reg(x[max_lags - k:], k, trend) aic = n * math.log(ssr / n) + 2 * m if best is None or aic < best[0 ]: best = (aic, k) lags = best[1 ] tau, _, n, _ = _adf_reg(x, lags, trend) return {" tau " : tau, " lags " : lags, " n " : n, " crit " : mackinnon_crit(trend, n)}` **Listing 17.2.** The augmented Dickey–Fuller test. code/firm/tsa/firm_tsa.py
3. **Run** `problem()` , `half_life_mc(0.99906, 12574)` , `spurious()` and `long_memory()` in `qm_tsa.py` , then `fig_tsa.py` .

**What to change next.** Fit an ARFIMA to the spread by the [log-periodogram](#def-qm-linear-time-series-spectral) $d$ and compare its forecast of the spread’s decay with the AR(1)’s; rerun the tests on weekly data and see which conclusions survive.

## 17.7 Build: the time-series toolkit

**Purpose.** Every persistence the miniature firm relies on (a spread’s [mean reversion](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou), a signal’s decay, a volatility’s memory) is measured and tested here before a strategy is built on it.

**Interface.** `acf`, `pacf`, `band`; `yule_walker(x, p)`; `arma_css(x, p, q)`; `ar1(x)`; `adf(x, lags, trend, max_lags)`; `mackinnon_crit(trend, n)`; `periodogram(x)`; `frac_diff(x, d, k)`; `hurst_rs(x)`, `hurst_gph(x)`.

**Rules.** Unit-root tests report their lag choice and are rerun over a range of lags; half-lives are reported with the bias correction and a simulated interval; no normal table for a coefficient near one.

**Acceptance tests.** `code/firm/tsa/tests/`: autocorrelations and partial autocorrelations of an AR(1); Yule–Walker for an AR(2); ARMA(1, 1) by conditional least squares; the [Dickey–Fuller test](#def-qm-linear-time-series-df) has its nominal size on random walks and rejects a stationary AR(1); MacKinnon’s formula; a flat [periodogram](#def-qm-linear-time-series-spectral) for [white noise](#def-qm-linear-time-series-stationary); [fractional differencing](#def-qm-linear-time-series-longmem) with $d = 1$ is differencing; Hurst estimates of [white noise](#def-qm-linear-time-series-stationary) and of an ARFIMA.

**Stretch.** The KPSS test, whose null is stationarity; exact Gaussian likelihood for ARMA by the Kalman filter of chapter 19; the local Whittle [estimator](https://one-course.com/books/quant/4/en/chapter/11-estimation#def-qm-estimation-estimator) of $d$.

Sources and further reading

- H. Wold, *A Study in the Analysis of Stationary Time Series* , Almqvist and Wiksell, 1938.
- D. A. Dickey and W. A. Fuller, “Distribution of the estimators for autoregressive time series with a unit root”, *Journal of the American Statistical Association* 74, 1979; S. E. Said and D. A. Dickey, *Biometrika* 71, 1984.
- J. G. MacKinnon, “Critical values for cointegration tests”, Queen’s Economics Department Working Paper 1227, 2010.
- C. W. J. Granger and P. Newbold, “Spurious regressions in econometrics”, *Journal of Econometrics* 2, 1974.
- M. G. Kendall, “Note on bias in the estimation of autocorrelation”, *Biometrika* 41, 1954.
- H. E. Hurst, “Long-term storage capacity of reservoirs”, *Transactions of the American Society of Civil Engineers* 116, 1951.
- B. B. Mandelbrot and J. W. Van Ness, “Fractional Brownian motions, fractional noises and applications”, *SIAM Review* 10, 1968.
- C. W. J. Granger and R. Joyeux, *Journal of Time Series Analysis* 1, 1980; J. R. M. Hosking, “Fractional differencing”, *Biometrika* 68, 1981; J. Geweke and S. Porter-Hudak, *Journal of Time Series Analysis* 4, 1983.
- Board of Governors of the Federal Reserve System, H.15 Selected Interest Rates, via FRED (series DGS2 and DGS10), accessed 24 September 2026.

## 17.8 Exercises

**Exercise 17.1 ★.**

An AR(1) has coefficient 0.98 on daily data. What is the [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) in days, and what is the autocorrelation at lag 20?

**Solution of Exercise 17.1.**

[Half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) $\ln(1/2)/\ln0.98 = 34.3$ days; $\rho(20) = 0.98^{20} = 0.67$.

**Exercise 17.2 ★.**

Is the MA(1) $X_t = \varepsilon_t + 2\varepsilon_{t-1}$ invertible? Find the invertible MA(1) with the same autocorrelations.

**Solution of Exercise 17.2.**

No: the root of $1 + 2z$ is $-1/2$, inside the unit circle. The MA(1) with coefficient $1/2$ and innovation variance four times larger, $X_t = \eta_t +
\frac12\eta_{t-1}$, has the same autocovariances ($\rho(1) = \vartheta/(1 + \vartheta^2) = 0.4$ in both cases) and is invertible.

**Exercise 17.3 ★.**

Write the first four fractional-differencing weights for $d = 0.3$.

**Solution of Exercise 17.3.**

$w_0 = 1$, $w_1 = -d = -0.3$, $w_2 = w_1(1 - d)/2 = -0.105$, $w_3 = w_2(2 - d)/3 = -0.0595$.

**Exercise 17.4 ★★.**

For an AR(1) with $\phi = 0.99$ estimated from 1 000 daily observations, what is Kendall’s bias, and how much does it change the [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou)?

**Solution of Exercise 17.4.**

$-(1 + 3 \times 0.99)/1\,000 = -0.0040$: the estimate averages 0.986 when the truth is 0.99. The [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) at 0.99 is 69.0 days; a naive analyst who estimates 0.99 should correct to 0.994, a [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) of 114.6 days. A bias of 0.4% in $\phi$ is a bias of 66% in the [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou).

**Exercise 17.5 ★★.**

Show that the autocorrelations of an AR(2) satisfy $\rho(h) = \phi_1\rho(h - 1) + \phi_2\rho(h - 2)$, and compute $\rho(1)$ and $\rho(2)$ for $\phi_1 = 0.5$, $\phi_2 =
-0.3$.

**Solution of Exercise 17.5.**

Multiply $X_t = \phi_1X_{t-1} + \phi_2X_{t-2} + \varepsilon_t$ (demeaned) by $X_{t-h}$, $h \ge 1$, and take expectations: $\gamma(h) = \phi_1\gamma(h - 1) + \phi_2\gamma(h - 2)$; divide by $\gamma(0)$. At $h = 1$, $\rho(1) = \phi_1 + \phi_2\rho(1)$, so $\rho(1) = 0.5/1.3 = 0.385$; then $\rho(2) = 0.5 \times 0.385 - 0.3 = -0.108$.

**Exercise 17.6 ★★.**

Using MacKinnon’s coefficients $(-2.86154, -2.8903, -4.234, -40.040)$, compute the 5% Dickey–Fuller critical value (with constant) for $T = 100$ and $T = 12\,574$.

**Solution of Exercise 17.6.**

$T = 100$: $-2.86154 - 0.028903 - 0.000423 - 0.000040 = -2.891$. $T = 12\,574$: $-2.862$.

**Exercise 17.7 ★★★.**

*Coding.* Regress one independent random walk of 500 steps on another 5 000 times. How often is $|t| > 1.96$? What happens if you regress the differences instead?

**Solution of Exercise 17.7.**

`spurious()`: $|t| > 1.96$ in 89% of the 5 000 regressions of levels, with median $R^2$ 0.17; regressing the differences, in 5.0%, the nominal rate. Differencing removes the common stochastic trend problem when the series are not cointegrated (chapter 20).

**Exercise 17.8 ★★★.**

*Find the flaw.* “We regressed the daily change of the spread on the lagged level; the coefficient is $-0.00125$ with $t = -2.79$, significant at 1%, so the spread mean-reverts and we will trade its reversion with a [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) of 552 days.”

**Solution of Exercise 17.8.**

The $t$-statistic of the lagged level in that regression is the Dickey–Fuller $\tau$, whose 1% and 5% critical values are $-3.43$ and $-2.86$: $-2.79$ is not significant at 5%. The [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) of 552 days is also biased down (739 after Kendall’s correction) and very uncertain (a 10–90% range of about 350 to 970 days if the truth were 739). A trade that waits years for reversion must be sized for the possibility that there is none.

## 17.9 Problem: The Half-Life of a Spread

**Problem 17.1.**

Weekend problem — how fast does the 2s10s spread mean-revert?

The data are the daily ten-year and two-year constant-maturity Treasury yields (FRED DGS10 and DGS2) from 1 June 1976 to 22 September 2026, on the 12 574 days with both; the spread is their difference in basis points.

**Part I — The series.**

1. What are the spread’s mean, range and last value?
2. What is its first autocorrelation, and what does its correlogram look like?
3. What are the standard deviation and first autocorrelations of its daily changes?
4. Which AR order does AIC choose for the changes, and does it matter economically?
5. What is the Wold representation of the changes, approximately?

**Part II — The AR(1).**

6. What are the AR(1) coefficient, its [standard error](https://one-course.com/books/quant/4/en/chapter/11-estimation#def-qm-estimation-estimator) , and the [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) ?
7. What $t$ -statistic against one, and what $p$ -value by the normal table?
8. What is Kendall’s correction, and the corrected [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) ?
9. What do the Dickey–Fuller and augmented tests say, with 0, 5, 10 and 20 lags?
10. What does the sample since 1990 give?

**Part III — What the data can tell.**

11. If the true [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) were 739 days, how would 12 574-day estimates be distributed?
12. How often would the [Dickey–Fuller test](#def-qm-linear-time-series-df) then fail to reject the [unit root](#def-qm-linear-time-series-unitroot) ?
13. What does the [log-periodogram](#def-qm-linear-time-series-spectral) say about the changes, and about the level?
14. Why might an AR(1) [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) be unstable across samples?
15. What would a steepener trader want to know that the [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) does not say?

**Part IV — Judgement.**

16. Is the spread stationary?
17. Which [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) would you use to size a mean-reversion trade, and with what caveat?
18. Why is the normal table’s $p$ -value of 0.003 misleading?
19. State the *named result* : the AR(1) [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) , its bias-corrected value, and whether a [unit root](#def-qm-linear-time-series-unitroot) can be rejected.
20. In one sentence: what does a unit-root test tell a trader?

**Solution of Problem 17.1.**

**1.** Mean 84.5 bp; range $-241$ to 291 bp; 25 bp on 22 September 2026. **2.** 0.9987; the correlogram declines almost linearly, still 0.64 after 250 days, the signature of a [unit root](#def-qm-linear-time-series-unitroot) or near [unit root](#def-qm-linear-time-series-unitroot). **3.** 4.6 bp; 0.015, 0.015, $-0.027$. **4.** AR(10), with every coefficient below 0.035 in absolute value: detectable in 12 574 days, worthless for prediction. **5.** Approximately [white noise](#def-qm-linear-time-series-stationary): $\Delta X_t \approx \varepsilon_t$ with $\psi_j$ close to zero for $j \ge 1$. **6.** $\hat\phi = 0.99875$, [standard error](https://one-course.com/books/quant/4/en/chapter/11-estimation#def-qm-estimation-estimator) 0.00045, [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) 552 trading days (2.2 years). **7.** $t = -2.79$, and a normal one-sided $p$-value of 0.003. **8.** $(1 + 3\hat\phi)/n = 0.00032$, so $\phi = 0.99906$ and a [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) of 739 days (2.9 years). **9.** With 0, 5, 10 and 20 lags: $-2.79$, $-2.74$, $-3.13$, $-3.55$, against critical values $-3.43$ (1%) and $-2.86$ (5%). Not rejected at 5% with few lags, rejected at 5% with ten and at 1% with twenty; AIC picks 34 lags and $-3.14$. **10.** $\tau = -2.01$: no evidence against a [unit root](#def-qm-linear-time-series-unitroot). **11.** Median 576 days, 10–90% range 352 to 969. **12.** In 57% of the histories. **13.** $H = 0.34$ for the changes, $d = -0.16 \pm 0.06$: over-differenced, so the level has $d \approx 0.84$, a slowly and hyperbolically mean-reverting process. **14.** Because the true memory is not geometric, the fitted AR(1) coefficient depends on the sample’s length and period; and near one the [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) magnifies every error. **15.** The distribution of the path to reversion: how far the spread can move against the position and for how long, which depends on the volatility and on the tails of the memory, not on the [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) alone. **16.** The data cannot say: consistent with a very persistent [stationary process](#def-qm-linear-time-series-stationary) and with a [unit root](#def-qm-linear-time-series-unitroot). **17.** The bias-corrected 739 days, treated as uncertain between one and four years, with the position sized to survive no reversion at all. **18.** Under a [unit root](#def-qm-linear-time-series-unitroot) the $t$-statistic follows the Dickey–Fuller law, whose 5% point is $-2.86$, not $-1.645$; the true $p$-value is about 0.06. **19.** Named result: *the [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) of a spread*: the AR(1) [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) of the 2s10s spread is 552 trading days, 739 after Kendall’s bias correction; the Dickey–Fuller statistic $-2.79$ does not reject a [unit root](#def-qm-linear-time-series-unitroot) at 5%, and augmented versions reject or not depending on the lag choice. **20.** Whether the data rule out a random walk, which, for a spread one hopes will revert, is usually the answer that it has not been ruled out.

## 17.10 Interview questions

**Interview question 17.1 ★ researcher, trader.**

What is the [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) of an AR(1) with coefficient $\phi$, and why is it hard to estimate when $\phi$ is near one?

**Solution of Interview question 17.1.**

$\ln(1/2)/\ln\phi \approx 0.69/(1 - \phi)$ periods. Near one, a small error in $\hat\phi$ is a large relative error in $1 - \hat\phi$, the estimate is biased down by about $(1 + 3\phi)/n$, and its distribution is skewed; the [half-life](https://one-course.com/books/quant/4/en/chapter/4-stochastic-differential-equations#def-qm-stochastic-differential-equations-ou) can be off by a factor of two.

*What the interviewer is looking for: The formula, the $1/(1 - \phi)$ sensitivity, and the small-sample bias.*

**Interview question 17.2 ★★ researcher.**

Why can’t you use the usual $t$-table to test $\phi = 1$?

**Solution of Interview question 17.2.**

Under $\phi = 1$ the regressor is a random walk; the [estimator](https://one-course.com/books/quant/4/en/chapter/11-estimation#def-qm-estimation-estimator) converges at rate $n$, not $\sqrt n$, and its $t$-statistic converges to a functional of [Brownian motion](https://one-course.com/books/quant/4/en/chapter/2-brownian-motion#def-qm-brownian-motion-bm) (the Dickey–Fuller law), shifted to the left: 5% critical value $-2.86$ with a constant, instead of $-1.645$.

*What the interviewer is looking for: Nonstandard asymptotics and the critical values.*

**Interview question 17.3 ★★ researcher, mle.**

You regress one price series on another and get $R^2 = 0.6$ and $t = 15$. What do you check?

**Solution of Interview question 17.3.**

Whether it is a [spurious regression](#def-qm-linear-time-series-unitroot): are both series unit-root processes, and are the residuals stationary (a cointegration test, chapter 20)? Regress changes on changes, or test the residuals with Engle–Granger critical values; look at the residuals’ autocorrelation.

*What the interviewer is looking for: [Spurious regression](#def-qm-linear-time-series-unitroot) as the first hypothesis, and cointegration as the test.*

**Interview question 17.4 ★★ researcher.**

How do you tell an AR process from an MA process from their correlograms?

**Solution of Interview question 17.4.**

An AR($p$) has a [partial autocorrelation function](#def-qm-linear-time-series-acf) that cuts off after lag $p$ and an [autocorrelation function](#def-qm-linear-time-series-acf) that decays; an MA($q$) has the reverse, an [autocorrelation function](#def-qm-linear-time-series-acf) that cuts off after lag $q$. ARMA has both decaying; use an [information criterion](#def-qm-linear-time-series-ic) to choose.

*What the interviewer is looking for: The cut-off rules and a model-selection tool.*

**Interview question 17.5 ★★ researcher, mle.**

What is [long memory](#def-qm-linear-time-series-longmem), and where do you see it in markets?

**Solution of Interview question 17.5.**

Autocorrelations that decay like a power of the lag and do not sum: the fractional $d$ between 0 and one half. In markets: volatility (absolute and squared returns), trading volume, order-flow signs, some spreads and interest rates; rarely in returns themselves.

*What the interviewer is looking for: Power-law decay, and volatility as the canonical example.*

**Interview question 17.6 ★★★ researcher.**

Why is the least-squares AR(1) coefficient biased downward in small samples, and how large is the bias?

**Solution of Interview question 17.6.**

The regression uses the sample mean, which absorbs part of each deviation, and the [estimator](https://one-course.com/books/quant/4/en/chapter/11-estimation#def-qm-estimation-estimator) is a ratio of correlated quadratic forms; both push it down. Kendall’s approximation with an estimated mean is $\E[\hat\phi] - \phi \approx -(1 + 3\phi)/n$: about $-4/n$ near one, so $-0.016$ with a year of daily data.

*What the interviewer is looking for: The mechanism and the $-(1 + 3\phi)/n$ size.*
