Quantitative Finance · Book 4 · Methods

Quantitative Methods

Quantitative Methods · Methods

17Linear Time Series

The spread between the ten-year and the two-year Treasury yields, the 2s10s that a steepener trades, has a daily autocorrelation of 0.9987 over fifty years. A regression of tomorrow’s spread on today’s gives a slope of 0.99875 with a standard error of 0.00045: a tt-statistic of −2.79-2.79 against a slope of one, a one-sided pp-value of 0.003 by the normal table. The spread seems to mean-revert, with a half-life of 552 trading days. But if the spread had a unit root, the normal table would be the wrong one: the right one puts the 5% critical value at −2.86-2.86, and the evidence falls short. This chapter is about the models that describe such series and the tests that tell them apart: stationarity and the autocorrelation function, autoregressive and moving-average models, the spectral view, unit roots and the regressions they make spurious, and long memory.

17.1 Stationarity and autocorrelation

Definition 17.1 (Stationary process, weak stationarity, white noise)

A process (Xt)(X_t) is a stationary process (strictly) if the joint law of (Xt1+h,…,Xtk+h)(X_{t_1 + h}, \dots, X_{t_k + h}) does not depend on hh. It has weak stationarity if its mean is constant and Cov⁡(Xt,Xt+h)\Cov(X_t, X_{t+h}) depends only on hh. White noise is a weakly stationary sequence with mean zero and no autocorrelation at any nonzero lag.

Definition 17.2 (Autocovariance, autocorrelation and partial autocorrelation functions)

For a weakly stationary process, the autocovariance function is γ(h)=Cov⁡(Xt,Xt+h)\gamma(h) = \Cov(X_t, X_{t+h}), and the autocorrelation function is ρ(h)=γ(h)/γ(0)\rho(h) = \gamma(h)/\gamma(0). The partial autocorrelation function at lag hh is the last coefficient ϕhh\phi_{hh} of the best linear predictor of Xt+hX_{t+h} from Xt+h−1,…,XtX_{t+h-1}, \dots, X_t: the correlation at lag hh once the intermediate lags are accounted for.

Under white noise the sample autocorrelations are approximately independent N(0,1/n)\mathcal N(0, 1/n), whence the ±2/n\pm 2/\sqrt n bands of every correlogram. The decomposition that justifies the models of the next section is Wold’s (1938): every weakly stationary, purely nondeterministic process is an infinite moving average Xt=μ+∑j≥0ψjεt−jX_t = \mu + \sum_{j \ge 0}\psi_j\varepsilon_{t-j} of its one-step prediction errors εt\varepsilon_t, a white noise, with ψ0=1\psi_0 = 1 and ∑jψj2<∞\sum_j\psi_j^2 < \infty.

The 2s10s spread’s correlogram (Figure 17.1) is the signature of a nearly nonstationary level: autocorrelations that start at 0.9987 and fall almost linearly over a year of lags. Its daily changes are nearly white noise, with a standard deviation of 4.6 basis points and autocorrelations of 0.015, 0.015 and −0.027-0.027 at the first three lags, within the band ±0.018\pm 0.018 for the first two.

Sample autocorrelations of the 2s10s Treasury spread, 1 June 1976 to 22 September 2026 (12 574 days). Left: the level, lags 1 to 250, from 0.9987 down. Right: the daily change, lags 1 to 40, with the ± 2/√ n band. Data: FRED series DGS10 and DGS2 (Board of Governors of the Federal Reserve System, H.15).
Figure 17.1. Sample autocorrelations of the 2s10s Treasury spread, 1 June 1976 to 22 September 2026 (12 574 days). Left: the level, lags 1 to 250, from 0.9987 down. Right: the daily change, lags 1 to 40, with the ±2/n\pm 2/\sqrt n band. Data: FRED series DGS10 and DGS2 (Board of Governors of the Federal Reserve System, H.15).

17.2 Autoregressive and moving-average models

Definition 17.3 (Autoregressive, moving-average and ARMA processes)

With (εt)(\varepsilon_t) white noise of variance σ2\sigma^2: an autoregressive process of order pp, AR(pp), satisfies Xt=c+∑i=1pϕiXt−i+εtX_t = c + \sum_{i=1}^p\phi_iX_{t-i} + \varepsilon_t; a moving-average process of order qq, MA(qq), is Xt=μ+εt+∑j=1qϑjεt−jX_t = \mu + \varepsilon_t + \sum_{j=1}^q\vartheta_j\varepsilon_{t-j}; an ARMA process combines them, ϕ(L)Xt=c+ϑ(L)εt\phi(L)X_t = c + \vartheta(L)\varepsilon_t with ϕ(z)=1−∑iϕizi\phi(z) = 1 - \sum_i\phi_iz^i, ϑ(z)=1+∑jϑjzj\vartheta(z) = 1 + \sum_j\vartheta_jz^j and LL the lag operator.

Proposition 17.4 (Stationarity and invertibility of ARMA)

The ARMA equation has a unique weakly stationary causal solution Xt=μ+∑j≥0ψjεt−jX_t = \mu + \sum_{j \ge 0}\psi_j\varepsilon_{t-j} if and only if every root of ϕ(z)\phi(z) lies outside the unit circle; it is invertible, εt=∑j≥0πjXt−j\varepsilon_t = \sum_{j \ge 0}\pi_jX_{t-j} up to a constant, if and only if every root of ϑ(z)\vartheta(z) does, the polynomials having no common root.

Proof. Factor ϕ(z)=∏i(1−z/zi)\phi(z) = \prod_i(1 - z/z_i). For ∣zi∣>1|z_i| > 1, (1−L/zi)−1=∑kzi−kLk(1 - L/z_i)^{-1} = \sum_kz_i^{-k}L^k is an absolutely summable filter, so ψ(z)=ϑ(z)/ϕ(z)\psi(z) = \vartheta(z)/\phi(z) has summable coefficients and defines the solution; if some ∣zi∣≤1|z_i| \le 1 no summable causal inverse exists (for ∣zi∣=1|z_i| = 1, none at all). The MA side is the same argument applied to ϑ\vartheta. ∎

For the AR(1) Xt=c+ϕXt−1+εtX_t = c + \phi X_{t-1} + \varepsilon_t, the condition is ∣ϕ∣<1|\phi| < 1, the autocorrelations are ϕh\phi^h, and a deviation from the mean decays by half in ln⁡(1/2)/ln⁡ϕ\ln(1/2)/\ln\phi periods, the half-life of chapter 4’s Ornstein–Uhlenbeck process sampled daily. An AR(pp) has a partial autocorrelation function that cuts off after lag pp; an MA(qq) has an autocorrelation function that cuts off after lag qq. The orders are chosen by an information criterion.

Definition 17.5 (Information criterion)

An information criterion ranks models by fit penalised for size: Akaike’s AIC=−2ℓ+2k\mathrm{AIC} = -2\ell + 2k and Schwarz’s BIC=−2ℓ+kln⁡n\mathrm{BIC} = -2\ell + k\ln n, for maximised log-likelihood ℓ\ell and kk parameters; the smallest wins.

AIC aims at prediction and tends to choose large models; BIC is consistent for the true order when there is one. On the 2s10s daily changes AIC chooses an AR(10) whose coefficients are all below 0.035 in absolute value: statistically detectable with 12 574 days, economically nothing.

17.3 The spectral view

Definition 17.6 (Spectral density, periodogram)

The spectral density of a weakly stationary process with summable autocovariances is f(ω)=12π∑hγ(h)e−ihωf(\omega) = \frac1{2\pi}\sum_h\gamma(h)e^{-ih\omega}, ω∈[−π,π]\omega \in [-\pi, \pi]: the decomposition of its variance by frequency, γ(0)=∫−ππf\gamma(0) = \int_{-\pi}^\pi f. The periodogram I(ωj)=12πn∣∑tXte−itωj∣2I(\omega_j) = \frac1{2\pi n}\bigl\lvert\sum_tX_te^{-it\omega_j}\bigr\rvert^2 at the Fourier frequencies ωj=2πj/n\omega_j = 2\pi j/n is its sample counterpart.

White noise has a flat spectrum, σ2/2π\sigma^2/2\pi; an AR(1) with ϕ\phi near one concentrates its variance at low frequencies, f(ω)=σ2/(2π∣1−ϕe−iω∣2)f(\omega) = \sigma^2/(2\pi|1 - \phi e^{-i\omega}|^2). The periodogram is asymptotically unbiased but not consistent: each ordinate is approximately f(ωj)f(\omega_j) times an exponential variable, however long the sample, so it must be smoothed or averaged. Its low-frequency end is where persistence shows, and it is the basis of the long-memory estimate of the last section. f(0)=12π∑hγ(h)f(0) = \frac1{2\pi}\sum_h\gamma(h) is, up to 2π2\pi, the long-run variance of chapter 11.

17.4 Unit roots

Definition 17.7 (Unit root, spurious regression)

A process has a unit root if its autoregressive polynomial has a root at z=1z = 1: its first difference is stationary but the level is not, like a random walk. A spurious regression is a regression between independent processes with unit roots whose usual tt-statistics and R2R^2 indicate a relation that does not exist.

Granger and Newbold (1974) showed the problem by simulation, and it is easy to repeat: regressing one random walk of 500 steps on another, independent one, the slope’s tt-statistic exceeds 1.96 in absolute value in 89% of 5 000 trials, with a median R2R^2 of 0.17. The residual inherits the unit root, the observations are not independent, and every formula that assumed they were is wrong. The same failure hits the test of ϕ=1\phi = 1 itself.

Definition 17.8 (Dickey–Fuller test)

The Dickey–Fuller test of a unit root regresses ΔXt\Delta X_t on Xt−1X_{t-1} (with a constant, and a trend if wanted) and compares the tt-statistic τ\tau of the coefficient on Xt−1X_{t-1} with the Dickey–Fuller distribution rather than the normal; the augmented version adds lagged differences ΔXt−k\Delta X_{t-k} to absorb short-run dependence (Said and Dickey, 1984).

Under the unit root and with a constant, τ\tau converges to (12(W(1)2−1)−W(1)∫01W)/(∫01W2−(∫01W)2)1/2\bigl(\frac12(W(1)^2 - 1) - W(1)\int_0^1W\bigr)/\bigl(\int_0^1W^2 - (\int_0^1W)^2\bigr)^{1/2} for a standard Brownian motion WW, a functional that is neither normal nor centred at zero (Figure 17.2). Its quantiles have no closed form; MacKinnon’s (2010) response surfaces give the 1%, 5% and 10% critical values −3.43-3.43, −2.86-2.86 and −2.57-2.57 for large samples. The chapter’s own simulation of 20 000 random walks of 500 steps gives −3.46-3.46, −2.89-2.89 and −2.57-2.57, against MacKinnon’s −3.44-3.44, −2.87-2.87 and −2.57-2.57 at that length.

The law of the Dickey–Fuller t-statistic (regression with a constant) under a unit root, from 20 000 simulated random walks of 500 steps, against the standard normal. The 5% critical values are -2.86 (solid line) and -1.645 (dotted): a test with the normal table rejects a true unit root far too often. Data: the chapter’s tutorial, seeded.
Figure 17.2. The law of the Dickey–Fuller tt-statistic (regression with a constant) under a unit root, from 20 000 simulated random walks of 500 steps, against the standard normal. The 5% critical values are −2.86-2.86 (solid line) and −1.645-1.645 (dotted): a test with the normal table rejects a true unit root far too often. Data: the chapter’s tutorial, seeded.

For the 2s10s spread, τ=−2.79\tau = -2.79 with no lagged differences: not significant at 5%, significant at 10%. The augmented test depends on the lags: −2.74-2.74 with five, −3.13-3.13 with ten, −3.55-3.55 with twenty; AIC, searching up to forty, chooses 34 and −3.14-3.14. On the sample since 1990 the Dickey–Fuller statistic is −2.01-2.01. The honest reading is that the data cannot decide: the spread mean-reverts too slowly for fifty years of daily data to tell it from a random walk.

Proposition 17.9 (Kendall’s bias of the AR(1) coefficient)

For a stationary AR(1) with an estimated mean, the least-squares coefficient satisfies E[ϕ^]−ϕ=−(1+3ϕ)/n+O(n−2)\E[\hat\phi] - \phi = -(1 + 3\phi)/n + O(n^{-2}) (Kendall, 1954).

The bias is small in the coefficient and large in the half-life, because the half-life divides by ln⁡ϕ\ln\phi, close to zero. The spread’s estimate 0.99875, corrected by (1+3×0.99875)/12 574=0.00032(1 + 3 \times 0.99875)/12\,574 = 0.00032, becomes 0.99906, and the half-life grows from 552 trading days (2.2 years) to 739 (2.9 years). The correction is an average: simulating 4 000 histories of 12 574 days from an AR(1) with ϕ=0.99906\phi = 0.99906, the estimated half-life has median 576 days and a 10–90% range of 352 to 969 days (Figure 17.3), and in 57% of the histories the Dickey–Fuller test does not reject the unit root at 5%, although there is none.

Sampling distribution of the AR(1) half-life estimated from 12 574 daily observations of an AR(1) whose true half-life is 739 days, over 4 000 simulated histories (estimates above 2 000 days piled in the last bin): median 576, 10–90% range 352 to 969. The 2s10s estimate of 552 days is a typical draw. Data: the chapter’s tutorial, seeded.
Figure 17.3. Sampling distribution of the AR(1) half-life estimated from 12 574 daily observations of an AR(1) whose true half-life is 739 days, over 4 000 simulated histories (estimates above 2 000 days piled in the last bin): median 576, 10–90% range 352 to 969. The 2s10s estimate of 552 days is a typical draw. Data: the chapter’s tutorial, seeded.

17.5 Long memory

Definition 17.10 (Long memory, fractional differencing, Hurst exponent)

A stationary process has long memory if its autocorrelations decay like a power, ρ(h)∼ch2d−1\rho(h) \sim ch^{2d-1} with 0<d<120 < d < \frac12, so that they are not summable. Fractional differencing applies (1−L)d=∑j≥0wjLj(1 - L)^d = \sum_{j \ge 0}w_jL^j with w0=1w_0 = 1, wj=wj−1(j−1−d)/jw_j = w_{j-1}(j - 1 - d)/j; the ARFIMA(0,d,0)(0, d, 0) process is (1−L)dXt=εt(1 - L)^dX_t = \varepsilon_t (Granger and Joyeux, 1980; Hosking, 1981). The Hurst exponent is H=d+12H = d + \frac12, measured by the growth of ranges or variances of partial sums like nHn^H (Hurst, 1951).

Definition 17.11 (Fractional Brownian motion)

Fractional Brownian motion with Hurst exponent H∈(0,1)H \in (0, 1) is the centred Gaussian process with BH(0)=0B_H(0) = 0 and E[BH(t)BH(s)]=12(t2H+s2H−∣t−s∣2H)\E[B_H(t)B_H(s)] = \frac12(t^{2H} + s^{2H} - |t - s|^{2H}) (Mandelbrot and Van Ness, 1968). For H=12H = \frac12 it is Brownian motion; for H>12H > \frac12 its increments are positively correlated with power-law decay, for H<12H < \frac12 negatively.

For H≠12H \ne \frac12 fractional Brownian motion is not a semimartingale, so the Itô calculus of chapter 3 does not apply to it, and a price that followed it would admit arbitrage; it is a model for volatility and for slowly mean-reverting quantities, not for prices. The difference between long and short memory is in the tail of the correlogram (Figure 17.4): an ARFIMA with d=0.3d = 0.3 and an AR(1) with the same first autocorrelation, 0.41, differ by a factor of a thousand at lag 10. On 20 000 simulated observations the rescaled-range estimate of HH is 0.78 and the log-periodogram (Geweke–Porter-Hudak) estimate 0.71, for a true 0.8; on white noise they give 0.53 and 0.46, the rescaled range being biased upward in short blocks.

Autocorrelations of a long-memory process (ARFIMA with d = 0.3: 20 000 simulated observations and the exact (h) = _i h(i - 1 + d)/(i - d), a straight line of slope 2d - 1 on these axes) and of an AR(1) with the same lag-one autocorrelation, which decays geometrically. Data: the chapter’s tutorial, seeded.
Figure 17.4. Autocorrelations of a long-memory process (ARFIMA with d=0.3d = 0.3: 20 000 simulated observations and the exact ρ(h)=∏i≤h(i−1+d)/(i−d)\rho(h) = \prod_{i \le h}(i - 1 + d)/(i - d), a straight line of slope 2d−12d - 1 on these axes) and of an AR(1) with the same lag-one autocorrelation, which decays geometrically. Data: the chapter’s tutorial, seeded.

The spread’s daily changes give a log-periodogram estimate H=0.34H = 0.34 from the lowest n=112\sqrt n = 112 frequencies, that is d=−0.16d = -0.16 with a standard error of π/24×112=0.06\pi/\sqrt{24 \times 112} = 0.06: the changes are over-differenced, and the level behaves like a fractionally integrated process with d≈0.84d \approx 0.84, mean reverting in the long run but with hyperbolic rather than geometric decay. That is one reason AR(1) half-lives of the spread are unstable across samples: the AR(1) is the wrong shape of memory.

17.6 Tutorial: the half-life of a spread

Goal. Estimate the 2s10s spread’s persistence three ways and decide what the data can say about a unit root. End state: Figures 17.1, 17.2 and 17.3; the half-lives 552 and 739 days and the Dickey–Fuller statistics.

  1. The AR(1) with its half-life and Kendall’s correction.

    def ar1(x) -> dict:
        """OLS of x_t on (1, x_{t-1}); half-life ln(1/2)/ln(rho); Kendall's correction rho + (1 + 3 rho)/n."""
        x = np.asarray(x, dtype=float)
        n = x.size
        X = np.column_stack([np.ones(n - 1), x[:-1]])
        b, *_ = np.linalg.lstsq(X, x[1:], rcond=None)
        e = x[1:] - X @ b
        s2 = float(e @ e) / (n - 3)
        xc = x[:-1] - x[:-1].mean()
        se = math.sqrt(s2 / float(xc @ xc))
        rho = float(b[1])
        rk = rho + (1 + 3 * rho) / n
    
        def hl(r):
            return math.log(0.5) / math.log(r) if 0 < r < 1 else math.inf
        return {"rho": rho, "se": se, "t_unit": (rho - 1) / se, "c": float(b[0]), "sigma": math.sqrt(s2),
                "half_life": hl(rho), "rho_kendall": rk, "half_life_kendall": hl(rk), "n": n}
    Listing 17.1. AR(1) by least squares, half-life and bias correction. code/firm/tsa/firm_tsa.py
  2. The augmented Dickey–Fuller regression and MacKinnon’s critical values.

    def _adf_reg(x, lags, trend):
        dx = np.diff(x)
        n = dx.size - lags
        cols = [x[lags:-1]]
        if trend in ("c", "ct"):
            cols.append(np.ones(n))
        if trend == "ct":
            cols.append(np.arange(n, dtype=float))
        for k in range(1, lags + 1):
            cols.append(dx[lags - k: dx.size - k])
        X = np.column_stack(cols)
        y = dx[lags:]
        b, *_ = np.linalg.lstsq(X, y, rcond=None)
        e = y - X @ b
        s2 = float(e @ e) / (n - X.shape[1])
        se = math.sqrt(s2 * np.linalg.inv(X.T @ X)[0, 0])
        return float(b[0]) / se, float(e @ e), n, X.shape[1]
    
    
    def adf(x, lags: int | None = None, trend: str = "c", max_lags: int | None = None) -> dict:
        """Augmented Dickey-Fuller tau for Delta x_t = gamma x_{t-1} + deterministics + sum a_k Delta x_{t-k} + e_t.
        lags=None chooses them by AIC up to max_lags (default 12 (n/100)^(1/4)) on a common sample."""
        x = np.asarray(x, dtype=float)
        if lags is None:
            max_lags = int(12 * (x.size / 100) ** 0.25) if max_lags is None else max_lags
            best = None
            for k in range(max_lags + 1):
                _, ssr, n, m = _adf_reg(x[max_lags - k:], k, trend)
                aic = n * math.log(ssr / n) + 2 * m
                if best is None or aic < best[0]:
                    best = (aic, k)
            lags = best[1]
        tau, _, n, _ = _adf_reg(x, lags, trend)
        return {"tau": tau, "lags": lags, "n": n, "crit": mackinnon_crit(trend, n)}
    Listing 17.2. The augmented Dickey–Fuller test. code/firm/tsa/firm_tsa.py
  3. Run problem(), half_life_mc(0.99906, 12574), spurious() and long_memory() in qm_tsa.py, then fig_tsa.py.

What to change next. Fit an ARFIMA to the spread by the log-periodogram dd and compare its forecast of the spread’s decay with the AR(1)’s; rerun the tests on weekly data and see which conclusions survive.

17.7 Build: the time-series toolkit

Purpose. Every persistence the miniature firm relies on (a spread’s mean reversion, a signal’s decay, a volatility’s memory) is measured and tested here before a strategy is built on it.

Interface. acf, pacf, band; yule_walker(x, p); arma_css(x, p, q); ar1(x); adf(x, lags, trend, max_lags); mackinnon_crit(trend, n); periodogram(x); frac_diff(x, d, k); hurst_rs(x), hurst_gph(x).

Rules. Unit-root tests report their lag choice and are rerun over a range of lags; half-lives are reported with the bias correction and a simulated interval; no normal table for a coefficient near one.

Acceptance tests. code/firm/tsa/tests/: autocorrelations and partial autocorrelations of an AR(1); Yule–Walker for an AR(2); ARMA(1, 1) by conditional least squares; the Dickey–Fuller test has its nominal size on random walks and rejects a stationary AR(1); MacKinnon’s formula; a flat periodogram for white noise; fractional differencing with d=1d = 1 is differencing; Hurst estimates of white noise and of an ARFIMA.

Stretch. The KPSS test, whose null is stationarity; exact Gaussian likelihood for ARMA by the Kalman filter of chapter 19; the local Whittle estimator of dd.

Sources and further reading

  • H. Wold, A Study in the Analysis of Stationary Time Series, Almqvist and Wiksell, 1938.
  • D. A. Dickey and W. A. Fuller, “Distribution of the estimators for autoregressive time series with a unit root”, Journal of the American Statistical Association 74, 1979; S. E. Said and D. A. Dickey, Biometrika 71, 1984.
  • J. G. MacKinnon, “Critical values for cointegration tests”, Queen’s Economics Department Working Paper 1227, 2010.
  • C. W. J. Granger and P. Newbold, “Spurious regressions in econometrics”, Journal of Econometrics 2, 1974.
  • M. G. Kendall, “Note on bias in the estimation of autocorrelation”, Biometrika 41, 1954.
  • H. E. Hurst, “Long-term storage capacity of reservoirs”, Transactions of the American Society of Civil Engineers 116, 1951.
  • B. B. Mandelbrot and J. W. Van Ness, “Fractional Brownian motions, fractional noises and applications”, SIAM Review 10, 1968.
  • C. W. J. Granger and R. Joyeux, Journal of Time Series Analysis 1, 1980; J. R. M. Hosking, “Fractional differencing”, Biometrika 68, 1981; J. Geweke and S. Porter-Hudak, Journal of Time Series Analysis 4, 1983.
  • Board of Governors of the Federal Reserve System, H.15 Selected Interest Rates, via FRED (series DGS2 and DGS10), accessed 24 September 2026.

17.8 Exercises

Exercise 17.1 ★

An AR(1) has coefficient 0.98 on daily data. What is the half-life in days, and what is the autocorrelation at lag 20?

Solution

Solution of Exercise 17.1.

Half-life ln⁡(1/2)/ln⁡0.98=34.3\ln(1/2)/\ln0.98 = 34.3 days; ρ(20)=0.9820=0.67\rho(20) = 0.98^{20} = 0.67.

Exercise 17.2 ★

Is the MA(1) Xt=εt+2εt−1X_t = \varepsilon_t + 2\varepsilon_{t-1} invertible? Find the invertible MA(1) with the same autocorrelations.

Solution

Solution of Exercise 17.2.

No: the root of 1+2z1 + 2z is −1/2-1/2, inside the unit circle. The MA(1) with coefficient 1/21/2 and innovation variance four times larger, Xt=ηt+12ηt−1X_t = \eta_t + \frac12\eta_{t-1}, has the same autocovariances (ρ(1)=ϑ/(1+ϑ2)=0.4\rho(1) = \vartheta/(1 + \vartheta^2) = 0.4 in both cases) and is invertible.

Exercise 17.3 ★

Write the first four fractional-differencing weights for d=0.3d = 0.3.

Solution

Solution of Exercise 17.3.

w0=1w_0 = 1, w1=−d=−0.3w_1 = -d = -0.3, w2=w1(1−d)/2=−0.105w_2 = w_1(1 - d)/2 = -0.105, w3=w2(2−d)/3=−0.0595w_3 = w_2(2 - d)/3 = -0.0595.

Exercise 17.4 ★★

For an AR(1) with ϕ=0.99\phi = 0.99 estimated from 1 000 daily observations, what is Kendall’s bias, and how much does it change the half-life?

Solution

Solution of Exercise 17.4.

−(1+3×0.99)/1 000=−0.0040-(1 + 3 \times 0.99)/1\,000 = -0.0040: the estimate averages 0.986 when the truth is 0.99. The half-life at 0.99 is 69.0 days; a naive analyst who estimates 0.99 should correct to 0.994, a half-life of 114.6 days. A bias of 0.4% in ϕ\phi is a bias of 66% in the half-life.

Exercise 17.5 ★★

Show that the autocorrelations of an AR(2) satisfy ρ(h)=ϕ1ρ(h−1)+ϕ2ρ(h−2)\rho(h) = \phi_1\rho(h - 1) + \phi_2\rho(h - 2), and compute ρ(1)\rho(1) and ρ(2)\rho(2) for ϕ1=0.5\phi_1 = 0.5, ϕ2=−0.3\phi_2 = -0.3.

Solution

Solution of Exercise 17.5.

Multiply Xt=ϕ1Xt−1+ϕ2Xt−2+εtX_t = \phi_1X_{t-1} + \phi_2X_{t-2} + \varepsilon_t (demeaned) by Xt−hX_{t-h}, h≥1h \ge 1, and take expectations: γ(h)=ϕ1γ(h−1)+ϕ2γ(h−2)\gamma(h) = \phi_1\gamma(h - 1) + \phi_2\gamma(h - 2); divide by γ(0)\gamma(0). At h=1h = 1, ρ(1)=ϕ1+ϕ2ρ(1)\rho(1) = \phi_1 + \phi_2\rho(1), so ρ(1)=0.5/1.3=0.385\rho(1) = 0.5/1.3 = 0.385; then ρ(2)=0.5×0.385−0.3=−0.108\rho(2) = 0.5 \times 0.385 - 0.3 = -0.108.

Exercise 17.6 ★★

Using MacKinnon’s coefficients (−2.86154,−2.8903,−4.234,−40.040)(-2.86154, -2.8903, -4.234, -40.040), compute the 5% Dickey–Fuller critical value (with constant) for T=100T = 100 and T=12 574T = 12\,574.

Solution

Solution of Exercise 17.6.

T=100T = 100: −2.86154−0.028903−0.000423−0.000040=−2.891-2.86154 - 0.028903 - 0.000423 - 0.000040 = -2.891. T=12 574T = 12\,574: −2.862-2.862.

Exercise 17.7 ★★★

Coding. Regress one independent random walk of 500 steps on another 5 000 times. How often is ∣t∣>1.96|t| > 1.96? What happens if you regress the differences instead?

Solution

Solution of Exercise 17.7.

spurious(): ∣t∣>1.96|t| > 1.96 in 89% of the 5 000 regressions of levels, with median R2R^2 0.17; regressing the differences, in 5.0%, the nominal rate. Differencing removes the common stochastic trend problem when the series are not cointegrated (chapter 20).

Exercise 17.8 ★★★

Find the flaw. “We regressed the daily change of the spread on the lagged level; the coefficient is −0.00125-0.00125 with t=−2.79t = -2.79, significant at 1%, so the spread mean-reverts and we will trade its reversion with a half-life of 552 days.”

Solution

Solution of Exercise 17.8.

The tt-statistic of the lagged level in that regression is the Dickey–Fuller τ\tau, whose 1% and 5% critical values are −3.43-3.43 and −2.86-2.86: −2.79-2.79 is not significant at 5%. The half-life of 552 days is also biased down (739 after Kendall’s correction) and very uncertain (a 10–90% range of about 350 to 970 days if the truth were 739). A trade that waits years for reversion must be sized for the possibility that there is none.

17.9 Problem: The Half-Life of a Spread

Problem 17.1

Weekend problem — how fast does the 2s10s spread mean-revert?

The data are the daily ten-year and two-year constant-maturity Treasury yields (FRED DGS10 and DGS2) from 1 June 1976 to 22 September 2026, on the 12 574 days with both; the spread is their difference in basis points.

Part I — The series.

  1. What are the spread’s mean, range and last value?
  2. What is its first autocorrelation, and what does its correlogram look like?
  3. What are the standard deviation and first autocorrelations of its daily changes?
  4. Which AR order does AIC choose for the changes, and does it matter economically?
  5. What is the Wold representation of the changes, approximately?

Part II — The AR(1).

  1. What are the AR(1) coefficient, its standard error, and the half-life?
  2. What tt-statistic against one, and what pp-value by the normal table?
  3. What is Kendall’s correction, and the corrected half-life?
  4. What do the Dickey–Fuller and augmented tests say, with 0, 5, 10 and 20 lags?
  5. What does the sample since 1990 give?

Part III — What the data can tell.

  1. If the true half-life were 739 days, how would 12 574-day estimates be distributed?
  2. How often would the Dickey–Fuller test then fail to reject the unit root?
  3. What does the log-periodogram say about the changes, and about the level?
  4. Why might an AR(1) half-life be unstable across samples?
  5. What would a steepener trader want to know that the half-life does not say?

Part IV — Judgement.

  1. Is the spread stationary?
  2. Which half-life would you use to size a mean-reversion trade, and with what caveat?
  3. Why is the normal table’s pp-value of 0.003 misleading?
  4. State the named result: the AR(1) half-life, its bias-corrected value, and whether a unit root can be rejected.
  5. In one sentence: what does a unit-root test tell a trader?
Solution

Solution of Problem 17.1.

1. Mean 84.5 bp; range −241-241 to 291 bp; 25 bp on 22 September 2026. 2. 0.9987; the correlogram declines almost linearly, still 0.64 after 250 days, the signature of a unit root or near unit root. 3. 4.6 bp; 0.015, 0.015, −0.027-0.027. 4. AR(10), with every coefficient below 0.035 in absolute value: detectable in 12 574 days, worthless for prediction. 5. Approximately white noise: ΔXt≈εt\Delta X_t \approx \varepsilon_t with ψj\psi_j close to zero for j≥1j \ge 1. 6. ϕ^=0.99875\hat\phi = 0.99875, standard error 0.00045, half-life 552 trading days (2.2 years). 7. t=−2.79t = -2.79, and a normal one-sided pp-value of 0.003. 8. (1+3ϕ^)/n=0.00032(1 + 3\hat\phi)/n = 0.00032, so ϕ=0.99906\phi = 0.99906 and a half-life of 739 days (2.9 years). 9. With 0, 5, 10 and 20 lags: −2.79-2.79, −2.74-2.74, −3.13-3.13, −3.55-3.55, against critical values −3.43-3.43 (1%) and −2.86-2.86 (5%). Not rejected at 5% with few lags, rejected at 5% with ten and at 1% with twenty; AIC picks 34 lags and −3.14-3.14. 10. τ=−2.01\tau = -2.01: no evidence against a unit root. 11. Median 576 days, 10–90% range 352 to 969. 12. In 57% of the histories. 13. H=0.34H = 0.34 for the changes, d=−0.16±0.06d = -0.16 \pm 0.06: over-differenced, so the level has d≈0.84d \approx 0.84, a slowly and hyperbolically mean-reverting process. 14. Because the true memory is not geometric, the fitted AR(1) coefficient depends on the sample’s length and period; and near one the half-life magnifies every error. 15. The distribution of the path to reversion: how far the spread can move against the position and for how long, which depends on the volatility and on the tails of the memory, not on the half-life alone. 16. The data cannot say: consistent with a very persistent stationary process and with a unit root. 17. The bias-corrected 739 days, treated as uncertain between one and four years, with the position sized to survive no reversion at all. 18. Under a unit root the tt-statistic follows the Dickey–Fuller law, whose 5% point is −2.86-2.86, not −1.645-1.645; the true pp-value is about 0.06. 19. Named result: the half-life of a spread: the AR(1) half-life of the 2s10s spread is 552 trading days, 739 after Kendall’s bias correction; the Dickey–Fuller statistic −2.79-2.79 does not reject a unit root at 5%, and augmented versions reject or not depending on the lag choice. 20. Whether the data rule out a random walk, which, for a spread one hopes will revert, is usually the answer that it has not been ruled out.

17.10 Interview questions

Interview question 17.1 ★ researcher, trader

What is the half-life of an AR(1) with coefficient ϕ\phi, and why is it hard to estimate when ϕ\phi is near one?

Solution

Solution of Interview question 17.1.

ln⁡(1/2)/ln⁡ϕ≈0.69/(1−ϕ)\ln(1/2)/\ln\phi \approx 0.69/(1 - \phi) periods. Near one, a small error in ϕ^\hat\phi is a large relative error in 1−ϕ^1 - \hat\phi, the estimate is biased down by about (1+3ϕ)/n(1 + 3\phi)/n, and its distribution is skewed; the half-life can be off by a factor of two.

What the interviewer is looking for: The formula, the 1/(1−ϕ)1/(1 - \phi) sensitivity, and the small-sample bias.

Interview question 17.2 ★★ researcher

Why can’t you use the usual tt-table to test ϕ=1\phi = 1?

Solution

Solution of Interview question 17.2.

Under ϕ=1\phi = 1 the regressor is a random walk; the estimator converges at rate nn, not n\sqrt n, and its tt-statistic converges to a functional of Brownian motion (the Dickey–Fuller law), shifted to the left: 5% critical value −2.86-2.86 with a constant, instead of −1.645-1.645.

What the interviewer is looking for: Nonstandard asymptotics and the critical values.

Interview question 17.3 ★★ researcher, mle

You regress one price series on another and get R2=0.6R^2 = 0.6 and t=15t = 15. What do you check?

Solution

Solution of Interview question 17.3.

Whether it is a spurious regression: are both series unit-root processes, and are the residuals stationary (a cointegration test, chapter 20)? Regress changes on changes, or test the residuals with Engle–Granger critical values; look at the residuals’ autocorrelation.

What the interviewer is looking for: Spurious regression as the first hypothesis, and cointegration as the test.

Interview question 17.4 ★★ researcher

How do you tell an AR process from an MA process from their correlograms?

Solution

Solution of Interview question 17.4.

An AR(pp) has a partial autocorrelation function that cuts off after lag pp and an autocorrelation function that decays; an MA(qq) has the reverse, an autocorrelation function that cuts off after lag qq. ARMA has both decaying; use an information criterion to choose.

What the interviewer is looking for: The cut-off rules and a model-selection tool.

Interview question 17.5 ★★ researcher, mle

What is long memory, and where do you see it in markets?

Solution

Solution of Interview question 17.5.

Autocorrelations that decay like a power of the lag and do not sum: the fractional dd between 0 and one half. In markets: volatility (absolute and squared returns), trading volume, order-flow signs, some spreads and interest rates; rarely in returns themselves.

What the interviewer is looking for: Power-law decay, and volatility as the canonical example.

Interview question 17.6 ★★★ researcher

Why is the least-squares AR(1) coefficient biased downward in small samples, and how large is the bias?

Solution

Solution of Interview question 17.6.

The regression uses the sample mean, which absorbs part of each deviation, and the estimator is a ratio of correlated quadratic forms; both push it down. Kendall’s approximation with an estimated mean is E[ϕ^]−ϕ≈−(1+3ϕ)/n\E[\hat\phi] - \phi \approx -(1 + 3\phi)/n: about −4/n-4/n near one, so −0.016-0.016 with a year of daily data.

What the interviewer is looking for: The mechanism and the −(1+3ϕ)/n-(1 + 3\phi)/n size.

Terms defined in this chapter

See all 2333 terms in the glossary