Quantitative Finance · Book 4 · Methods

Quantitative Methods

Quantitative Methods · Methods

11Estimation

A researcher presents a strategy: over five years its daily P&L averaged 4.0 basis points with a standard error of 1.2, a tt-statistic of 3.4, comfortably significant. A colleague notices that each day’s position is held for five days, so consecutive days share four fifths of their positions, and recomputes the standard error allowing for the autocorrelation that creates: 2.3. The tt-statistic falls to 1.7, and the strategy is no longer distinguishable from luck. Nothing was wrong with the average; the formula for its uncertainty assumed independent days. This chapter is about estimators and their uncertainty: maximum likelihood and its asymptotics, what happens when the model is wrong, M-estimators and moments, the delta method that gives the Sharpe ratio its standard error, and the heteroskedasticity- and autocorrelation-consistent variances that repaired the researcher’s tt-statistic.

11.1 Maximum likelihood

Definition 11.1 (Estimator, unbiased, consistent, standard error)

An estimator θ^n\hat\theta_n of a parameter θ0\theta_0 is a function of the data X1,…,XnX_1, \dots, X_n. It is an unbiased estimator if E[θ^n]=θ0\E[\hat\theta_n] = \theta_0, and a consistent estimator if θ^n→Pθ0\hat\theta_n \xrightarrow{\P} \theta_0 as n→∞n \to \infty. Its standard error se(θ^n)\mathrm{se}(\hat\theta_n) is (an estimate of) its standard deviation.

The sample variance with divisor nn is biased and consistent; with divisor n−1n - 1 it is unbiased. Consistency is what matters with the sample sizes of finance, and so does the standard error: an estimate without one is an anecdote.

Definition 11.2 (Likelihood function, maximum likelihood estimator, score, Fisher information)

For independent observations with density f(x;θ)f(x; \theta), the likelihood function is θ↦∏if(Xi;θ)\theta \mapsto \prod_if(X_i; \theta) and the log-likelihood ℓn(θ)=∑iln⁡f(Xi;θ)\ell_n(\theta) = \sum_i\ln f(X_i; \theta). The maximum likelihood estimator (MLE) maximises it. The score function is ∇θln⁡f(X;θ)\nabla_\theta\ln f(X; \theta), and the Fisher information is I(θ)=E[∇ln⁡f ∇ln⁡f⊤]=−E[∇2ln⁡f]\mathcal I(\theta) = \E[\nabla\ln f\,\nabla\ln f^\top] = -\E[\nabla^2\ln f] per observation.

Theorem 11.3 (Asymptotics of maximum likelihood)

Under regularity conditions (identifiability, smoothness, a true value in the interior), the MLE is consistent and

n(θ^n−θ0)→dN(0,I(θ0)−1).\sqrt n(\hat\theta_n - \theta_0) \xrightarrow{d} \mathcal N\bigl(0, \mathcal I(\theta_0)^{-1}\bigr).

No unbiased estimator has a smaller variance than I(θ0)−1/n\mathcal I(\theta_0)^{-1}/n (the Cramér–Rao bound), so the MLE is asymptotically efficient.

Partial proof. The score at θ0\theta_0 has mean zero and variance I\mathcal I, so n−1/2∇ℓn(θ0)→N(0,I)n^{-1/2}\nabla\ell_n(\theta_0) \to \mathcal N(0, \mathcal I) by the central limit theorem. Expanding 0=∇ℓn(θ^n)=∇ℓn(θ0)+∇2ℓn(θ~)(θ^n−θ0)0 = \nabla\ell_n(\hat\theta_n) = \nabla\ell_n(\theta_0) + \nabla^2\ell_n(\tilde\theta)(\hat\theta_n - \theta_0) and using −n−1∇2ℓn→I-n^{-1}\nabla^2\ell_n \to \mathcal I gives n(θ^n−θ0)=I−1n−1/2∇ℓn(θ0)+oP(1)\sqrt n(\hat\theta_n - \theta_0) = \mathcal I^{-1}n^{-1/2}\nabla\ell_n(\theta_0) + o_{\P}(1). Consistency and the Cramér–Rao bound are in Casella and Berger (2002). ∎

Example 11.4 (The rate of arrivals)

For nn exponential waiting times with rate λ\lambda, ℓn(λ)=nln⁡λ−λ∑ti\ell_n(\lambda) = n\ln\lambda - \lambda\sum t_i, so λ^=n/∑ti\hat\lambda = n/\sum t_i, I(λ)=1/λ2\mathcal I(\lambda) = 1/\lambda^2, and se(λ^)=λ^/n\mathrm{se}(\hat\lambda) = \hat\lambda/\sqrt n. A hundred gaps summing to 50 seconds give 2±0.22 \pm 0.2 arrivals per second. The Hawkes fits of chapter 7 are the same computation with a harder likelihood.

11.2 Misspecification and the sandwich

Models are approximations; the question is what maximum likelihood estimates when the density is wrong.

Definition 11.5 (Kullback–Leibler divergence, quasi-maximum likelihood)

The Kullback–Leibler divergence of a density ff from the true density pp is DKL(p ∥ f)=Ep[ln⁡(p(X)/f(X))]≥0D_{\mathrm{KL}}(p\,\|\,f) = \E_p[\ln(p(X)/f(X))] \ge 0, zero only if f=pf = p. Quasi-maximum likelihood maximises a likelihood that is not believed to be the true one.

Proposition 11.6 (What a wrong likelihood estimates)

Under regularity conditions the quasi-MLE converges to the pseudo-true value θ∗\theta^* that minimises DKL(p ∥ fθ)D_{\mathrm{KL}}(p\,\|\,f_\theta), and n(θ^n−θ∗)→dN(0,A−1BA−1)\sqrt n(\hat\theta_n - \theta^*) \xrightarrow{d} \mathcal N(0, A^{-1}BA^{-1}) with A=−E[∇2ln⁡fθ∗]A = -\E[\nabla^2\ln f_{\theta^*}] and B=Var⁡(∇ln⁡fθ∗)B = \Var(\nabla\ln f_{\theta^*}).

Proof. n−1ℓn(θ)→Ep[ln⁡fθ]=Ep[ln⁡p]−DKL(p ∥ fθ)n^{-1}\ell_n(\theta) \to \E_p[\ln f_\theta] = \E_p[\ln p] - D_{\mathrm{KL}}(p\,\|\,f_\theta), maximised at θ∗\theta^*. The expansion of the previous proof holds with I\mathcal I replaced by AA in the Hessian and by BB in the score’s variance; they coincide only when the model is right (the information equality). ∎

Definition 11.7 (Sandwich variance)

The sandwich variance of an estimator defined by an estimating equation is A−1BA−1A^{-1}BA^{-1}: the bread AA, the expected derivative of the estimating equation, around the meat BB, the variance of the equation itself. Estimated by the sample Hessian and the outer product of the per-observation scores, it is the robust (Huber–White) variance.

Fit a normal model to 2 000 draws of a Student tt with 6 degrees of freedom. The estimate of the log standard deviation is still consistent for the log standard deviation, but the Hessian standard error, 0.0158, assumes normal tails; the sandwich gives 0.0230, and over 4 000 simulated samples the actual spread is 0.0237 (Figure 11.1). With fatter tails still (4 degrees of freedom, infinite fourth moment) the variance of the sample variance does not exist, and no standard error is right.

The log standard deviation estimated by a normal likelihood from 2 000 Student-t_6 observations: its actual sampling distribution (steps) against the normal approximations with the Hessian standard error (dashed), too narrow, and with the sandwich (solid). Data: the chapter’s tutorial, seeded.
Figure 11.1. The log standard deviation estimated by a normal likelihood from 2 000 Student-t6t_6 observations: its actual sampling distribution (steps) against the normal approximations with the Hessian standard error (dashed), too narrow, and with the sandwich (solid). Data: the chapter’s tutorial, seeded.

11.3 M-estimators and the method of moments

Definition 11.8 (M-estimator, method of moments, generalised method of moments)

An M-estimator maximises (or zeroes the derivative of) a sample average n−1∑im(Xi;θ)n^{-1}\sum_im(X_i; \theta); maximum likelihood takes m=ln⁡fm = \ln f, least squares m=−(y−x⊤θ)2m = -(y - x^\top\theta)^2. The method of moments solves n−1∑ig(Xi)=Eθ[g(X)]n^{-1}\sum_ig(X_i) = \E_\theta[g(X)] for as many moments as parameters. The generalised method of moments (GMM) minimises gˉn(θ)⊤Wgˉn(θ)\bar g_n(\theta)^\top W\bar g_n(\theta) for gˉn(θ)=n−1∑ig(Xi;θ)\bar g_n(\theta) = n^{-1}\sum_ig(X_i; \theta) with more moment conditions than parameters and a weight matrix WW.

All three have the sandwich asymptotics of Proposition 11.6, with AA the derivative of the estimating equations and BB their long-run variance; GMM is efficient when WW is the inverse of that variance (Hansen, 1982), and the minimised objective then tests the overidentifying conditions. For an Ornstein–Uhlenbeck process sampled daily, the method of moments on the first autocorrelation, κ^=−ln⁡ρ^(1)/Δt\hat\kappa = -\ln\hat\rho(1)/\Delta t, coincides with the conditional maximum likelihood of the autoregression (chapter 17), and both inherit its small-sample bias.

11.4 The delta method and the Sharpe ratio

Method 11.9 (Delta method)

If n(θ^n−θ0)→dN(0,V)\sqrt n(\hat\theta_n - \theta_0) \xrightarrow{d} \mathcal N(0, V) and gg is differentiable at θ0\theta_0, then n(g(θ^n)−g(θ0))→dN(0,∇g⊤V∇g)\sqrt n(g(\hat\theta_n) - g(\theta_0)) \xrightarrow{d} \mathcal N(0, \nabla g^\top V\nabla g). The delta method linearises a smooth function of an estimate and propagates its variance.

Definition 11.10 (Sharpe ratio)

The Sharpe ratio of a strategy is its expected excess return per unit of standard deviation, SR=μ/σ\mathrm{SR} = \mu/\sigma, per period; annualised by periods a year\sqrt{\text{periods a year}} when returns are independent across periods. Its estimate is SR^=xˉ/s\widehat{\mathrm{SR}} = \bar x/s.

Proposition 11.11 (Standard error of the Sharpe ratio)

For independent normal returns, n(SR^−SR)→dN(0,1+12SR2)\sqrt n(\widehat{\mathrm{SR}} - \mathrm{SR}) \xrightarrow{d} \mathcal N(0, 1 + \tfrac12\mathrm{SR}^2) per period.

Proof. (xˉ,s2)(\bar x, s^2) are independent with asymptotic variances σ2\sigma^2 and 2σ42\sigma^4. The gradient of μ/v\mu/\sqrt v is (1/σ,−μ/(2σ3))(1/\sigma, -\mu/(2\sigma^3)), so the delta method gives 1+μ2⋅2σ4/(4σ6)=1+12SR21 + \mu^2\cdot2\sigma^4/(4\sigma^6) = 1 + \tfrac12\mathrm{SR}^2. ∎

Annualised, the standard error is about (1+SR12/2)/years\sqrt{(1 + \mathrm{SR}_1^2/2)/\text{years}}, since the per-day Sharpe ratio SR1\mathrm{SR}_1 is small: five years estimate an annual Sharpe ratio to ±0.45\pm 0.45, whatever its size. The researcher’s strategy has SR^=1.51\widehat{\mathrm{SR}} = 1.51 with that standard error under independence (Lo, 2002, derives the general case). With non-normal returns BB changes; with autocorrelated returns, the annualisation by 252\sqrt{252} and the variance both change, which is the next section’s subject: the HAC version of the same delta method gives 0.87.

11.5 Serial correlation: HAC variances

The mean of a stationary series with autocovariances γ(k)\gamma(k) has

Var⁡(xˉ)=1n∑∣k∣<n(1−∣k∣n)γ(k)≈1n∑kγ(k)=long-run variancen.\Var(\bar x) = \frac1n\sum_{|k|<n}\Bigl(1 - \frac{|k|}n\Bigr)\gamma(k) \approx \frac1n\sum_k\gamma(k) = \frac{\text{long-run variance}}n.

The researcher’s P&L is the average of five overlapping five-day positions, so its autocorrelations are 1−k/51 - k/5 for k≤4k \le 4 (Figure 11.2), and the long-run variance is 1+2(0.8+0.6+0.4+0.2)=51 + 2(0.8 + 0.6 + 0.4 + 0.2) = 5 times the variance: the naive standard error is 5=2.24\sqrt5 = 2.24 times too small.

Definition 11.12 (HAC estimator)

A heteroskedasticity- and autocorrelation-consistent (HAC estimator) estimates a long-run variance by weighting sample autocovariances; the Newey–West estimator γ^(0)+2∑k=1L(1−kL+1)γ^(k)\hat\gamma(0) + 2\sum_{k=1}^L(1 - \frac k{L+1})\hat\gamma(k) uses Bartlett weights, which keep it nonnegative, with a bandwidth LL growing slowly with nn (a common rule is L=⌊4(n/100)2/9⌋L = \lfloor4(n/100)^{2/9}\rfloor).

The researcher’s five years of daily P&L. Left: sample autocorrelations and the theoretical 1 - k/5 of five overlapping positions. Right: the Newey–West standard error of the mean P&L against the number of lags: 1.19 bp with none (the iid formula), 2.32 at the rule’s seven lags, against a true 2.65. Data: the chapter’s tutorial, seeded.
Figure 11.2. The researcher’s five years of daily P&L. Left: sample autocorrelations and the theoretical 1−k/51 - k/5 of five overlapping positions. Right: the Newey–West standard error of the mean P&L against the number of lags: 1.19 bp with none (the iid formula), 2.32 at the rule’s seven lags, against a true 2.65. Data: the chapter’s tutorial, seeded.

With seven lags, Newey–West gives 2.32 basis points and a tt-statistic of 1.73: the colleague’s number. It still understates the true 2.65, because the estimated autocovariances are biased toward zero in a sample of 1 260 days. Over 2 000 simulated histories of the strategy, nominal 95% intervals built on the iid standard error contain the true mean 62.5% of the time; with Newey–West, 92.2% with seven lags and 93.5% with twenty (Figure 11.3). A HAC standard error is necessary, not sufficient; the model-based alternative, when the overlap is known, uses the factor 5 directly.

Coverage of nominal 95% confidence intervals for the mean daily P&L of the overlapping strategy, over 2 000 simulated five-year histories: the iid standard error covers 62.5% of the time, Newey–West 92–94%. Data: the chapter’s tutorial, seeded.
Figure 11.3. Coverage of nominal 95% confidence intervals for the mean daily P&L of the overlapping strategy, over 2 000 simulated five-year histories: the iid standard error covers 62.5% of the time, Newey–West 92–94%. Data: the chapter’s tutorial, seeded.

The same arithmetic tells a researcher how long to wait. An annual Sharpe ratio SR\mathrm{SR} estimated from independent daily returns reaches a tt-statistic of 2 after about 4/SR24/\mathrm{SR}^2 years: one year at a Sharpe ratio of 2, four at 1, sixteen at 0.5. If the long-run variance is mm times the variance, as with the overlapping positions (m=5m = 5), the naive annualisation overstates the Sharpe ratio by m\sqrt m and the wait is mm times longer (Figure 11.4): the researcher’s naive 1.51 needs 1.8 years, the honest reading 8.8. Five years of history cannot separate an honest Sharpe ratio of 0.68 from zero.

Years of daily data before a strategy’s Sharpe ratio reaches a t-statistic of 2, against the Sharpe ratio annualised by √252: with independent returns and with the researcher’s five overlapping positions (long-run variance five times the variance). Dots: the researcher’s 1.51. Data: the chapter’s tutorial module.
Figure 11.4. Years of daily data before a strategy’s Sharpe ratio reaches a tt-statistic of 2, against the Sharpe ratio annualised by 252\sqrt{252}: with independent returns and with the researcher’s five overlapping positions (long-run variance five times the variance). Dots: the researcher’s 1.51. Data: the chapter’s tutorial module.

11.6 Tutorial: the tt-statistic that halved

Goal. Rebuild the researcher’s and the colleague’s standard errors, measure how often each is right, and see the sandwich correct a misspecified likelihood. End state: Figures 11.1, 11.2 and 11.3 and the tt-statistics 3.38 and 1.73.

  1. The long-run variance with Bartlett weights, and the standard error of a mean.

    def long_run_variance(x, lags: int | None = None) -> float:
        """sum over |k| <= L of (1 - |k| / (L + 1)) gamma(k): Newey-West, always nonnegative."""
        x = np.asarray(x, dtype=float)
        n = x.size
        L = nw_lags(n) if lags is None else lags
        d = x - x.mean()
        lrv = d @ d / n
        for k in range(1, L + 1):
            lrv += 2 * (1 - k / (L + 1)) * (d[k:] @ d[:-k]) / n
        return float(lrv)
    
    
    def mean_se(x, kind: str = "iid", lags: int | None = None) -> float:
        x = np.asarray(x, dtype=float)
        n = x.size
        if kind == "iid":
            return float(x.std(ddof=1) / math.sqrt(n))
        return float(math.sqrt(long_run_variance(x, lags) / n))
    Listing 11.1. Newey–West long-run variance and the standard error of a mean. code/firm/estim/firm_estim.py
  2. The Sharpe ratio with its iid and HAC standard errors, by the delta method.

    def sharpe(x, periods: int = 252, lags: int | None = None) -> dict:
        """Annualised Sharpe ratio sqrt(periods) mean / sd, its iid standard error sqrt((1 + SR_1^2 / 2) / n)
        (per period, Lo 2002, normal returns) scaled by sqrt(periods), and a HAC standard error by the
        delta method applied to the long-run covariance of (x, x^2)."""
        x = np.asarray(x, dtype=float)
        n = x.size
        m, s = x.mean(), x.std(ddof=1)
        sr1 = m / s
        se_iid = math.sqrt((1 + 0.5 * sr1**2) / n)
        z = np.column_stack([x - m, (x - m) ** 2 - s**2])
        L = nw_lags(n) if lags is None else lags
        S = z.T @ z / n
        for k in range(1, L + 1):
            g = z[k:].T @ z[:-k] / n
            S += (1 - k / (L + 1)) * (g + g.T)
        grad = np.array([1 / s, -m / (2 * s**3)])               # d SR / d (mean, variance)
        se_hac = math.sqrt(grad @ S @ grad / n)
        root = math.sqrt(periods)
        return {"sr": root * sr1, "se_iid": root * se_iid, "se_hac": root * se_hac}
    Listing 11.2. The Sharpe ratio and its two standard errors. code/firm/estim/firm_estim.py
  3. Run problem(), coverage(), misspecified_fit() and fig_estimation.py.

What to change next. Give the P&L volatility clustering and compare the iid, White and Newey–West standard errors; estimate the mean from non-overlapping five-day returns and compare the efficiency.

11.7 Build: the estimation toolkit

Purpose. Every estimate the miniature firm reports (a mean return, a Sharpe ratio, a fitted parameter) carries a standard error computed here, robust by default.

Interface. minimize(f, x0); gradient, hessian; mle(loglik_obs, x0) returning the estimate and its Hessian, outer-product and sandwich covariances; long_run_variance(x, lags); nw_lags(n); mean_se(x, kind); sharpe(x, periods); delta_method(g, theta, cov).

Rules. Per-observation log-likelihoods (so scores can be formed); the sandwich is reported whenever a likelihood is maximised; HAC bandwidth by the rule unless stated; no standard error without its method named in the output.

Acceptance tests. code/firm/estim/tests/: the MLE of normal and exponential samples and their standard errors; the sandwich equals the Hessian variance for a correct model; Newey–West recovers the long-run variance of a moving average; the Sharpe standard error matches simulation.

Stretch. Andrews’ automatic bandwidth; GMM with an efficient weight matrix and the overidentification test.

Sources and further reading

  • P. J. Huber, “The behavior of maximum likelihood estimates under nonstandard conditions”, Fifth Berkeley Symposium, 1967.
  • H. White, “Maximum likelihood estimation of misspecified models”, Econometrica 50, 1982.
  • W. K. Newey and K. D. West, “A simple, positive semi-definite, heteroskedasticity and autocorrelation consistent covariance matrix”, Econometrica 55, 1987.
  • L. P. Hansen, “Large sample properties of generalized method of moments estimators”, Econometrica 50, 1982.
  • A. W. Lo, “The statistics of Sharpe ratios”, Financial Analysts Journal 58, 2002.
  • G. Casella and R. L. Berger, Statistical Inference, Duxbury, 2nd ed., 2002.

11.8 Exercises

Exercise 11.1 ★

A hundred waiting times between trades sum to 50 seconds. Estimate the arrival rate and its standard error.

Solution

Solution of Exercise 11.1.

λ^=100/50=2\hat\lambda = 100/50 = 2 per second; se=λ^/100=0.2\mathrm{se} = \hat\lambda/\sqrt{100} = 0.2.

Exercise 11.2 ★

A strategy’s annualised Sharpe ratio is estimated at 1.0 from four years of independent daily returns. What is its standard error?

Solution

Solution of Exercise 11.2.

Per day SR1=1/252=0.063\mathrm{SR}_1 = 1/\sqrt{252} = 0.063 and se=(1+0.002)/1 008=0.0315\mathrm{se} = \sqrt{(1 + 0.002)/1\,008} = 0.0315; annualised, 0.0315252=0.500.0315\sqrt{252} = 0.50.

Exercise 11.3 ★

Is the sample variance with divisor nn unbiased? Consistent?

Solution

Solution of Exercise 11.3.

Biased: its mean is (n−1)σ2/n(n - 1)\sigma^2/n. Consistent: the bias and the variance vanish as n→∞n \to \infty.

Exercise 11.4 ★★

Daily P&L follows an AR(1) with coefficient 0.2. By what factor does the naive standard error of the mean understate the truth?

Solution

Solution of Exercise 11.4.

The long-run variance is σ2(1+ρ)/(1−ρ)=1.5σ2\sigma^2(1 + \rho)/(1 - \rho) = 1.5\sigma^2, so the naive standard error is 1.5=1.22\sqrt{1.5} = 1.22 times too small.

Exercise 11.5 ★★

A signal was right on 110 of 200 trades. Using the Fisher information of a Bernoulli variable, give the standard error of the hit rate, and the tt-statistic against 50%.

Solution

Solution of Exercise 11.5.

I(p)=1/(p(1−p))\mathcal I(p) = 1/(p(1 - p)), so se=0.55×0.45/200=0.035\mathrm{se} = \sqrt{0.55 \times 0.45/200} = 0.035 and t=0.05/0.035=1.42t = 0.05/0.035 = 1.42: not significant.

Exercise 11.6 ★★

For nn independent normal observations, use the delta method to find the standard error of ln⁡σ^\ln\hat\sigma, and evaluate it at n=2 000n = 2\,000.

Solution

Solution of Exercise 11.6.

Var⁡(σ^2)≈2σ4/n\Var(\hat\sigma^2) \approx 2\sigma^4/n and ln⁡σ^=12ln⁡σ^2\ln\hat\sigma = \tfrac12\ln\hat\sigma^2, so Var⁡(ln⁡σ^)≈14⋅2σ4/(nσ4)=1/(2n)\Var(\ln\hat\sigma) \approx \tfrac14 \cdot 2\sigma^4/(n\sigma^4) = 1/(2n): se=1/4 000=0.0158\mathrm{se} = 1/\sqrt{4\,000} = 0.0158, the Hessian value of the chapter, correct for normal data only.

Exercise 11.7 ★★★

Coding. Over 2 000 simulated histories of the overlapping strategy, measure the coverage of 95% intervals built on the iid, Newey–West(7) and Newey–West(20) standard errors.

Solution

Solution of Exercise 11.7.

62.5% with the iid standard error, 92.2% with Newey–West at seven lags and 93.5% at twenty.

Exercise 11.8 ★★★

Find the flaw. “Each day we record our strategy’s trailing 20-day return. Over three years the Sharpe ratio of these daily figures, annualised by 252\sqrt{252}, is 1.4, with a standard error of 1/3=0.581/\sqrt3 = 0.58: significant at the 5% level.”

Solution

Solution of Exercise 11.8.

Consecutive trailing 20-day returns share 19 days: the series is heavily autocorrelated, so its standard deviation understates the uncertainty of its mean, and there are only about 3×252/20=383 \times 252/20 = 38 independent observations. Annualising a 20-day return by 252\sqrt{252} also inflates the Sharpe ratio by 20\sqrt{20}: use non-overlapping 20-day returns annualised by 252/20\sqrt{252/20}, or daily P&L, with a HAC standard error.

11.9 Problem: The tt-Statistic That Halved

Problem 11.1

Weekend problem — overlapping positions and the standard error of a mean

A strategy opens a position every day and holds it five days, so each day’s P&L is the average of five overlapping positions. Over five years (1 260 days) its daily P&L, in basis points of capital, is the chapter’s seeded history.

Part I — The naive view.

  1. What are the sample mean and standard deviation?
  2. What are the iid standard error and tt-statistic?
  3. What are the annualised Sharpe ratio and its iid standard error?
  4. What are the first four sample autocorrelations, against theory?
  5. Why are consecutive days correlated?

Part II — Robust standard errors.

  1. What is the long-run variance factor of the P&L, and the true standard error of the mean?
  2. How many lags does the rule give, and what do Newey–West and the tt-statistic become?
  3. With twenty lags?
  4. What is the HAC standard error of the Sharpe ratio?
  5. How often do iid 95% intervals contain the true mean, over 2 000 simulated histories?

Part III — What the data can say.

  1. Knowing the overlap, what model-based standard error would you use?
  2. What tt-statistic should the researcher expect, on average, with a true mean of 4 bp?
  3. How many years would a tt-statistic of 3 need, on average?
  4. Why is Newey–West still slightly too small?
  5. What would evaluating non-overlapping five-day P&L change?

Part IV — Judgement.

  1. What other dependence in daily P&L would the iid formula miss?
  2. The researcher tried twenty variants before this one. What else must change in the evaluation?
  3. What should the report on the strategy state?
  4. State the named result: the naive and the robust tt-statistics.
  5. In one sentence: when is the iid standard error of a mean wrong?
Solution

Solution of Problem 11.1.

1. Mean 4.03 bp, standard deviation 42.3 bp. 2. 42.3/1 260=1.1942.3/\sqrt{1\,260} = 1.19 bp; t=3.38t = 3.38. 3. SR^=1.51\widehat{\mathrm{SR}} = 1.51, iid standard error 0.45. 4. 0.80, 0.59, 0.37, 0.16 against 0.8, 0.6, 0.4, 0.2. 5. Consecutive days hold four fifths of the same positions. 6. 1+2(0.8+0.6+0.4+0.2)=51 + 2(0.8 + 0.6 + 0.4 + 0.2) = 5; true standard error 425/1 260=2.6542\sqrt{5/1\,260} = 2.65 bp. 7. ⌊4×12.62/9⌋=7\lfloor4 \times 12.6^{2/9}\rfloor = 7 lags: 2.32 bp and t=1.73t = 1.73. 8. 2.33 bp: the Bartlett weights beyond lag four add almost nothing. 9. 0.87, about twice the iid value. 10. 62.5% of the time. 11. σ^5/n=42.35/1 260=2.66\hat\sigma\sqrt{5/n} = 42.3\sqrt{5/1\,260} = 2.66 bp. 12. 4/2.65=1.514/2.65 = 1.51. 13. n=(3×425/4)2=4 960n = (3 \times 42\sqrt5/4)^2 = 4\,960 days, about 20 years. 14. Sample autocovariances are biased toward zero in finite samples, and the Bartlett weights shrink the ones that matter (lags 1 to 4) further. 15. Non-overlapping returns are independent, so the iid formula is right for them, at a small loss of efficiency; the estimate of the mean is almost unchanged. 16. Volatility clustering (heteroskedasticity), regime changes and slow trends in the signal’s edge. 17. The significance threshold: twenty variants tried make a tt of 1.7, or even 3.4, far less surprising (chapter 12). 18. The mean, its HAC standard error and bandwidth, the tt-statistic, the Sharpe ratio with its robust standard error, and the number of variants tried. 19. Named result: the tt-statistic that halved: the iid tt-statistic of 3.38 becomes 1.73 with Newey–West standard errors (and 1.51 with the true long-run variance), because overlapping positions make the long-run variance five times the variance. 20. When the observations are dependent or not identically distributed, as with overlapping holdings.

11.10 Interview questions

Interview question 11.1 ★ researcher, mle

What is maximum likelihood, and why is it the default estimator?

Solution

Solution of Interview question 11.1.

Choose the parameter that makes the observed data most probable. It is consistent, asymptotically normal and efficient (it reaches the Cramér–Rao bound), invariant to reparametrisation, and gives standard errors from the curvature of the log-likelihood.

What the interviewer is looking for: the definition and the efficiency property.

Interview question 11.2 ★★ researcher, trader

A strategy shows a Sharpe ratio of 1.0 over one year. How confident are you that it is positive?

Solution

Solution of Interview question 11.2.

With independent daily returns its standard error is about 1/years=11/\sqrt{\text{years}} = 1: the estimate is one standard error from zero, a one-sided probability of about 84% that the true value is positive under a flat prior, far from proof.

What the interviewer is looking for: se≈1/T\mathrm{se} \approx 1/\sqrt T for annual Sharpe ratios.

Interview question 11.3 ★★ researcher

What is a sandwich variance, and when do you need it?

Solution

Solution of Interview question 11.3.

A−1BA−1A^{-1}BA^{-1}, with AA the expected Hessian of the objective and BB the variance of its gradient. When the likelihood is misspecified, or with least squares under heteroskedasticity, A≠BA \ne B and the usual A−1A^{-1} is wrong; the sandwich is not.

What the interviewer is looking for: bread and meat, and the information equality.

Interview question 11.4 ★★ researcher

You compute monthly returns every day from overlapping windows and regress them on a signal. What is wrong with the usual standard errors, and how do you fix them?

Solution

Solution of Interview question 11.4.

Overlapping windows make the dependent variable autocorrelated up to the window length, so the residuals are too: the usual standard errors are too small by roughly the square root of the overlap. Use Newey–West with at least as many lags as the overlap (or Hansen–Hodrick), or non-overlapping observations.

What the interviewer is looking for: overlap-induced autocorrelation and HAC.

Interview question 11.5 ★★ researcher, mle

What is the Fisher information, and what does the Cramér–Rao bound say?

Solution

Solution of Interview question 11.5.

The variance of the score, equal to minus the expected Hessian of the log-likelihood: the curvature, the information one observation carries about θ\theta. Cramér–Rao: an unbiased estimator’s variance is at least I−1/n\mathcal I^{-1}/n.

What the interviewer is looking for: curvature equals information.

Interview question 11.6 ★★★ researcher

If your likelihood is wrong, what does its maximiser estimate?

Solution

Solution of Interview question 11.6.

The parameter θ∗\theta^* that minimises the Kullback–Leibler divergence from the true distribution to the model: the best approximation within the model. Its standard errors are the sandwich, not the inverse Hessian.

What the interviewer is looking for: pseudo-true value and robust variance.

Terms defined in this chapter

See all 2333 terms in the glossary