Quantitative Methods · Methods
11Estimation
A researcher presents a strategy: over five years its daily P&L averaged 4.0 basis points with a standard error of 1.2, a -statistic of 3.4, comfortably significant. A colleague notices that each day’s position is held for five days, so consecutive days share four fifths of their positions, and recomputes the standard error allowing for the autocorrelation that creates: 2.3. The -statistic falls to 1.7, and the strategy is no longer distinguishable from luck. Nothing was wrong with the average; the formula for its uncertainty assumed independent days. This chapter is about estimators and their uncertainty: maximum likelihood and its asymptotics, what happens when the model is wrong, M-estimators and moments, the delta method that gives the Sharpe ratio its standard error, and the heteroskedasticity- and autocorrelation-consistent variances that repaired the researcher’s -statistic.
11.1 Maximum likelihood
Definition 11.1 (Estimator, unbiased, consistent, standard error)
An estimator of a parameter is a function of the data . It is an unbiased estimator if , and a consistent estimator if as . Its standard error is (an estimate of) its standard deviation.
The sample variance with divisor is biased and consistent; with divisor it is unbiased. Consistency is what matters with the sample sizes of finance, and so does the standard error: an estimate without one is an anecdote.
Definition 11.2 (Likelihood function, maximum likelihood estimator, score, Fisher information)
For independent observations with density , the likelihood function is and the log-likelihood . The maximum likelihood estimator (MLE) maximises it. The score function is , and the Fisher information is per observation.
Theorem 11.3 (Asymptotics of maximum likelihood)
Under regularity conditions (identifiability, smoothness, a true value in the interior), the MLE is consistent and
No unbiased estimator has a smaller variance than (the Cramér–Rao bound), so the MLE is asymptotically efficient.
Partial proof. The score at has mean zero and variance , so by the central limit theorem. Expanding and using gives . Consistency and the Cramér–Rao bound are in Casella and Berger (2002). ∎
Example 11.4 (The rate of arrivals)
For exponential waiting times with rate , , so , , and . A hundred gaps summing to 50 seconds give arrivals per second. The Hawkes fits of chapter 7 are the same computation with a harder likelihood.
11.2 Misspecification and the sandwich
Models are approximations; the question is what maximum likelihood estimates when the density is wrong.
Definition 11.5 (Kullback–Leibler divergence, quasi-maximum likelihood)
The Kullback–Leibler divergence of a density from the true density is , zero only if . Quasi-maximum likelihood maximises a likelihood that is not believed to be the true one.
Proposition 11.6 (What a wrong likelihood estimates)
Under regularity conditions the quasi-MLE converges to the pseudo-true value that minimises , and with and .
Proof. , maximised at . The expansion of the previous proof holds with replaced by in the Hessian and by in the score’s variance; they coincide only when the model is right (the information equality). ∎
Definition 11.7 (Sandwich variance)
The sandwich variance of an estimator defined by an estimating equation is : the bread , the expected derivative of the estimating equation, around the meat , the variance of the equation itself. Estimated by the sample Hessian and the outer product of the per-observation scores, it is the robust (Huber–White) variance.
Fit a normal model to 2 000 draws of a Student with 6 degrees of freedom. The estimate of the log standard deviation is still consistent for the log standard deviation, but the Hessian standard error, 0.0158, assumes normal tails; the sandwich gives 0.0230, and over 4 000 simulated samples the actual spread is 0.0237 (Figure 11.1). With fatter tails still (4 degrees of freedom, infinite fourth moment) the variance of the sample variance does not exist, and no standard error is right.
11.3 M-estimators and the method of moments
Definition 11.8 (M-estimator, method of moments, generalised method of moments)
An M-estimator maximises (or zeroes the derivative of) a sample average ; maximum likelihood takes , least squares . The method of moments solves for as many moments as parameters. The generalised method of moments (GMM) minimises for with more moment conditions than parameters and a weight matrix .
All three have the sandwich asymptotics of Proposition 11.6, with the derivative of the estimating equations and their long-run variance; GMM is efficient when is the inverse of that variance (Hansen, 1982), and the minimised objective then tests the overidentifying conditions. For an Ornstein–Uhlenbeck process sampled daily, the method of moments on the first autocorrelation, , coincides with the conditional maximum likelihood of the autoregression (chapter 17), and both inherit its small-sample bias.
11.4 The delta method and the Sharpe ratio
Method 11.9 (Delta method)
If and is differentiable at , then . The delta method linearises a smooth function of an estimate and propagates its variance.
Definition 11.10 (Sharpe ratio)
The Sharpe ratio of a strategy is its expected excess return per unit of standard deviation, , per period; annualised by when returns are independent across periods. Its estimate is .
Proposition 11.11 (Standard error of the Sharpe ratio)
For independent normal returns, per period.
Proof. are independent with asymptotic variances and . The gradient of is , so the delta method gives . ∎
Annualised, the standard error is about , since the per-day Sharpe ratio is small: five years estimate an annual Sharpe ratio to , whatever its size. The researcher’s strategy has with that standard error under independence (Lo, 2002, derives the general case). With non-normal returns changes; with autocorrelated returns, the annualisation by and the variance both change, which is the next section’s subject: the HAC version of the same delta method gives 0.87.
11.5 Serial correlation: HAC variances
The mean of a stationary series with autocovariances has
The researcher’s P&L is the average of five overlapping five-day positions, so its autocorrelations are for (Figure 11.2), and the long-run variance is times the variance: the naive standard error is times too small.
Definition 11.12 (HAC estimator)
A heteroskedasticity- and autocorrelation-consistent (HAC estimator) estimates a long-run variance by weighting sample autocovariances; the Newey–West estimator uses Bartlett weights, which keep it nonnegative, with a bandwidth growing slowly with (a common rule is ).
With seven lags, Newey–West gives 2.32 basis points and a -statistic of 1.73: the colleague’s number. It still understates the true 2.65, because the estimated autocovariances are biased toward zero in a sample of 1 260 days. Over 2 000 simulated histories of the strategy, nominal 95% intervals built on the iid standard error contain the true mean 62.5% of the time; with Newey–West, 92.2% with seven lags and 93.5% with twenty (Figure 11.3). A HAC standard error is necessary, not sufficient; the model-based alternative, when the overlap is known, uses the factor 5 directly.
The same arithmetic tells a researcher how long to wait. An annual Sharpe ratio estimated from independent daily returns reaches a -statistic of 2 after about years: one year at a Sharpe ratio of 2, four at 1, sixteen at 0.5. If the long-run variance is times the variance, as with the overlapping positions (), the naive annualisation overstates the Sharpe ratio by and the wait is times longer (Figure 11.4): the researcher’s naive 1.51 needs 1.8 years, the honest reading 8.8. Five years of history cannot separate an honest Sharpe ratio of 0.68 from zero.
11.6 Tutorial: the -statistic that halved
Goal. Rebuild the researcher’s and the colleague’s standard errors, measure how often each is right, and see the sandwich correct a misspecified likelihood. End state: Figures 11.1, 11.2 and 11.3 and the -statistics 3.38 and 1.73.
The long-run variance with Bartlett weights, and the standard error of a mean.
def long_run_variance(x, lags: int | None = None) -> float: """sum over |k| <= L of (1 - |k| / (L + 1)) gamma(k): Newey-West, always nonnegative.""" x = np.asarray(x, dtype=float) n = x.size L = nw_lags(n) if lags is None else lags d = x - x.mean() lrv = d @ d / n for k in range(1, L + 1): lrv += 2 * (1 - k / (L + 1)) * (d[k:] @ d[:-k]) / n return float(lrv) def mean_se(x, kind: str = "iid", lags: int | None = None) -> float: x = np.asarray(x, dtype=float) n = x.size if kind == "iid": return float(x.std(ddof=1) / math.sqrt(n)) return float(math.sqrt(long_run_variance(x, lags) / n))Listing 11.1. Newey–West long-run variance and the standard error of a mean. code/firm/estim/firm_estim.py The Sharpe ratio with its iid and HAC standard errors, by the delta method.
def sharpe(x, periods: int = 252, lags: int | None = None) -> dict: """Annualised Sharpe ratio sqrt(periods) mean / sd, its iid standard error sqrt((1 + SR_1^2 / 2) / n) (per period, Lo 2002, normal returns) scaled by sqrt(periods), and a HAC standard error by the delta method applied to the long-run covariance of (x, x^2).""" x = np.asarray(x, dtype=float) n = x.size m, s = x.mean(), x.std(ddof=1) sr1 = m / s se_iid = math.sqrt((1 + 0.5 * sr1**2) / n) z = np.column_stack([x - m, (x - m) ** 2 - s**2]) L = nw_lags(n) if lags is None else lags S = z.T @ z / n for k in range(1, L + 1): g = z[k:].T @ z[:-k] / n S += (1 - k / (L + 1)) * (g + g.T) grad = np.array([1 / s, -m / (2 * s**3)]) # d SR / d (mean, variance) se_hac = math.sqrt(grad @ S @ grad / n) root = math.sqrt(periods) return {"sr": root * sr1, "se_iid": root * se_iid, "se_hac": root * se_hac}Listing 11.2. The Sharpe ratio and its two standard errors. code/firm/estim/firm_estim.py - Run
problem(),coverage(),misspecified_fit()andfig_estimation.py.
What to change next. Give the P&L volatility clustering and compare the iid, White and Newey–West standard errors; estimate the mean from non-overlapping five-day returns and compare the efficiency.
11.7 Build: the estimation toolkit
Purpose. Every estimate the miniature firm reports (a mean return, a Sharpe ratio, a fitted parameter) carries a standard error computed here, robust by default.
Interface. minimize(f, x0); gradient, hessian; mle(loglik_obs, x0) returning the estimate and its Hessian, outer-product and sandwich covariances; long_run_variance(x, lags); nw_lags(n); mean_se(x, kind); sharpe(x, periods); delta_method(g, theta, cov).
Rules. Per-observation log-likelihoods (so scores can be formed); the sandwich is reported whenever a likelihood is maximised; HAC bandwidth by the rule unless stated; no standard error without its method named in the output.
Acceptance tests. code/firm/estim/tests/: the MLE of normal and exponential samples and their standard errors; the sandwich equals the Hessian variance for a correct model; Newey–West recovers the long-run variance of a moving average; the Sharpe standard error matches simulation.
Stretch. Andrews’ automatic bandwidth; GMM with an efficient weight matrix and the overidentification test.
Sources and further reading
- P. J. Huber, “The behavior of maximum likelihood estimates under nonstandard conditions”, Fifth Berkeley Symposium, 1967.
- H. White, “Maximum likelihood estimation of misspecified models”, Econometrica 50, 1982.
- W. K. Newey and K. D. West, “A simple, positive semi-definite, heteroskedasticity and autocorrelation consistent covariance matrix”, Econometrica 55, 1987.
- L. P. Hansen, “Large sample properties of generalized method of moments estimators”, Econometrica 50, 1982.
- A. W. Lo, “The statistics of Sharpe ratios”, Financial Analysts Journal 58, 2002.
- G. Casella and R. L. Berger, Statistical Inference, Duxbury, 2nd ed., 2002.
11.8 Exercises
Exercise 11.1 ★
A hundred waiting times between trades sum to 50 seconds. Estimate the arrival rate and its standard error.
Solution
Solution of Exercise 11.1.
per second; .
Exercise 11.2 ★
A strategy’s annualised Sharpe ratio is estimated at 1.0 from four years of independent daily returns. What is its standard error?
Solution
Solution of Exercise 11.2.
Per day and ; annualised, .
Exercise 11.3 ★
Is the sample variance with divisor unbiased? Consistent?
Solution
Solution of Exercise 11.3.
Biased: its mean is . Consistent: the bias and the variance vanish as .
Exercise 11.4 ★★
Daily P&L follows an AR(1) with coefficient 0.2. By what factor does the naive standard error of the mean understate the truth?
Solution
Solution of Exercise 11.4.
The long-run variance is , so the naive standard error is times too small.
Exercise 11.5 ★★
A signal was right on 110 of 200 trades. Using the Fisher information of a Bernoulli variable, give the standard error of the hit rate, and the -statistic against 50%.
Solution
Solution of Exercise 11.5.
, so and : not significant.
Exercise 11.6 ★★
For independent normal observations, use the delta method to find the standard error of , and evaluate it at .
Solution
Solution of Exercise 11.6.
and , so : , the Hessian value of the chapter, correct for normal data only.
Exercise 11.7 ★★★
Coding. Over 2 000 simulated histories of the overlapping strategy, measure the coverage of 95% intervals built on the iid, Newey–West(7) and Newey–West(20) standard errors.
Solution
Solution of Exercise 11.7.
62.5% with the iid standard error, 92.2% with Newey–West at seven lags and 93.5% at twenty.
Exercise 11.8 ★★★
Find the flaw. “Each day we record our strategy’s trailing 20-day return. Over three years the Sharpe ratio of these daily figures, annualised by , is 1.4, with a standard error of : significant at the 5% level.”
Solution
Solution of Exercise 11.8.
Consecutive trailing 20-day returns share 19 days: the series is heavily autocorrelated, so its standard deviation understates the uncertainty of its mean, and there are only about independent observations. Annualising a 20-day return by also inflates the Sharpe ratio by : use non-overlapping 20-day returns annualised by , or daily P&L, with a HAC standard error.
11.9 Problem: The -Statistic That Halved
Problem 11.1
Weekend problem — overlapping positions and the standard error of a mean
A strategy opens a position every day and holds it five days, so each day’s P&L is the average of five overlapping positions. Over five years (1 260 days) its daily P&L, in basis points of capital, is the chapter’s seeded history.
Part I — The naive view.
- What are the sample mean and standard deviation?
- What are the iid standard error and -statistic?
- What are the annualised Sharpe ratio and its iid standard error?
- What are the first four sample autocorrelations, against theory?
- Why are consecutive days correlated?
Part II — Robust standard errors.
- What is the long-run variance factor of the P&L, and the true standard error of the mean?
- How many lags does the rule give, and what do Newey–West and the -statistic become?
- With twenty lags?
- What is the HAC standard error of the Sharpe ratio?
- How often do iid 95% intervals contain the true mean, over 2 000 simulated histories?
Part III — What the data can say.
- Knowing the overlap, what model-based standard error would you use?
- What -statistic should the researcher expect, on average, with a true mean of 4 bp?
- How many years would a -statistic of 3 need, on average?
- Why is Newey–West still slightly too small?
- What would evaluating non-overlapping five-day P&L change?
Part IV — Judgement.
- What other dependence in daily P&L would the iid formula miss?
- The researcher tried twenty variants before this one. What else must change in the evaluation?
- What should the report on the strategy state?
- State the named result: the naive and the robust -statistics.
- In one sentence: when is the iid standard error of a mean wrong?
Solution
Solution of Problem 11.1.
1. Mean 4.03 bp, standard deviation 42.3 bp. 2. bp; . 3. , iid standard error 0.45. 4. 0.80, 0.59, 0.37, 0.16 against 0.8, 0.6, 0.4, 0.2. 5. Consecutive days hold four fifths of the same positions. 6. ; true standard error bp. 7. lags: 2.32 bp and . 8. 2.33 bp: the Bartlett weights beyond lag four add almost nothing. 9. 0.87, about twice the iid value. 10. 62.5% of the time. 11. bp. 12. . 13. days, about 20 years. 14. Sample autocovariances are biased toward zero in finite samples, and the Bartlett weights shrink the ones that matter (lags 1 to 4) further. 15. Non-overlapping returns are independent, so the iid formula is right for them, at a small loss of efficiency; the estimate of the mean is almost unchanged. 16. Volatility clustering (heteroskedasticity), regime changes and slow trends in the signal’s edge. 17. The significance threshold: twenty variants tried make a of 1.7, or even 3.4, far less surprising (chapter 12). 18. The mean, its HAC standard error and bandwidth, the -statistic, the Sharpe ratio with its robust standard error, and the number of variants tried. 19. Named result: the -statistic that halved: the iid -statistic of 3.38 becomes 1.73 with Newey–West standard errors (and 1.51 with the true long-run variance), because overlapping positions make the long-run variance five times the variance. 20. When the observations are dependent or not identically distributed, as with overlapping holdings.
11.10 Interview questions
Interview question 11.1 ★ researcher, mle
What is maximum likelihood, and why is it the default estimator?
Solution
Solution of Interview question 11.1.
Choose the parameter that makes the observed data most probable. It is consistent, asymptotically normal and efficient (it reaches the Cramér–Rao bound), invariant to reparametrisation, and gives standard errors from the curvature of the log-likelihood.
What the interviewer is looking for: the definition and the efficiency property.
Interview question 11.2 ★★ researcher, trader
A strategy shows a Sharpe ratio of 1.0 over one year. How confident are you that it is positive?
Solution
Solution of Interview question 11.2.
With independent daily returns its standard error is about : the estimate is one standard error from zero, a one-sided probability of about 84% that the true value is positive under a flat prior, far from proof.
What the interviewer is looking for: for annual Sharpe ratios.
Interview question 11.3 ★★ researcher
What is a sandwich variance, and when do you need it?
Solution
Solution of Interview question 11.3.
, with the expected Hessian of the objective and the variance of its gradient. When the likelihood is misspecified, or with least squares under heteroskedasticity, and the usual is wrong; the sandwich is not.
What the interviewer is looking for: bread and meat, and the information equality.
Interview question 11.4 ★★ researcher
You compute monthly returns every day from overlapping windows and regress them on a signal. What is wrong with the usual standard errors, and how do you fix them?
Solution
Solution of Interview question 11.4.
Overlapping windows make the dependent variable autocorrelated up to the window length, so the residuals are too: the usual standard errors are too small by roughly the square root of the overlap. Use Newey–West with at least as many lags as the overlap (or Hansen–Hodrick), or non-overlapping observations.
What the interviewer is looking for: overlap-induced autocorrelation and HAC.
Interview question 11.5 ★★ researcher, mle
What is the Fisher information, and what does the Cramér–Rao bound say?
Solution
Solution of Interview question 11.5.
The variance of the score, equal to minus the expected Hessian of the log-likelihood: the curvature, the information one observation carries about . Cramér–Rao: an unbiased estimator’s variance is at least .
What the interviewer is looking for: curvature equals information.
Interview question 11.6 ★★★ researcher
If your likelihood is wrong, what does its maximiser estimate?
Solution
Solution of Interview question 11.6.
The parameter that minimises the Kullback–Leibler divergence from the true distribution to the model: the best approximation within the model. Its standard errors are the sandwich, not the inverse Hessian.
What the interviewer is looking for: pseudo-true value and robust variance.