Quantitative Methods · Methods
15Robust Statistics and Heavy Tails
One bad print in the feed, the euro’s reference rate recorded as 1.5172 dollars instead of 1.1572 on one day in March 2026, multiplies the standard deviation of the last 250 daily EUR/USD returns by seven, from 0.34% to 2.43%, and moves their median absolute deviation by 1%. A risk system that estimates volatility with the standard deviation now reports a currency as volatile as a small biotechnology stock; one that uses the median absolute deviation has not noticed. The same data hold a second lesson: over 27 years EUR/USD moved more than four standard deviations in a day seventeen times, where a normal law expects fewer than one. This chapter is about estimators that bad data cannot capture and about the tails that good data do contain: the influence function and the breakdown point, winsorising and ranks, copulas and tail dependence, heavy tails and the extreme-value theory that extrapolates them, and the estimation of the tail index on real exchange rates.
15.1 Robust losses and the influence function
Definition 15.1 (Influence function, breakdown point)
For an estimator written as a functional of the data’s law, the influence function at is : the effect on the estimate of an infinitesimal contamination at the point (Hampel, 1974). The breakdown point is the smallest fraction of the sample that, replaced by arbitrary values, can carry the estimate arbitrarily far.
The mean has , unbounded; the variance has , worse. One observation placed far enough moves either wherever it likes.
Proposition 15.2 (Breakdown points of the mean and the median)
The sample mean has breakdown point , which tends to 0; the sample median has breakdown point , which tends to , the largest possible for an equivariant location estimator.
Proof. Replacing one observation by moves the mean by , unbounded in . The median stays between the smallest and largest of the uncontaminated observations as long as they are a majority; once half the sample is replaced, the contaminated values can be placed on both sides of the median position and carry it. No equivariant estimator can do better: with half the sample shifted by , the data are equally consistent with a sample and with its shift, and the estimate cannot track both. ∎
Definition 15.3 (Huber loss, median absolute deviation)
The Huber loss with threshold is for and otherwise; the Huber M-estimator of location minimises for a robust scale , so its estimating function is bounded (Huber, 1964). The median absolute deviation is , the constant making it consistent for the standard deviation of a normal law.
The Huber estimator is the mean near the centre and the median in the tails. With it keeps 95% of the mean’s efficiency when the data are normal (), and its influence function is , bounded by standard units. The MAD has breakdown point one half but is inefficient and assumes symmetry; the estimator of Rousseeuw and Croux (1993), a scaled lower quartile of all pairwise distances , keeps the breakdown point without the symmetry. Figure 15.1 shows what the influence functions mean for one bad print: as the error on one day’s log price grows from 0 to 27%, the 250-day standard deviation grows without limit, while the MAD and move once, by 1.3% and 1.5%, and then not at all.
15.2 Winsorising, trimming and rank methods
Definition 15.4 (Winsorisation, trimmed mean)
Winsorisation at level replaces the observations below the -quantile by that quantile and those above the -quantile by that one. The -trimmed mean discards the smallest and largest observations and averages the rest.
Both have breakdown point . Winsorising is the common repair for a cross-section of signals or returns before a regression (chapter 16): it keeps every observation’s sign and rank and caps its leverage. It is also a decision about the data that the report must state, because a winsorised Sharpe ratio is not the strategy’s Sharpe ratio. For a bad print, the better repair is upstream: filtering clearly erroneous trades against a robust band around a robust centre before any statistic is computed.
Definition 15.5 (Rank correlation, Spearman’s rho, Kendall’s tau)
A rank correlation measures dependence through the ranks of the observations only, so that it is unchanged by increasing transformations of either variable. Spearman’s rank correlation is the correlation of the ranks. Kendall’s tau is the probability that two independent pairs are concordant minus the probability that they are discordant, .
Definition 15.6 (Copula, tail dependence coefficient)
A copula is a joint distribution function on with uniform margins. The lower tail dependence coefficient of with continuous margins is , and the upper one is defined symmetrically.
Theorem 15.7 (Sklar)
Every joint distribution function with margins can be written for a copula , unique when the margins are continuous; conversely, any copula and any margins define a joint law this way.
For continuous margins the proof is one line: is the law of , whose margins are uniform. The theorem separates what a joint law says about each variable from what it says about their dependence, and rank correlations and tail dependence are properties of the copula alone. For the Gaussian copula with correlation , Spearman’s rho is , Kendall’s tau , and both tail dependence coefficients are zero: joint crashes become asymptotically independent, however high is. Credit portfolio models inherit that property from the Gaussian copula, a point One Quant Book 6 takes up with the Student- copula, whose tail dependence is positive.
The dollar and the pound, both priced in euros, have a daily Pearson correlation of 0.42 over 1999–2026. Spearman’s rho is 0.406, exactly what a Gaussian copula with that correlation implies (0.406), and Kendall’s tau 0.286 against an implied 0.277. The tails disagree: on the 5% of days when the dollar fell most against the euro, the pound was in its own worst 5% on 32% of them, against 20% for the Gaussian copula; at 1%, 21% against 10% (Figure 15.2). The average dependence is Gaussian; the joint extremes are not.
15.3 Heavy tails
Definition 15.8 (Heavy-tailed distribution, tail index)
A distribution is heavy-tailed (in the Pareto sense) if for a slowly varying ( as for every ). The exponent is its tail index.
Proposition 15.9 (Moments beyond the tail index)
If has tail index , then for and for .
Proof. . For large , a slowly varying satisfies for any (Potter’s bounds), so the integrand lies between : integrable at infinity when (take small) and not when . ∎
A Student with degrees of freedom has tail index . Mandelbrot’s 1963 study of cotton prices proposed stable laws, whose variance is infinite; the Hill plot below gives EUR/USD’s daily losses a tail index near four, so a finite variance and an unreliable fourth moment, in line with Cont’s survey of stylised facts, which finds tail indices above two and below five for most return series studied. The practical consequence is that the kurtosis, a fourth moment, is estimated badly or not at all: EUR/USD’s sample kurtosis over 27 years is 6.3, and a few days decide it. A quantile-quantile plot shows the tails directly (Figure 15.3).
15.4 Extreme-value theory
Definition 15.10 (Extreme-value theory, GEV distribution)
Extreme-value theory studies the laws of maxima and of exceedances of high thresholds. The generalised extreme value distribution with shape has distribution function for (and for ), up to location and scale.
Theorem 15.11 (Fisher–Tippett–Gnedenko)
If, for iid and constants , , the law of converges to a non-degenerate limit, that limit is a generalised extreme value distribution. Heavy tails with index give ; normal, exponential and similar tails give ; bounded tails give .
Definition 15.12 (Generalised Pareto distribution, peaks over threshold)
The generalised Pareto distribution with shape and scale has for (and for ). The peaks-over-threshold method fits it by maximum likelihood to the excesses of the observations above a high threshold .
Theorem 15.13 (Pickands–Balkema–de Haan)
For a distribution in the domain of attraction of , the law of the excess given is approximated, uniformly in the excess, by a generalised Pareto law with shape and some scale , the error vanishing as tends to the upper end of the support.
The theorem licenses an extrapolation: with of observations above , the loss exceeded with probability is
For EUR/USD daily losses above their 95% quantile (, 355 excesses), maximum likelihood gives and , and the one-in-a-thousand-days loss is 2.47%. The fitted depends on the threshold (0.03 above the 90% quantile, 0.21 above the 99%), the quantile much less (2.43% to 2.49% over the same thresholds): the extrapolation is steadier than the parameter.
15.5 Estimating the tail index
Definition 15.14 (Hill estimator)
With the order statistics of a positive sample, the Hill estimator of the tail index from the largest is (Hill, 1975).
It is the maximum likelihood estimator of for an exact Pareto tail above , and when with at a suitable rate, . The choice of is a bias-variance trade: small is noisy, large reaches into the body of the law, where the Pareto form fails. The Hill plot shows the trade (Figure 15.4): for EUR/USD losses, is 4.9 at (), near 3.9 from to 200, and drifts down to 2.6 at , where the body is being fitted.
Three models then answer the risk manager’s question, the one-in-a-thousand-days loss (Figure 15.5). The normal law with the sample volatility of 0.58% puts it at 1.80%, and 35 days of the 7 098 lost more, against 7 expected. A Student fitted by maximum likelihood (, scale 0.44%) puts it at 2.74%. The generalised Pareto fit puts it at 2.47%, exceeded on 6 days; the empirical quantile is 2.27%.
Now put the bad print of the hook into the full history. The normal quantile rises to 2.27% (+26%), because one pair of returns inflates the standard deviation by a quarter. The generalised Pareto quantile rises to 2.85% (+15%): the bad loss is the largest excess by far and lifts the fitted shape from 0.07 to 0.22. The Student moves least, to 2.91% (+6%), its likelihood bounding the influence of any single point. No tail estimate is robust in the sense of the first section; the tail is where the information is scarce, and one observation is a large share of it. The defence is upstream, in the filter.
15.6 Tutorial: the bad tick
Goal. Estimate EUR/USD’s volatility and tail robustly, then corrupt one print and measure what moves. End state: Figures 15.1, 15.4 and 15.5 and the three one-in-a-thousand-days losses with and without the bad print.
The Hill estimator from the largest losses.
def hill(x, k: int) -> tuple[float, float]: """Hill estimator of the tail index alpha from the k largest positive values: alpha_hat and its asymptotic standard error alpha_hat / sqrt(k).""" x = np.sort(np.asarray(x, dtype=float))[::-1] if k < 2 or x[k] <= 0: raise ValueError("need k >= 2 and the (k+1)-th largest value positive") gamma = float(np.mean(np.log(x[:k])) - math.log(x[k])) a = 1.0 / gamma return a, a / math.sqrt(k)Listing 15.1. The Hill estimator with its asymptotic standard error. code/firm/robust/firm_robust.py The generalised Pareto likelihood, fit and quantile.
def gpd_nll(xi: float, beta: float, y: np.ndarray) -> float: if beta <= 0: return math.inf if abs(xi) < 1e-9: return float(y.size * math.log(beta) + np.sum(y) / beta) z = 1 + xi * y / beta if np.any(z <= 0): return math.inf return float(y.size * math.log(beta) + (1 + 1 / xi) * np.sum(np.log(z))) def gpd_fit(excess) -> tuple[float, float]: """Maximum likelihood (xi, beta) of the generalised Pareto law for threshold excesses y > 0.""" y = np.asarray(excess, dtype=float) m, v = float(y.mean()), float(y.var()) xi0 = 0.5 * (1 - m * m / v) # method of moments start beta0 = 0.5 * m * (m * m / v + 1) th = _nelder_mead(lambda t: gpd_nll(t[0], math.exp(t[1]), y), [xi0, math.log(max(beta0, 1e-12))]) return float(th[0]), float(math.exp(th[1])) def gpd_quantile(u: float, xi: float, beta: float, p_u: float, p: float) -> float: """VaR at level p: u + beta / xi ((p_u / (1 - p))^xi - 1), P(X > u) = p_u.""" r = p_u / (1 - p) if abs(xi) < 1e-9: return u + beta * math.log(r) return u + beta / xi * (r**xi - 1)Listing 15.2. Peaks over threshold: likelihood, fit and quantile. code/firm/robust/firm_robust.py - Run
scale_table(False),scale_table(True),tail_quantiles(False),tail_quantiles(True)anddependence()inqm_robust.py, thenfig_robust.py.
What to change next. Build the filter: flag a print whose return exceeds ten MADs of the trailing 60 days and is reversed the next day, and check that it catches the bad print without flagging 2008; estimate the upper tail (gains) and compare.
15.7 Build: the robust statistics module
Purpose. Robust scales for every volatility the miniature firm estimates from prints, and the tail estimates its risk limits rest on.
Interface. mad, qn_scale, huber_location; winsorize, trimmed_mean; spearman, kendall_tau, tail_dependence(x, y, q); hill(x, k); gpd_fit(excess), gpd_quantile(u, xi, beta, p_u, p).
Rules. Scales are reported with the estimator named; a tail quantile is reported with its threshold, its number of excesses and its sensitivity to the threshold; no memory for pairwise statistics.
Acceptance tests. code/firm/robust/tests/: MAD and consistent at the normal and stable under 10% gross errors; equal to its brute-force definition; the rank-correlation formulas of the Gaussian copula; Hill on an exact Pareto; the GPD fit recovers its parameters and its quantile is exceeded at the stated rate.
Stretch. Block maxima with the GEV; a declustered peaks-over-threshold for dependent losses; the Student- copula fitted to the dollar and the pound.
Sources and further reading
- R. Cont, “Empirical properties of asset returns: stylized facts and statistical issues”, Quantitative Finance 1, 2001.
- P. J. Huber, “Robust estimation of a location parameter”, Annals of Mathematical Statistics 35, 1964.
- F. R. Hampel, “The influence curve and its role in robust estimation”, Journal of the American Statistical Association 69, 1974.
- P. J. Rousseeuw and C. Croux, “Alternatives to the median absolute deviation”, Journal of the American Statistical Association 88, 1993.
- A. Sklar, “Fonctions de répartition à dimensions et leurs marges”, Publications de l’Institut de Statistique de l’Université de Paris 8, 1959.
- R. A. Fisher and L. H. C. Tippett, “Limiting forms of the frequency distribution of the largest or smallest member of a sample”, Proceedings of the Cambridge Philosophical Society 24, 1928; B. Gnedenko, Annals of Mathematics 44, 1943.
- A. A. Balkema and L. de Haan, “Residual life time at great age”, Annals of Probability 2, 1974; J. Pickands, “Statistical inference using extreme order statistics”, Annals of Statistics 3, 1975.
- B. M. Hill, “A simple general approach to inference about the tail of a distribution”, Annals of Statistics 3, 1975.
- B. Mandelbrot, “The variation of certain speculative prices”, Journal of Business 36, 1963.
- European Central Bank, euro foreign exchange reference rates against the dollar and the pound (ECB Data Portal, series
EXR.D.USD.EUR.SP00.AandEXR.D.GBP.EUR.SP00.A), accessed 24 September 2026.
15.8 Exercises
Exercise 15.1 ★
What are the breakdown points of the 10%-trimmed mean and of the interquartile range?
Solution
Solution of Exercise 15.1.
10% for the trimmed mean (more than the trimmed share on one side carries it away); 25% for the interquartile range, which depends only on the quartiles.
Exercise 15.2 ★
A Pareto tail has index 3. Which of the mean, the variance and the kurtosis exist?
Solution
Solution of Exercise 15.2.
Moments of order below 3 exist: the mean and the variance. The kurtosis, a fourth moment, does not; the third moment is the borderline case and is infinite for an exact Pareto tail, since diverges logarithmically.
Exercise 15.3 ★
For a Gaussian copula with correlation 0.42, compute Spearman’s rho and Kendall’s tau.
Solution
Solution of Exercise 15.3.
; .
Exercise 15.4 ★★
Show that the Huber estimating function with has 95% efficiency at the normal law: evaluate .
Solution
Solution of Exercise 15.4.
. . The asymptotic variance is , so the efficiency relative to the mean is .
Exercise 15.5 ★★
Show that the Hill estimator is the maximum likelihood estimator of for the largest observations of an exact Pareto law above .
Solution
Solution of Exercise 15.5.
Above the Pareto density is . The log-likelihood of the largest is ; its derivative vanishes at , the Hill estimator.
Exercise 15.6 ★★
With , 355 excesses out of 7 098 days, and , compute the loss exceeded once in ten thousand days.
Solution
Solution of Exercise 15.6.
(with the unrounded fit; ). Two days in 27 years lost more (3.7% and 4.7%), about what the fit implies (0.7 expected); so few observations are why the extrapolation’s uncertainty is large.
Exercise 15.7 ★★★
Coding. Refit the generalised Pareto law to EUR/USD losses above the 90%, 95%, 97.5% and 99% quantiles, and tabulate and the one-in-a-thousand-days loss.
Solution
Solution of Exercise 15.7.
Above the 90%, 95%, 97.5% and 99% quantiles (710, 355, 178 and 71 excesses): , , , ; one-in-a-thousand-days loss 2.45%, 2.47%, 2.49%, 2.43%. The shape wanders, the quantile inside the data does not.
Exercise 15.8 ★★★
Find the flaw. “Our risk model estimates each currency’s volatility with the standard deviation of the last 250 days, but we winsorise the returns at the 1% and 99% quantiles first, so bad prints cannot hurt us.”
Solution
Solution of Exercise 15.8.
With 250 days, winsorising at 1% caps the two or three largest moves on each side. It happens to neutralise the chapter’s single bad print (the corrupted standard deviation becomes 0.341% against a clean 0.339%), but it also caps genuine large moves every day, biasing the clean volatility down (0.328% against 0.339%), most in the stress periods that matter; and three bad prints in a year pass through. The standard deviation of winsorised data also needs a consistency correction. Filter prints upstream, and use a robust scale by design.
15.9 Problem: The Bad Tick
Problem 15.1
Weekend problem — one erroneous print and the one-in-a-thousand-days loss
The data are the ECB’s daily euro reference rates against the dollar, 4 January 1999 to 23 September 2026 (7 098 daily returns). In the corrupted version, the print of 24 March 2026 reads 1.5172 instead of 1.1572.
Part I — Scale.
- What log-price error does the bad print make, and which returns does it affect?
- What are the standard deviation, MAD and of the last 250 returns, clean and corrupted?
- What are the mean, the median and the Huber location of those returns, clean and corrupted?
- How many days in the full history lost more than four standard deviations, against a normal law’s expectation?
- What is the sample kurtosis, and why is it a poor summary?
Part II — The tail.
- What does the Hill plot give for the tail index, and where?
- What are the threshold, the number of excesses and the fitted and of the peaks-over-threshold fit?
- What degrees of freedom and scale does the Student- fit give?
- What are the three one-in-a-thousand-days losses, and the empirical one?
- How many days exceeded the normal and the GPD quantiles, against the 7 expected?
Part III — The bad print in the tail.
- What do the three quantiles become with the bad print?
- Why does the normal quantile move most?
- What happens to the fitted ?
- Why does the Student move least?
- Would a robust scale in the normal model have saved it?
Part IV — Judgement.
- Which quantile would you put in the risk system, and with what caveat?
- Where should the bad print have been stopped?
- How does the dollar’s tail dependence with the pound change a two-currency limit?
- State the named result: the three quantiles and the error one bad print introduces into each.
- In one sentence: what can and cannot be made robust?
Solution
Solution of Problem 15.1.
1. : the returns of 24 and 25 March 2026 become about and . 2. Clean: 0.339%, 0.275%, 0.300%. Corrupted: 2.434%, 0.279%, 0.305%. 3. The mean () and the median () are unchanged, the first by luck, since the two bad returns cancel; the Huber location barely moves ( to ), its bounded influence at work. 4. Seventeen moves beyond four standard deviations (six losses, eleven gains), against 0.45 expected under a normal law. 5. 6.3; a fourth moment of a law whose tail index is near four is barely finite, and a handful of days decide the estimate. 6. About 3.9 on a plateau for between 100 and 200 largest losses (4.9 at , drifting to 2.6 at ). 7. , 355 excesses, , . 8. , scale 0.44%. 9. Normal 1.80%, Student 2.74%, GPD 2.47%; empirical 2.27%. 10. 35 days beyond the normal quantile, 6 beyond the GPD quantile. 11. Normal 2.27% (+27%), Student 2.91% (+6%), GPD 2.85% (+15%). 12. The normal quantile is the standard deviation times 3.09, and the standard deviation has an unbounded influence function; the pair of returns raises it from 0.58% to 0.74%. 13. It rises from 0.07 to 0.22: the bad loss is by far the largest excess, and the shape parameter is fitted mostly by the largest excesses. 14. Its likelihood weights an observation by , which falls with the size of the residual: the likelihood is itself a robust M-estimator. 15. Partly: a MAD-based normal quantile ignores the bad print but is even further from the true tail (the normal shape is wrong); robust scale does not make a thin-tailed model fit a heavy-tailed variable. 16. The GPD or Student- quantile, about 2.5% to 2.7%, reported with its threshold and sensitivity, never the normal 1.8%. 17. At ingestion: a filter that flags a print whose return exceeds many robust scales and reverses the next day, before any statistic is computed. 18. Joint losses are more frequent than a Gaussian copula implies (32% against 20% in the joint 5% tail), so a limit on the sum of the two exposures computed with correlation 0.42 understates their joint stress loss. 19. Named result: the bad tick: the one-in-a-thousand-days EUR/USD loss is 1.80% (normal), 2.74% (Student ) and 2.47% (GPD above the 95% quantile); one transposed-digit print raises them by 27%, 6% and 15%. 20. Estimates of the centre and the scale of the bulk can be made robust; estimates of the tail cannot, because in the tail every observation is information, so the protection must come from cleaning the data.
15.10 Interview questions
Interview question 15.1 ★ researcher, developer
Your volatility estimate for a quiet currency jumps sevenfold overnight. What do you check first?
Solution
Solution of Interview question 15.1.
The data: a bad print or a stale-then-corrected price produces a pair of large opposite returns. Look at the largest returns in the window, compare with another source, and compare the standard deviation with the MAD: a sevenfold gap between them is a data problem, not a market.
What the interviewer is looking for: Data before model; the paired-reversal signature; a robust scale as a diagnostic.
Interview question 15.2 ★★ researcher, risk
What is a breakdown point? Give the breakdown points of the mean, the median and the MAD.
Solution
Solution of Interview question 15.2.
The smallest fraction of arbitrary contamination that can carry the estimate arbitrarily far. Mean: , tending to zero. Median: one half. MAD: one half.
What the interviewer is looking for: The definition, and the contrast between zero and one half.
Interview question 15.3 ★★ researcher, risk
What does a tail index of 3 imply for the moments of returns, and for estimating kurtosis?
Solution
Solution of Interview question 15.3.
Moments of order below 3 exist and above 3 do not: finite variance, infinite kurtosis. The sample kurtosis then does not converge; it grows with the sample and is dominated by the largest observation, so it should not be used as a parameter.
What the interviewer is looking for: The moment rule, and the consequence for sample moments.
Interview question 15.4 ★★ risk, researcher
How would you estimate a 99.9% daily loss quantile from 25 years of data?
Solution
Solution of Interview question 15.4.
About six observations lie beyond the 99.9% quantile in 25 years: the empirical quantile is noisy, a normal fit is badly biased. Fit a generalised Pareto law above a high threshold (peaks over threshold), check stability across thresholds, compare with a Student- fit, handle volatility clustering by fitting to returns scaled by a volatility forecast, and report the uncertainty.
What the interviewer is looking for: Extreme-value theory with threshold diagnostics, and awareness of dependence.
Interview question 15.5 ★★ researcher, mle
Why use rank correlation instead of Pearson correlation, and what does neither tell you?
Solution
Solution of Interview question 15.5.
Rank correlations are invariant to increasing transformations and bounded in influence, so outliers and heavy tails do not dominate them. Neither Pearson nor rank correlation describes the tails: two pairs with the same rank correlation can have very different joint-crash probabilities.
What the interviewer is looking for: Invariance and robustness of ranks, and tail dependence as a separate quantity.
Interview question 15.6 ★★★ researcher, risk
What does Sklar’s theorem say, and why does the Gaussian copula understate joint crashes?
Solution
Solution of Interview question 15.6.
Any joint law is its margins composed with a copula, unique for continuous margins, so dependence can be modelled separately from the margins. The Gaussian copula has zero tail dependence for any correlation below one: extreme joint moves become asymptotically independent, while in markets they cluster, as a Student- copula allows.
What the interviewer is looking for: The decomposition, and tail dependence zero versus positive.
Terms defined in this chapter
- Copula, tail dependence coefficient
- Extreme-value theory, GEV distribution
- Generalised Pareto distribution, peaks over threshold
- Heavy-tailed distribution, tail index
- Hill estimator
- Huber loss, median absolute deviation
- Influence function, breakdown point
- Rank correlation, Spearman’s rho, Kendall’s tau
- Winsorisation, trimmed mean