Quantitative Methods · Methods
12Testing and Multiple Testing
A research team tries two hundred variants of a momentum signal and presents the best: its backtest -statistic is 3.1, a one-sided -value of one in a thousand. Had none of the variants any edge, and had they been independent, the largest of two hundred -statistics would still exceed 3.1 with probability 17.6%, about one time in six. The variants are not independent, which helps; the team also chose the sample period, the universe and the cost model, which does not. This chapter is about what a test can say: size and power and the lemma that makes the likelihood ratio the right statistic; the many quiet choices that multiply the tests a researcher runs; the corrections that control the chance of any false discovery, or their proportion; and the tests built for the search over strategies, which respect the correlation between them.
12.1 Tests, size and power
Definition 12.1 (Hypothesis test, null hypothesis, -value)
A hypothesis test is a rule that, from data , decides whether to reject a null hypothesis (a statement about the law of , such as “the strategy’s mean return is zero”) in favour of an alternative . It rejects when a statistic exceeds a critical value. The -value of an observed is : the probability, under the null, of a statistic at least as extreme as the one seen.
A -value is not the probability that the null is true, and one minus it is not the probability that the strategy works: both of those need a prior (chapter 14). It is a statement about the data under one hypothesis.
Definition 12.2 (Size, power)
The size of a test is the largest probability of rejecting when the null is true (a false positive, or type I error). The power of a test at an alternative is the probability of rejecting when that alternative is true; one minus the power is the type II error.
Theorem 12.3 (Neyman–Pearson lemma)
For a simple null (density ) against a simple alternative (), the test that rejects when , with chosen so that its size is , is most powerful: no test of size at most has higher power.
Proof. Let be the likelihood-ratio test and any test of size at most (both indicator functions of the rejection region). Pointwise : where , ; where , . Integrating, . ∎
For the mean of normal returns, the ratio against any positive mean is increasing in the sample mean, so the one-sided -test is the most powerful test of “no edge” against every positive edge at once. With composite hypotheses and nuisance parameters, the lemma becomes a recipe.
Definition 12.4 (Likelihood-ratio test)
The likelihood-ratio test of a null that imposes restrictions on rejects for large , where is the log-likelihood, its maximiser and its maximiser under the restrictions.
By Wilks’ theorem, under the null and regularity conditions (the true parameter inside the parameter space), is asymptotically . The condition matters in practice: testing that a Hawkes excitation (chapter 7) is zero puts the null on the boundary , and the limit becomes an equal mixture of a point mass at zero and a , which halves the -value.
Definition 12.5 (Kolmogorov–Smirnov test)
The Kolmogorov–Smirnov test of the null that are iid with a fully specified continuous cdf uses , the empirical cdf. Under the null, , whatever ; the 5% critical value is .
A thousand daily returns drawn from a Student with four degrees of freedom, scaled to 1% volatility, give against the normal law with the same volatility: above the critical value , a -value of 0.0010. Two cautions come with the test. It is weakest in the tails, where is small even when the tails are wrong; and when the mean and variance are estimated from the same data, is smaller than the tabulated law assumes, so the critical values must be simulated for the fitted family (the Lilliefors correction).
The sharpest use of size and power in a research group is the question of how long to wait.
Proposition 12.6 (Sample size for a Sharpe ratio)
With years of independent returns and an annual Sharpe ratio , the one-sided test of at size has power approximately , and reaches power after
Proof. By the delta method of chapter 11, the annualised estimate is approximately when the daily ratio is small. The test rejects when , which has probability ; setting it to gives . ∎
At 5% size and 80% power, a Sharpe ratio of 2 needs 1.5 years, a Sharpe ratio of 1 needs 6.2 years and a Sharpe ratio of 0.5 needs 24.7 years (Figure 12.1). Five years of a genuine Sharpe ratio of 1 are detected 72% of the time: more than one in four good strategies fails its first honest test.
12.2 The garden of forking paths
Definition 12.7 (Multiple testing, data snooping, garden of forking paths)
Multiple testing is running many tests and reporting some of them. Data snooping is choosing a model, a strategy or a parameter by searching over the same data that then test it, so that the reported statistic is the best of a search presented as a single test. The garden of forking paths is the implicit version: every choice the analyst would have made differently had the data been different (the sample period, the universe, the filter, the cost model) multiplies the tests actually run, even when only one analysis is ever performed.
The name is Gelman and Loken’s (2013), and the point is theirs: a researcher who never runs two backtests can still snoop, because the one backtest was shaped by looking at the data. A search is only corrected for if it is counted.
Proposition 12.8 (The maximum of independent statistics)
If are independent standard normal, , and as .
Proof. . For the growth, , and the maximum sits where , that is to leading order. ∎
The asymptote is slow: for the expected maximum is 2.75, not . That maximum of two hundred noise statistics clears the conventional bar almost always, and the hook’s 3.1 about one time in six (Figure 12.2). A single test at has ; the best of two hundred independent tests at has .
Real searches are correlated. The tutorial’s family is two hundred trend rules that hold the sign of the trailing -day return, for lookbacks days, applied to ten years of a driftless random walk with 1% daily volatility: no rule has any edge. Neighbouring lookbacks share almost all their signals (the best variant’s daily P&L has correlation 0.95 with its neighbour’s), distant ones much less (0.28 with the 5-day rule, 0.20 with the 204-day rule). In the history shown, 33 of the 200 variants have a one-sided -value below 5%, the mean -statistic across the family is 1.04 rather than zero, since every variant sees the same trending episodes, and the best, at a 42-day lookback, has and an annualised Sharpe ratio of 0.98 (Figure 12.3). That history is the 336th we simulated: we drew histories until the best variant came near the hook’s 3.1, which is the chapter’s subject enacted on the chapter itself.
12.3 Family-wise error and false discoveries
Definition 12.9 (Family-wise error rate, Bonferroni correction, Holm procedure)
For hypotheses of which are true, the family-wise error rate (FWER) of a procedure is the probability that it rejects at least one true null. The Bonferroni correction rejects when . The Holm procedure orders the -values and rejects , where is the last index with for every .
Proposition 12.10 (Bonferroni and Holm control the family-wise error rate)
Under any dependence between the tests, both procedures have FWER at most ; Holm rejects everything Bonferroni rejects.
Proof. Bonferroni: the union bound over the true nulls gives . Holm: let be the rank of the smallest -value among the true nulls. At most hypotheses precede it, so , and Holm rejects a true null only if , which by the union bound over the true nulls has probability at most . Holm’s thresholds are never below . ∎
Controlling the chance of any false discovery is the right goal when one false discovery is expensive, as when a single strategy goes live. When a research group screens hundreds of signals for a combined model, a few false ones among many true ones cost little, and demanding none throws away most of the true ones.
Definition 12.11 (False discovery rate, Benjamini–Hochberg procedure)
If a procedure makes rejections of which are true nulls, its false discovery rate is . The Benjamini–Hochberg procedure at level rejects for the largest with .
Theorem 12.12 (Benjamini–Hochberg)
If the -values of the true nulls are independent of each other and of the others, the Benjamini–Hochberg procedure has false discovery rate exactly . It remains at most under positive regression dependence, and at most under any dependence if is replaced by (Benjamini and Yekutieli, 2001).
Example 12.13 (Five -values)
Take and . Bonferroni’s threshold rejects two. Holm compares them in order with : and , then stops it at two. Benjamini–Hochberg compares with : the largest with is , so it rejects four. As adjusted -values (the smallest level at which each is rejected): Bonferroni , Holm , Benjamini–Hochberg .
The trade shows at scale. Among a thousand candidate signals of which a hundred are real (with -statistics centred on 3), the naive 5% test makes 136 discoveries of which 45 are false; Bonferroni and Holm make about 18, with a false one in 5% of screens; Benjamini–Hochberg at 5% makes 63, of which 2.9 are false on average, a false discovery rate of 4.4% against the theorem’s (Figure 12.4). It pays for its power with a family-wise error rate of 92%: almost every screen contains a false signal, and that is the design.
On the momentum family neither control is kind. Bonferroni and Holm both give the best variant an adjusted -value of ; Benjamini–Hochberg gives 0.076, smaller because its step-up borrows strength from the 33 other small -values. No variant survives either at 5%. But both corrections were built for the worst case of dependence, and two hundred variants that share most of their signals are far from two hundred independent chances.
12.4 Reality-check tests for strategy search
Definition 12.14 (Reality Check, superior predictive ability test)
For strategies with performance series in excess of a benchmark, the Reality Check (White, 2000) tests with the statistic , whose null distribution is estimated by the bootstrap of . The superior predictive ability test (Hansen, 2005) studentises each and recentres only the strategies that are not clearly worse than the benchmark, so that adding poor strategies to the search does not make the test more conservative.
Both tests replace the union bound by the actual law of the maximum, and the bootstrap (chapter 13) keeps the dependence between strategies and over time. When the statistics are approximately jointly normal, a Gaussian simulation with their estimated correlation does the same job at no cost.
Method 12.15 (Max-statistic -values)
Given -statistics and the correlation matrix of the underlying P&L series: draw for ; the adjusted -value of is the share of draws with . The step-down version (Romano and Wolf, 2005) takes the maximum, for the -th largest statistic, only over the hypotheses not yet rejected. Both control the family-wise error rate; the effective number of independent trials of the search is the that solves .
For the momentum family, 200 000 draws from the estimated correlation give the best variant an adjusted -value of 0.030, against 0.187 for Bonferroni and 0.171 for two hundred independent tests. The search over two hundred correlated lookbacks amounts to 32 independent trials. The max-statistic 5% critical value is 2.92, well below Bonferroni’s 3.48, and three variants clear it. The check that matters: over 2 000 fresh ten-year histories, the best variant reaches in 2.6% of them, close to the simulated 3.0%.
So the corrected test rejects, at 5%, a strategy that has no edge. It is allowed to, 3% of the time; and this history was the 336th we drew. The correction accounts for the two hundred lookbacks it was told about and for nothing else: not the histories we discarded, not the other signal families the team tried last year. Harvey, Liu and Zhu (2016) reach a similar conclusion for published equity factors, arguing that a new factor should clear a -statistic of about 3 rather than 2.
Definition 12.16 (Deflated Sharpe ratio)
For a Sharpe ratio (per period) estimated from returns with skewness and kurtosis , the probability that the true ratio exceeds is approximately
The deflated Sharpe ratio (Bailey and López de Prado, 2014) is this probability with the expected maximum of null Sharpe ratios of variance : , the Euler–Mascheroni constant.
For the best momentum variant the probability that its true Sharpe ratio is positive is 0.999, the single-test view. Deflated for two hundred trials (with , the null variance), the benchmark becomes an annual 0.87 and the deflated Sharpe ratio 0.63; for the 32 effective trials, 0.67 and 0.84. Neither reaches the 0.95 a desk would want, and the number of trials, which the researcher controls and the reader rarely sees, moves the answer more than the data do.
12.5 Tutorial: the best of two hundred
Goal. Search two hundred worthless trend rules, pick the best, and measure how surprising it is under five corrections. End state: Figures 12.3 and 12.2 and the adjusted -values 0.187, 0.076 and 0.030.
Adjusted -values by Holm’s step-down and Benjamini–Hochberg’s step-up.
def holm(p) -> np.ndarray: """Step-down: the k-th smallest p-value (k = 1..m) is multiplied by m - k + 1, then made monotone.""" p = _check(p) m = p.size order = np.argsort(p) adj = np.maximum.accumulate((m - np.arange(m)) * p[order]) out = np.empty(m) out[order] = np.minimum(1.0, adj) return out def benjamini_hochberg(p, dependence_factor: float = 1.0) -> np.ndarray: """Step-up: the k-th smallest p-value is multiplied by m / k, then made monotone from the top.""" p = _check(p) m = p.size order = np.argsort(p) ranked = p[order] * m * dependence_factor / np.arange(1, m + 1) adj = np.minimum.accumulate(ranked[::-1])[::-1] out = np.empty(m) out[order] = np.minimum(1.0, adj) return outListing 12.1. Holm and Benjamini–Hochberg adjusted -values. code/firm/multitest/firm_multitest.py Max-statistic -values from a Gaussian null with the estimated correlation, or from any matrix of null draws.
def _gaussian_null(corr: np.ndarray, n_sim: int, seed: int) -> np.ndarray: w, v = np.linalg.eigh(0.5 * (corr + corr.T)) root = v * np.sqrt(np.clip(w, 0.0, None)) rng = np.random.default_rng(seed) out = np.empty((n_sim, corr.shape[0])) for a in range(0, n_sim, 10_000): # blocks keep memory flat b = min(n_sim, a + 10_000) out[a:b] = rng.standard_normal((b - a, corr.shape[0])) @ root.T return out def maxt_pvalues(t, corr=None, n_sim: int = 100_000, seed: int = 0, null_draws=None, stepdown: bool = False) -> np.ndarray: """One-sided max-statistic adjusted p-values: P(max_j Z_j >= t_i) under the joint null. The null is Gaussian with correlation `corr` (identity if None), or the rows of `null_draws` (n_sim x m, e.g. bootstrap statistics centred on the null). With stepdown=True the maximum for the k-th largest statistic runs over the hypotheses not yet rejected (Romano-Wolf), which is uniformly less conservative and still controls the family-wise error rate. """ t = np.asarray(t, dtype=float) m = t.size if null_draws is None: null_draws = _gaussian_null(np.eye(m) if corr is None else np.asarray(corr, float), n_sim, seed) null_draws = np.asarray(null_draws, dtype=float) if not stepdown: mx = np.sort(null_draws.max(axis=1)) return 1.0 - np.searchsorted(mx, t, side="left") / mx.size order = np.argsort(-t) adj = np.empty(m) for k, i in enumerate(order): mx = null_draws[:, order[k:]].max(axis=1) adj[k] = np.mean(mx >= t[i]) adj = np.maximum.accumulate(adj) out = np.empty(m) out[order] = adj return outListing 12.2. Single-step and step-down max-statistic -values. code/firm/multitest/firm_multitest.py - Run
problem()inqm_testing.py, thenfamily_exceedance(3.11)to check the simulated -value on fresh histories, andfig_testing.py.
What to change next. Add the mirror-image reversal rules and see the effective number of trials grow; replace the Gaussian null by the stationary bootstrap of chapter 13 through the null_draws hook.
12.6 Build: the multiple-testing module
Purpose. Every strategy the miniature firm promotes from research carries the size of the search behind it and an adjusted -value; this module computes them.
Interface. bonferroni(p), holm(p), benjamini_hochberg(p), benjamini_yekutieli(p) returning adjusted -values; maxt_pvalues(t, corr, n_sim, seed, null_draws, stepdown); effective_trials(p_single, p_family); expected_max_sr(n_trials, sr_var); deflated_sharpe(sr, n_obs, skew, kurt, sr0); norm_cdf, norm_ppf.
Rules. Adjusted -values, not reject flags, so that the level is chosen by the caller; one-sided statistics for strategy search; the null draws are a parameter, so that the bootstrap replaces the Gaussian without touching the callers.
Acceptance tests. code/firm/multitest/tests/: the five--value example; FWER and FDR control by simulation; the max-statistic -value equals the independent formula for and the naive -value for perfectly correlated tests; the expected-maximum approximation within 1.5%.
Stretch. Hansen’s SPA test with the stationary bootstrap; the Romano–Wolf step-down with bootstrap null draws; the probability of backtest overfitting by combinatorial cross-validation.
Sources and further reading
- J. Neyman and E. S. Pearson, “On the problem of the most efficient tests of statistical hypotheses”, Philosophical Transactions of the Royal Society A 231, 1933.
- S. Holm, “A simple sequentially rejective multiple test procedure”, Scandinavian Journal of Statistics 6, 1979.
- Y. Benjamini and Y. Hochberg, “Controlling the false discovery rate: a practical and powerful approach to multiple testing”, Journal of the Royal Statistical Society B 57, 1995.
- Y. Benjamini and D. Yekutieli, “The control of the false discovery rate in multiple testing under dependency”, Annals of Statistics 29, 2001.
- H. White, “A reality check for data snooping”, Econometrica 68, 2000.
- P. R. Hansen, “A test for superior predictive ability”, Journal of Business and Economic Statistics 23, 2005.
- J. P. Romano and M. Wolf, “Stepwise multiple testing as formalized data snooping”, Econometrica 73, 2005.
- C. R. Harvey, Y. Liu and H. Zhu, “…and the cross-section of expected returns”, Review of Financial Studies 29, 2016.
- D. H. Bailey and M. López de Prado, “The deflated Sharpe ratio”, Journal of Portfolio Management 40(5), 2014.
- A. Gelman and E. Loken, “The garden of forking paths”, working paper, 2013.
12.7 Exercises
Exercise 12.1 ★
A backtest has . What is its one-sided -value, and its Bonferroni-adjusted -value if it is the best of twenty variants?
Solution
Solution of Exercise 12.1.
; Bonferroni over twenty variants, : not significant at 5%.
Exercise 12.2 ★
A test rejects when . Under the alternative of interest . What are its size and its power?
Solution
Solution of Exercise 12.2.
Size ; power .
Exercise 12.3 ★
Apply Bonferroni and Holm at 5% to .
Solution
Solution of Exercise 12.3.
Bonferroni’s threshold is 0.0125: it rejects the first two. Holm compares the ordered -values with 0.0125, 0.0167, 0.025 and 0.05 in turn: 0.001, 0.012 and 0.02 pass and 0.3 does not, so it makes three rejections. The adjusted -values are, for Bonferroni, 0.004, 0.048, 0.08 and 1; for Holm, 0.004, 0.036, 0.04 and 0.3.
Exercise 12.4 ★★
How many years of independent returns does a test of 5% size need to detect an annual Sharpe ratio of 0.7 with 90% power?
Solution
Solution of Exercise 12.4.
years.
Exercise 12.5 ★★
Daily returns are normal with mean zero. The first 500 days have sample variance , the next 500 . Compute the likelihood-ratio statistic for a common variance and its -value.
Solution
Solution of Exercise 12.5.
With known zero mean, with the pooled (in ): . Against , : the variance changed.
Exercise 12.6 ★★
For ten independent noise strategies, what is the probability that the best has , and what is the expected maximum by the formula of the deflated Sharpe ratio?
Solution
Solution of Exercise 12.6.
. The formula gives (the exact expected maximum is 1.54).
Exercise 12.7 ★★★
Coding. Repeat the thousand-signal screen with null statistics that share a common factor (correlation 0.5). Does Benjamini–Hochberg still control the false discovery rate at 5%? How often does the realised proportion of false discoveries exceed 10%?
Solution
Solution of Exercise 12.7.
With 500 screens (fdr_correlated()): the false discovery rate is 3.0%, within the 5% bound (positive dependence), with about 63 discoveries per screen. But the realised proportion is volatile: it exceeds 10% in 7.8% of screens, when the common factor pushes all the null statistics up together. The procedure controls an average, not each screen.
Exercise 12.8 ★★★
Find the flaw. “We tested our signal separately on twelve markets. It is significant at 5% in three of them, where we now trade it; the three backtests are our estimate of its performance.”
Solution
Solution of Exercise 12.8.
Twelve tests at 5% produce 0.6 false positives on average; three or more occur with probability 2.0% if the signal is useless everywhere and the markets independent, so the count is mild evidence of something. The flaw is the selection: the three winners were chosen because their backtests were good, so their in-sample performance is biased upward (the winner’s curse), and the markets are not independent. Report the family result, and estimate performance on data not used to choose the markets.
12.8 Problem: Two Hundred Momentum Variants
Problem 12.1
Weekend problem — how surprising is the best of a correlated search?
A team backtests two hundred trend rules (hold the sign of the trailing -day return, ) on ten years of a market that, unknown to them, is a driftless random walk with 1% daily volatility. The history is the chapter’s seeded one.
Part I — The search.
- Which lookback wins, with what -statistic and annualised Sharpe ratio?
- What is its naive one-sided -value?
- How many of the two hundred pass a one-sided 5% test?
- What is the mean -statistic across the family, and why is it not near zero?
- What is the correlation of the best rule’s P&L with its neighbour’s, and with the 5-day and 204-day rules’?
Part II — Corrections.
- What are the Bonferroni-adjusted -value and the Bonferroni critical value?
- What does Holm give for the best, and why the same?
- What does Benjamini–Hochberg give, and why smaller?
- What would the adjusted -value be if the two hundred were independent?
- What are the max-statistic adjusted -value and critical value, from the estimated correlation?
Part III — What the search amounts to.
- What is the effective number of independent trials?
- How many variants clear the max-statistic threshold?
- Over 2 000 fresh histories, how often does the best rule reach ?
- What are the probabilistic Sharpe ratio and the deflated Sharpe ratio with 200 and with the effective number of trials?
- Do the answers change if daily returns have Student- tails with four degrees of freedom?
Part IV — Judgement.
- The rules have no edge by construction. Why does the corrected test reject?
- What would you require before trading the 42-day rule?
- What should the research report record about the search?
- State the named result: the adjusted -value of the best of two hundred correlated variants and the effective number of trials.
- In one sentence: what does a multiple-testing correction protect against, and what not?
Solution
Solution of Problem 12.1.
1. Lookback 42 days: , annualised Sharpe ratio 0.98. 2. . 3. 33 of 200. 4. 1.04: every rule trades the same history, so its trending episodes lift the whole family together. 5. 0.95 with the 43-day rule; 0.28 with the 5-day and 0.20 with the 204-day rule. 6. ; critical value . 7. 0.187: Holm multiplies the smallest -value by , as Bonferroni does; it gains only further down the list. 8. 0.076: the step-up takes the minimum of over the ranks above, and the 33 small -values make those small. 9. . 10. 0.030 from 200 000 Gaussian draws with the estimated correlation; the 5% critical value of the maximum is 2.92. 11. trials. 12. Three. 13. In 2.6% of them, close to the simulated 3.0% (the sampling error of 2 000 histories is 0.4 points). 14. The probability that the true Sharpe ratio is positive is 0.999. Deflated for 200 trials: benchmark 0.87 a year, deflated Sharpe ratio 0.63; for 32 trials: 0.67 and 0.84. 15. Barely: with Student- returns the best rule reaches 3.11 in 2.35% of 2 000 fresh histories; the -statistics are still close to normal after 2 520 days. 16. A 5% test rejects a true null 5% of the time, and the best rule’s adjusted -value is 3%; besides, the history is the 336th drawn, a search the correction was not told about. 17. A holdout never used in the search, an economic reason for trend at that horizon, performance net of costs, and an adjusted -value that counts every family the team tried. 18. The number of variants and families tried, their correlation, the adjusted -values and the method, the sample period and what was fixed before looking at the data. 19. Named result: the best of two hundred: the best of two hundred correlated momentum rules, naive , has a max-statistic adjusted -value of 0.030 (Bonferroni: 0.187); the search amounts to 32 independent trials. 20. It protects against the searches it is told about, and against nothing the researcher did not count.
12.9 Interview questions
Interview question 12.1 ★ researcher, trader
What is a -value, and what is it not?
Solution
Solution of Interview question 12.1.
The probability, if the null is true, of a statistic at least as extreme as the one observed. It is not the probability that the null is true, not the probability that the result will repeat, and not a measure of the size of the effect.
What the interviewer is looking for: The conditional direction (, not ), and that effect size needs its own number.
Interview question 12.2 ★★ researcher, trader
You backtested a hundred strategies and the best has . Is it real?
Solution
Solution of Interview question 12.2.
Probably not. If all were noise and independent, the best of 100 exceeds 2.8 with probability ; correlation lowers that, but then ask how many variants and choices came before the hundred. Adjust for the search (max-statistic or Bonferroni), then test on a holdout.
What the interviewer is looking for: The one-in-four arithmetic, the effect of correlation, and a holdout as the real test.
Interview question 12.3 ★★ researcher, mle
Bonferroni or Benjamini–Hochberg: when would you use which?
Solution
Solution of Interview question 12.3.
Bonferroni or Holm (family-wise error) when one false positive is costly, such as the one strategy that goes live; Benjamini–Hochberg (false discovery rate) when screening many candidates for a combined model, where a few false ones among many true ones are acceptable. Under strong dependence use the max-statistic version or Benjamini–Yekutieli.
What the interviewer is looking for: Matching the error rate to the cost of a false discovery; Holm dominates Bonferroni.
Interview question 12.4 ★★ researcher, trader, risk
How long must a strategy run before you can confirm a Sharpe ratio of 1?
Solution
Solution of Interview question 12.4.
The annual estimate has standard error about , so : four years for , and at 5% size with 80% power years. Autocorrelated returns and fat tails make it longer; a search makes it longer still.
What the interviewer is looking for: The rule and the power calculation, not only .
Interview question 12.5 ★★★ researcher
Your two hundred variants are highly correlated. How do you correct for the search?
Solution
Solution of Interview question 12.5.
Use the law of the maximum rather than the union bound: estimate the correlation of the variants’ P&L and simulate the maximum of correlated normals, or bootstrap the maximum jointly (Reality Check, SPA, Romano–Wolf step-down). Report the effective number of trials. Bonferroni is valid but very conservative here.
What the interviewer is looking for: Max-statistic over union bound, joint resampling that keeps the dependence, and counting families as well as variants.
Interview question 12.6 ★★ mle, researcher
What is the garden of forking paths, and where does it hide in a machine-learning pipeline?
Solution
Solution of Interview question 12.6.
The implicit multiplicity of choices that depend on the data even when one analysis is run. In machine learning: feature choices made after looking at validation scores, early stopping and hyperparameters tuned on the test set, reruns with new seeds until the curve looks good, and a test set reused across many model iterations until it becomes a training set.
What the interviewer is looking for: Concrete leak points, and the discipline of a holdout touched once.
Interview question 12.7 ★★ researcher, risk
You fit a distribution to returns and a Kolmogorov–Smirnov test does not reject it. What have you learned?
Solution
Solution of Interview question 12.7.
Little. The test is weak in the tails, which is where return distributions differ; with parameters fitted to the same data its tabulated critical values are too lenient (Lilliefors), so it rejects even less often; and failing to reject is not acceptance. Compare tail quantiles or exceedance counts directly, with simulated critical values.
What the interviewer is looking for: Weak tail power, the fitted-parameter correction, and absence of evidence versus evidence of fit.
Terms defined in this chapter
- Deflated Sharpe ratio
- False discovery rate, Benjamini–Hochberg procedure
- Family-wise error rate, Bonferroni correction, Holm procedure
- Hypothesis test, null hypothesis, p-value
- Kolmogorov–Smirnov test
- Likelihood-ratio test
- Multiple testing, data snooping, garden of forking paths
- Reality Check, superior predictive ability test
- Size, power