Quantitative Finance · Book 4 · Methods

Quantitative Methods

Quantitative Methods · Methods

13Resampling

A strategy that sells volatility has earned a Sharpe ratio of 1.1 over five years, and the head of desk asks for a 95% interval before doubling its capital. Resampling the five years day by day gives an interval from 0.17 to 2.25: comfortably positive. Resampling blocks of about a month, which keeps together the quiet weeks and the violent ones, gives −0.25-0.25 to 2.962.96: the data cannot exclude that the strategy has no edge. Only the second interval is honest, and even it is too narrow. This chapter is about estimating the uncertainty of a statistic by recomputing it on the data themselves: the bootstrap and why it works, the block versions that dependent data need, permutation tests and the exchangeability that makes them exact, the older jackknife, and the bootstrap version of chapter 12’s test for the best of many strategies.

13.1 The bootstrap

Definition 13.1 (Empirical distribution function, bootstrap)

The empirical distribution function of X1,…,XnX_1, \dots, X_n is Fn(x)=1n#{i:Xi≤x}F_n(x) = \frac1n\#\{i: X_i \le x\}, the law that puts mass 1/n1/n on each observation. The bootstrap estimates the sampling law of a statistic θ^=T(X1,…,Xn)\hat\theta = T(X_1, \dots, X_n) by the law of θ^∗=T(X1∗,…,Xn∗)\hat\theta^* = T(X_1^*, \dots, X_n^*), where the Xi∗X_i^* are drawn independently from FnF_n, that is, with replacement from the data; in practice by computing θ^∗\hat\theta^* on BB resamples.

The principle is a substitution: the unknown law FF is replaced by FnF_n, and the sampling experiment that cannot be repeated in the world is repeated on the computer. The bootstrap standard error of a mean is s(n−1)/n/ns\sqrt{(n - 1)/n}/\sqrt n, the usual one; its use is for statistics with no formula, or with a formula resting on assumptions one does not trust.

Definition 13.2 (Bootstrap percentile interval, coverage probability)

The bootstrap percentile interval at level 1−α1 - \alpha runs from the α/2\alpha/2 to the 1−α/21 - \alpha/2 quantile of the bootstrap draws θ^∗\hat\theta^*. The coverage probability of an interval procedure is the probability, over repeated samples from the true law, that its interval contains the true parameter.

A nominal 95% interval is a claim about coverage, and coverage can be measured by simulation: draw many histories from a known law, build the interval on each, count. That is the test every interval in this chapter is put to.

The bootstrap is consistent (the law of n(θ^∗−θ^)\sqrt n(\hat\theta^* - \hat\theta) given the data approaches that of n(θ^−θ)\sqrt n(\hat\theta - \theta)) for smooth functionals of FF with iid data: means, variances, correlations, Sharpe ratios, regression coefficients. It fails where the statistic depends on the extremes. For the maximum of a uniform sample, each resample misses the sample maximum with probability (1−1/n)n→e−1(1 - 1/n)^n \to e^{-1}, so the bootstrap law of n(max⁡∗−max⁡)n(\max^* - \max) has an atom of mass 1−e−1=0.6321 - e^{-1} = 0.632 at zero, while the true law of n(θ−max⁡)n(\theta - \max) is continuous. Tail quantiles, drawdowns and the largest loss share the problem: the bootstrap cannot resample a crash the history does not contain.

13.2 Dependent data: block and stationary bootstraps

Resampling days independently destroys whatever links one day to the next. For a strategy whose P&L is serially correlated, the iid bootstrap then estimates the variance of the mean by γ^(0)/n\hat\gamma(0)/n where the truth is the long-run variance over nn (chapter 11), and its intervals are too narrow.

Definition 13.3 (Moving-block bootstrap, stationary bootstrap)

The moving-block bootstrap (Künsch, 1989) builds a resample by concatenating blocks of bb consecutive observations, Xj+1,…,Xj+bX_{j+1}, \dots, X_{j+b}, with starting points jj drawn uniformly (indices taken modulo nn in the circular version). The stationary bootstrap (Politis and Romano, 1994) draws the block lengths independently from a geometric law of mean bb: each resampled day continues the current block with probability 1−1/b1 - 1/b and starts a new block at a uniform position with probability 1/b1/b.

Proposition 13.4 (The stationary bootstrap is stationary)

Conditionally on the data, the stationary-bootstrap series X1∗,X2∗,…X_1^*, X_2^*, \dots is strictly stationary.

Proof. Let ItI_t be the index of Xt∗X_t^*. I1I_1 is uniform on {1,…,n}\{1, \dots, n\}. Given It−1I_{t-1}, ItI_t is It−1+1I_{t-1} + 1 (modulo nn) with probability 1−1/b1 - 1/b and uniform with probability 1/b1/b; both maps send the uniform law to itself, so every ItI_t is uniform. The sequence (It)(I_t) is a Markov chain started in its stationary law, hence strictly stationary, and so is Xt∗=XItX_t^* = X_{I_t}. ∎

The fixed-length blocks of the moving-block bootstrap make the resample non-stationary: the joins between blocks fall at fixed dates. The random lengths remove that artefact at the price of slightly more variance, and both need a block length. Too short, and the dependence is cut; too long, and there are few distinct blocks to recombine. The mean squared error of the variance estimate is minimised by a block length growing like n1/3n^{1/3}, with a constant that depends on the autocovariances; Politis and White (2004), with the correction of Patton, Politis and White (2009), estimate it from the sample autocovariances with a flat-top lag window. That is the rule optimal_block_length implements.

Example 13.5 (The volatility seller)

Each day the strategy earns a premium and pays the day’s squared return, xt=μ+1−rt2x_t = \mu + 1 - r_t^2, with rt=σtεtr_t = \sigma_t\varepsilon_t and ln⁡σt\ln\sigma_t an AR(1) with coefficient 0.9 and stationary standard deviation 0.4, scaled so that E[σt2]=1\E[\sigma_t^2] = 1. The market returns rtr_t are uncorrelated, but the squared returns are not: Cov⁡(rt2,rt+k2)=e0.64⋅0.9k−1\Cov(r_t^2, r_{t+k}^2) = e^{0.64 \cdot 0.9^k} - 1, and the P&L’s long-run variance is 3.9 times its variance. The premium μ\mu is set for a population Sharpe ratio of 1.1. The seeded five-year history has a Sharpe ratio of 1.11, a skewness of −4.7-4.7, a kurtosis of 34, and autocorrelations of 0.13 at one day, 0.10 at five, −0.02-0.02 at twenty.

The automatic rule chooses a mean block of 18.6 days for the stationary bootstrap on this history (21.2 for the circular one), and the bootstrap distributions separate (Figure 13.1): the iid bootstrap gives a standard error of 0.53 and the interval (0.17,2.25)(0.17, 2.25); the moving-block bootstrap 0.84 and (−0.34,2.93)(-0.34, 2.93); the stationary bootstrap 0.83 and (−0.25,2.96)(-0.25, 2.96). The delta method of chapter 11 with the sample skewness and kurtosis agrees with the iid bootstrap (0.52): both assume independent days.

Bootstrap distributions of the volatility seller’s five-year Sharpe ratio (sample value 1.11, dashed): resampling days independently (shaded, standard error 0.53) and resampling random blocks of mean length 18.6 days (standard error 0.83). Data: the chapter’s tutorial, 4 000 resamples, seeded.
Figure 13.1. Bootstrap distributions of the volatility seller’s five-year Sharpe ratio (sample value 1.11, dashed): resampling days independently (shaded, standard error 0.53) and resampling random blocks of mean length 18.6 days (standard error 0.83). Data: the chapter’s tutorial, 4 000 resamples, seeded.

Which is right is a question for simulation. Over 4 000 histories of the same process, the five-year Sharpe ratio has a standard deviation of 1.10: larger than any of the bootstrap estimates. Averaged over sixty histories, the stationary-bootstrap standard error rises with the block length from 0.59 (one-day blocks, the iid bootstrap) to a plateau of about 0.97 between 40 and 90 days, and never reaches the truth (Figure 13.2). The shortfall is the tails: histories that happen to contain no violent week cannot resample one.

The stationary bootstrap’s standard error of the five-year Sharpe ratio against the mean block length, averaged over sixty simulated histories of the volatility seller, and the true standard deviation (1.10, from 4 000 histories). One-day blocks are the iid bootstrap (0.59); the plateau near 0.97 is the most any block length recovers. Data: the chapter’s tutorial, seeded.
Figure 13.2. The stationary bootstrap’s standard error of the five-year Sharpe ratio against the mean block length, averaged over sixty simulated histories of the volatility seller, and the true standard deviation (1.10, from 4 000 histories). One-day blocks are the iid bootstrap (0.59); the plateau near 0.97 is the most any block length recovers. Data: the chapter’s tutorial, seeded.

Coverage is the verdict (Figure 13.3). Over 400 simulated histories, nominal 95% iid intervals contain the true Sharpe ratio of 1.1 in 68.5% of them; moving-block and stationary intervals with the automatic block length in 82.75%; stationary intervals with 40-day blocks in 84.0%. Remove the volatility clustering (an AR coefficient of zero) and the iid interval covers 91.5%, the remaining shortfall being the heavy tails.

Coverage of nominal 95% bootstrap intervals for the volatility seller’s Sharpe ratio over 400 simulated five-year histories: iid 68.5%, block bootstraps with the automatic length 82.75%, stationary with 40-day blocks 84.0%; the dashed line is the nominal 95%. Data: the chapter’s tutorial, seeded.
Figure 13.3. Coverage of nominal 95% bootstrap intervals for the volatility seller’s Sharpe ratio over 400 simulated five-year histories: iid 68.5%, block bootstraps with the automatic length 82.75%, stationary with 40-day blocks 84.0%; the dashed line is the nominal 95%. Data: the chapter’s tutorial, seeded.

13.3 Permutation tests

Definition 13.6 (Exchangeability, permutation test)

Random variables Y1,…,YnY_1, \dots, Y_n are exchangeable (exchangeability) if the law of (Yπ(1),…,Yπ(n))(Y_{\pi(1)}, \dots, Y_{\pi(n)}) is the same for every permutation π\pi. A permutation test of the null that YY is independent of XX compares a statistic T(X,Y)T(X, Y) with its values over random permutations of YY; the pp-value is (1+#{T(X,Yπ)≥T(X,Y)})/(1+P)(1 + \#\{T(X, Y_\pi) \ge T(X, Y)\})/(1 + P) for PP permutations.

Proposition 13.7 (Exactness under exchangeability)

If, under the null, YY is exchangeable and independent of XX, the permutation pp-value satisfies P(p≤α)≤α\P(p \le \alpha) \le \alpha for every α\alpha and every sample size.

Proof. Under the null, the observed YY and the PP permuted copies are exchangeable given XX, so the rank of T(X,Y)T(X, Y) among the P+1P + 1 values is uniform (ties broken against rejection). The pp-value is that rank from the top divided by P+1P + 1, and P(rank≤k)≤k/(P+1)\P(\text{rank} \le k) \le k/(P + 1). ∎

The condition is the whole test. Daily P&L is not exchangeable when it is serially dependent, and a persistent signal is not either. Take the volatility seller’s P&L and a signal built from a different, independent market: its trailing twenty-day mean squared return. Nothing links them. Yet over 200 such independent pairs, the permutation test of their correlation rejects at 5% in 21.5% of them, because permuting destroys the persistence that makes spurious correlations between two slow series large. Rotating one series by a random offset instead (a circular shift, which keeps its dependence) rejects in 2.5%. The same trap catches tests of a signal against overlapping returns, and the cure is the same: resample in a way that preserves what the null does not break.

13.4 The jackknife

Definition 13.8 (Jackknife)

The jackknife recomputes a statistic with each observation left out in turn, θ^(i)\hat\theta_{(i)}, and with θˉ=1n∑iθ^(i)\bar\theta = \frac1n\sum_i\hat\theta_{(i)} estimates its bias by (n−1)(θˉ−θ^)(n - 1)(\bar\theta - \hat\theta) and its standard error by n−1n∑i(θ^(i)−θˉ)2\sqrt{\frac{n-1}n\sum_i(\hat\theta_{(i)} - \bar\theta)^2}.

Proposition 13.9 (Jackknife bias correction)

If E[θ^n]=θ+a/n+O(n−2)\E[\hat\theta_n] = \theta + a/n + O(n^{-2}), the bias-corrected estimate θ^−(n−1)(θˉ−θ^)\hat\theta - (n - 1)(\bar\theta - \hat\theta) has bias O(n−2)O(n^{-2}).

Proof. Write E[θ^n]=θ+a/n+c/n2+O(n−3)\E[\hat\theta_n] = \theta + a/n + c/n^2 + O(n^{-3}). Each leave-one-out estimate uses n−1n - 1 observations, so E[θˉ]=θ+a/(n−1)+c/(n−1)2+O(n−3)\E[\bar\theta] = \theta + a/(n - 1) + c/(n - 1)^2 + O(n^{-3}). The corrected estimate is nθ^−(n−1)θˉn\hat\theta - (n - 1)\bar\theta, with mean θ+a−a+c(1/n−1/(n−1))+O(n−2)=θ+O(n−2)\theta + a - a + c\bigl(1/n - 1/(n - 1)\bigr) + O(n^{-2}) = \theta + O(n^{-2}): the 1/n1/n terms cancel exactly. ∎

For the plug-in variance 1n∑(Xi−Xˉ)2\frac1n\sum(X_i - \bar X)^2, whose bias is exactly −σ2/n-\sigma^2/n, the correction returns the unbiased s2s^2 exactly. On the volatility seller’s history the jackknife standard error of the Sharpe ratio is 0.53, the iid bootstrap’s number, and for the same reason: leaving out single days, it treats them as independent. A block jackknife, which deletes blocks, is its dependent version. The jackknife is cheap (nn recomputations, no randomness) and good for bias; the bootstrap gives whole distributions.

13.5 Bootstrapping the reality check

Chapter 12’s max-statistic test simulated the null of two hundred momentum rules from a Gaussian law with their estimated correlation. The Reality Check does the same with the data: resample days jointly for all two hundred rules (so that the correlation between rules is kept), with a stationary bootstrap (so that any dependence over time is kept), centre each rule’s resampled mean on its sample mean (so that the resamples obey the null), and take the maximum tt-statistic of each resample.

With 2 000 stationary-bootstrap resamples, the momentum rules’ P&L being nearly independent over time (the automatic block length is 1.3 days), the best rule’s adjusted pp-value is 0.0275, against 0.030 from the Gaussian simulation, and the 5% critical value of the maximum is 2.85 against 2.92 (Figure 13.4). The Gaussian shortcut was adequate here; with fat-tailed or serially dependent P&L it would not be, and the bootstrap needs no change.

The null law of the best t-statistic among chapter 12’s two hundred momentum rules, estimated by 2 000 stationary-bootstrap resamples of their joint P&L and by 200 000 Gaussian draws with their estimated correlation. At the best rule’s 3.11 (vertical line): 0.0275 and 0.030. Data: the chapter’s tutorial, seeded.
Figure 13.4. The null law of the best tt-statistic among chapter 12’s two hundred momentum rules, estimated by 2 000 stationary-bootstrap resamples of their joint P&L and by 200 000 Gaussian draws with their estimated correlation. At the best rule’s 3.11 (vertical line): 0.0275 and 0.030. Data: the chapter’s tutorial, seeded.

13.6 Tutorial: the interval that was too narrow

Goal. Bootstrap the volatility seller’s Sharpe ratio three ways and measure each interval’s coverage. End state: Figures 13.1, 13.2 and 13.3; coverages 68.5% and 82.75%.

  1. Resampling indices for fixed and geometric blocks, generated for all resamples at once so that a many-column P&L matrix is resampled consistently.

    def block_indices(n: int, n_boot: int, block: int, rng: np.random.Generator) -> np.ndarray:
        """Moving-block bootstrap (Kunsch): concatenate blocks of `block` consecutive days, wrapping at the end."""
        k = -(-n // block)
        starts = rng.integers(0, n, size=(n_boot, k))
        idx = (starts[:, :, None] + np.arange(block)[None, None, :]).reshape(n_boot, k * block)[:, :n]
        return idx % n
    
    
    def stationary_indices(n: int, n_boot: int, mean_block: float, rng: np.random.Generator) -> np.ndarray:
        """Stationary bootstrap (Politis-Romano): each day starts a new block with probability 1 / mean_block,
        at a uniform position; otherwise it continues the current block, wrapping circularly."""
        p = 1.0 / max(1.0, mean_block)
        idx = np.empty((n_boot, n), dtype=np.int64)
        idx[:, 0] = rng.integers(0, n, n_boot)
        new = rng.random((n_boot, n)) < p
        jumps = rng.integers(0, n, size=(n_boot, n))
        for t in range(1, n):
            idx[:, t] = np.where(new[:, t], jumps[:, t], (idx[:, t - 1] + 1) % n)
        return idx
    Listing 13.1. Moving-block and stationary bootstrap indices. code/firm/resample/firm_resample.py
  2. The block length from the sample autocovariances with a flat-top window.

    def optimal_block_length(x) -> tuple[float, float]:
        """Automatic mean block length for the stationary and circular bootstraps: Politis and White (2004)
        with the constants corrected by Patton, Politis and White (2009). The bandwidth M is twice the first
        lag after which K_n consecutive autocorrelations are insignificant."""
        x = np.asarray(x, dtype=float)
        n = x.size
        xc = x - x.mean()
        k_n = max(5, int(math.ceil(math.sqrt(math.log10(n)))))
        m_max = int(math.ceil(math.sqrt(n))) + k_n
        acov = np.array([xc[: n - k] @ xc[k:] / n for k in range(m_max + 1)])
        rho = acov / acov[0]
        crit = 2.0 * math.sqrt(math.log10(n) / n)
        m_hat = m_max - k_n
        for m in range(1, m_max - k_n + 1):
            if np.all(np.abs(rho[m: m + k_n]) < crit):
                m_hat = m - 1
                break
        big_m = min(2 * max(m_hat, 1), m_max)
        k = np.arange(-big_m, big_m + 1)
        lam = _flat_top(k / big_m)
        r = acov[np.abs(k)]
        g = float(np.sum(lam * np.abs(k) * r))
        s0 = float(np.sum(lam * r))
        if g == 0.0:
            return 1.0, 1.0
        b_sb = (2.0 * g * g / (2.0 * s0 * s0)) ** (1 / 3) * n ** (1 / 3)
        b_cb = (2.0 * g * g / (4.0 / 3.0 * s0 * s0)) ** (1 / 3) * n ** (1 / 3)
        cap = math.ceil(min(3 * math.sqrt(n), n / 3))
        return float(min(max(b_sb, 1.0), cap)), float(min(max(b_cb, 1.0), cap))
    Listing 13.2. Automatic mean block length (Politis–White, corrected). code/firm/resample/firm_resample.py
  3. Run problem(), coverage() and coverage(block=40) in qm_resampling.py, then fig_resampling.py.

What to change next. Replace the percentile interval by a studentised bootstrap (resample the tt-statistic, not the ratio) and measure its coverage; set the AR coefficient of log volatility to 0.97 and watch every interval degrade.

13.7 Build: the resampling module

Purpose. Intervals and pp-values for the miniature firm’s statistics that have no trustworthy formula: Sharpe ratios of strategies with dependent P&L, drawdowns, the best of a search.

Interface. iid_indices, block_indices, stationary_indices (each (n, n_boot, …, rng) returning an integer array of shape (n_boot, n)); optimal_block_length(x); bootstrap(stat, x, idx); percentile_interval(draws, level); permutation_test, circular_shift_test; jackknife(stat, x).

Rules. Indices, not resampled data, so that every column of a P&L matrix is resampled with the same days; a generator is always passed in; no interval without its coverage having been measured on a model of the data.

Acceptance tests. code/firm/resample/tests/: geometric block lengths of the requested mean; uniform marginals at every position of a stationary resample; the bootstrap standard error of a mean; the block-length rule on white noise and on an AR(1); the exact size of a permutation test; the jackknife reproduces s2s^2 exactly.

Stretch. Studentised and bias-corrected accelerated intervals; the block jackknife; Hansen’s SPA test on top of stationary_indices and firm.multitest.

Sources and further reading

  • B. Efron, “Bootstrap methods: another look at the jackknife”, Annals of Statistics 7, 1979.
  • H. R. Künsch, “The jackknife and the bootstrap for general stationary observations”, Annals of Statistics 17, 1989.
  • D. N. Politis and J. P. Romano, “The stationary bootstrap”, Journal of the American Statistical Association 89, 1994.
  • D. N. Politis and H. White, “Automatic block-length selection for the dependent bootstrap”, Econometric Reviews 23, 2004; correction by A. Patton, D. N. Politis and H. White, Econometric Reviews 28, 2009.
  • M. H. Quenouille, “Approximate tests of correlation in time-series”, Journal of the Royal Statistical Society B 11, 1949.

13.8 Exercises

Exercise 13.1 ★

What is the probability that a given day is absent from an iid bootstrap resample of nn days? Evaluate it for n=1 260n = 1\,260 and in the limit.

Solution

Solution of Exercise 13.1.

Each of the nn draws misses the day with probability 1−1/n1 - 1/n, so (1−1/n)n(1 - 1/n)^n: 0.3677 for n=1 260n = 1\,260, and e−1=0.3679e^{-1} = 0.3679 in the limit. A resample contains about 63% of the distinct days.

Exercise 13.2 ★

In a stationary bootstrap with mean block length 20 days, what is the probability that a block lasts more than 60 days?

Solution

Solution of Exercise 13.2.

The block continues each day with probability 0.950.95, so it lasts more than 60 days with probability 0.9560=0.0460.95^{60} = 0.046.

Exercise 13.3 ★

A sample has five observations. How many ordered bootstrap resamples are there, and how many distinct ones as multisets?

Solution

Solution of Exercise 13.3.

55=3 1255^5 = 3\,125 ordered resamples; (95)=126\binom{9}{5} = 126 multisets (choosing five draws among five values with repetition).

Exercise 13.4 ★★

Show that the jackknife bias correction of the plug-in variance 1n∑(Xi−Xˉ)2\frac1n\sum(X_i - \bar X)^2 gives s2=1n−1∑(Xi−Xˉ)2s^2 = \frac1{n-1}\sum(X_i - \bar X)^2.

Solution

Solution of Exercise 13.4.

Let S=∑j(Xj−Xˉ)2S = \sum_j(X_j - \bar X)^2 and di=Xi−Xˉd_i = X_i - \bar X. Removing XiX_i leaves the sum of squares S−nn−1di2S - \frac{n}{n-1}d_i^2 about the new mean, so θ^(i)=(S−nn−1di2)/(n−1)\hat\theta_{(i)} = \bigl(S - \frac{n}{n-1}d_i^2\bigr)/(n - 1). Since ∑idi2=S\sum_i d_i^2 = S, the average is θˉ=(S−Sn−1)/(n−1)=(n−2)S(n−1)2\bar\theta = (S - \frac{S}{n-1})/(n - 1) = \frac{(n-2)S}{(n-1)^2}. The corrected estimate is nθ^−(n−1)θˉ=S−(n−2)Sn−1=Sn−1=s2n\hat\theta - (n - 1)\bar\theta = S - \frac{(n-2)S}{n-1} = \frac{S}{n-1} = s^2.

Exercise 13.5 ★★

For an AR(1) with coefficient 0.5 and n=1 260n = 1\,260, compute the optimal mean block length of the stationary bootstrap from b=(G/g(0))2/3n1/3b = (G/g(0))^{2/3}n^{1/3}, with G=∑k∣k∣γ(k)G = \sum_k|k|\gamma(k) and g(0)=∑kγ(k)g(0) = \sum_k\gamma(k).

Solution

Solution of Exercise 13.5.

With γ(k)=γ(0)ϕ∣k∣\gamma(k) = \gamma(0)\phi^{|k|}: G=2γ(0)ϕ/(1−ϕ)2G = 2\gamma(0)\phi/(1 - \phi)^2 and g(0)=γ(0)(1+ϕ)/(1−ϕ)g(0) = \gamma(0)(1 + \phi)/(1 - \phi), so G/g(0)=2ϕ/(1−ϕ2)=4/3G/g(0) = 2\phi/(1 - \phi^2) = 4/3, and b=(4/3)2/3×1 2601/3=1.211×10.80=13.1b = (4/3)^{2/3} \times 1\,260^{1/3} = 1.211 \times 10.80 = 13.1 days.

Exercise 13.6 ★★

For n=100n = 100 uniform observations, what is the probability that a bootstrap resample contains the sample maximum? What does this say about bootstrapping the maximum?

Solution

Solution of Exercise 13.6.

1−0.99100=0.6341 - 0.99^{100} = 0.634. The bootstrap maximum equals the sample maximum 63% of the time: its law has an atom, while the true law of the maximum is continuous, so the bootstrap is inconsistent for extremes.

Exercise 13.7 ★★★

Coding. Over 200 independent pairs of a volatility seller’s P&L and a trailing-variance signal from another market, measure how often a permutation test and a circular-shift test of their correlation reject at 5%.

Solution

Solution of Exercise 13.7.

spurious_rejections() with 199 permutations or shifts per pair: the permutation test rejects in 21.5% of the 200 pairs, four times its nominal size, because both series are persistent and shuffling destroys that; the circular-shift test, which keeps each series’ dependence, rejects in 2.5%.

Exercise 13.8 ★★★

Find the flaw. “We bootstrapped our five years of daily P&L ten thousand times. The 5th percentile of the resampled Sharpe ratios is 0.4, so with 95% confidence the strategy’s Sharpe ratio exceeds 0.4.”

Solution

Solution of Exercise 13.8.

Three flaws. Resampling days independently ignores the clustering of the P&L; for the chapter’s volatility seller such intervals cover the truth 68.5% of the time, not 95%. The bootstrap cannot produce losses larger than the worst day in the sample, so even block intervals are too narrow for a strategy with rare large losses. And the percentile is a statement about this history’s resamples, not a probability about the true Sharpe ratio; if the strategy was chosen among many, chapter 12’s correction applies on top.

13.9 Problem: The Interval That Was Too Narrow

Problem 13.1

Weekend problem — a Sharpe ratio interval for a volatility seller

A strategy earns a premium each day and pays the day’s squared market return; the market’s log volatility follows an AR(1) with coefficient 0.9 and standard deviation 0.4. The population Sharpe ratio is 1.1, and five years of daily P&L are the chapter’s seeded history.

Part I — The history.

  1. What are the sample Sharpe ratio, skewness and kurtosis?
  2. What are the autocorrelations at one, five and twenty days?
  3. Why is the P&L autocorrelated when the market’s returns are not?
  4. What is the population long-run variance of the P&L, in multiples of its variance?
  5. What standard error does the delta method give, using the sample skewness and kurtosis?

Part II — Three intervals.

  1. What are the iid bootstrap interval and standard error?
  2. What block lengths does the automatic rule choose?
  3. What are the moving-block interval and standard error?
  4. What are the stationary-bootstrap interval and standard error?
  5. What do the jackknife standard error and bias give?

Part III — Coverage.

  1. What is the true standard deviation of the five-year Sharpe ratio?
  2. How often do iid intervals contain 1.1, over 400 histories?
  3. And stationary intervals, with the automatic length and with 40-day blocks?
  4. Why does no block length reach 95%?
  5. What happens to the iid interval without volatility clustering?

Part IV — Judgement.

  1. What interval do you give the head of desk?
  2. Should the capital be doubled on this evidence?
  3. What would you add to the analysis before deciding?
  4. State the named result: the coverage of the nominal 95% iid and stationary-bootstrap intervals.
  5. In one sentence: what must a resampling scheme preserve?
Solution

Solution of Problem 13.1.

1. Sharpe ratio 1.11, skewness −4.7-4.7, kurtosis 34. 2. 0.13, 0.10 and −0.02-0.02. 3. The strategy pays rt2r_t^2, and squared returns are autocorrelated because volatility is persistent: quiet days follow quiet days, losses follow losses. 4. 3.9 times the variance (1+2∑k(e0.64⋅0.9k−1)/(3e0.64−1)1 + 2\sum_k(e^{0.64 \cdot 0.9^k} - 1)/(3e^{0.64} - 1)). 5. 0.52: the skewness term −γ3SR-\gamma_3\mathrm{SR} adds to the variance, but the formula still assumes independent days. 6. (0.17,2.25)(0.17, 2.25), standard error 0.53. 7. Mean block 18.6 days for the stationary bootstrap, 21.2 for the circular block bootstrap. 8. (−0.34,2.93)(-0.34, 2.93), standard error 0.84. 9. (−0.25,2.96)(-0.25, 2.96), standard error 0.83. 10. Standard error 0.53 (the iid answer: single-day deletions ignore the dependence) and bias 0.04, so a corrected ratio of 1.06. 11. 1.10, from 4 000 simulated histories. 12. 68.5% of the time. 13. 82.75% with the automatic block length (moving-block the same), 84.0% with 40-day blocks. 14. The bootstrap standard error plateaus near 0.97 against a truth of 1.10: histories without a violent episode cannot resample one, and five years hold few independent volatility clusters. 15. It covers 91.5%; the rest of the shortfall is the heavy tails of the P&L. 16. The stationary interval, (−0.25,2.96)(-0.25, 2.96), stated as a lower bound on the uncertainty, with its measured coverage on a model of the strategy. 17. No: the interval contains zero, and the true uncertainty is larger still; a doubling would rest on the iid interval, which is the wrong one. 18. Stress tests with volatility regimes not in the sample, the strategy’s history through past volatility spikes, a longer backtest, and the search that produced the strategy. 19. Named result: the interval that was too narrow: nominal 95% bootstrap intervals for the volatility seller’s Sharpe ratio contain the truth 68.5% of the time when days are resampled independently and 82.75% with the stationary bootstrap’s automatic blocks. 20. It must preserve every dependence the null hypothesis or the estimand does not deny: within days across assets, and across days in time.

13.10 Interview questions

Interview question 13.1 ★ researcher, mle

What does the bootstrap do, and when does it fail?

Solution

Solution of Interview question 13.1.

It estimates the sampling law of a statistic by recomputing it on resamples drawn with replacement from the data, the empirical law standing in for the true one. It fails for statistics driven by extremes (the maximum, tail quantiles), for dependent data resampled as if independent, and for very small samples.

What the interviewer is looking for: The plug-in principle, and at least two failure modes.

Interview question 13.2 ★★ researcher, trader

Why not bootstrap a strategy’s daily returns independently to get an interval for its Sharpe ratio?

Solution

Solution of Interview question 13.2.

Daily P&L is dependent: volatility clusters, positions overlap, edges come in regimes. Independent resampling estimates the short-run variance, and the interval is too narrow; in the chapter’s example it covers 68.5% instead of 95%. Resample blocks (stationary bootstrap), and check coverage by simulation.

What the interviewer is looking for: Long-run versus short-run variance, block resampling, coverage as the test.

Interview question 13.3 ★★ researcher

How do you choose the block length of a block bootstrap?

Solution

Solution of Interview question 13.3.

The error of the variance estimate is minimised by a length growing like n1/3n^{1/3} with a constant set by the autocovariances; estimate it (Politis–White with the 2009 correction), then check sensitivity by plotting the result against the block length. Persistent, heavy-tailed data need longer blocks than the rule suggests.

What the interviewer is looking for: The n1/3n^{1/3} rate, an automatic rule, and a sensitivity check.

Interview question 13.4 ★★ researcher, mle

You test whether a signal predicts returns by shuffling the signal. When does that test lie?

Solution

Solution of Interview question 13.4.

When the series are serially dependent: shuffling makes the signal exchangeable although it is not, and spurious correlation between two persistent series is far more common than the shuffles show. In the chapter a permutation test rejects 21.5% of independent pairs at 5%. Use circular shifts or block permutations.

What the interviewer is looking for: Exchangeability as the condition, and a resampling scheme that keeps the dependence.

Interview question 13.5 ★★ researcher

Jackknife or bootstrap: what does each give you?

Solution

Solution of Interview question 13.5.

The jackknife: nn deterministic leave-one-out recomputations, a bias correction that removes the 1/n1/n term, and a standard error for smooth statistics. The bootstrap: a whole distribution, hence intervals and tests, at the cost of randomness and more computation. Both need blocks for dependent data.

What the interviewer is looking for: Bias versus distribution, and the shared need for blocks.

Interview question 13.6 ★★★ researcher

How would you compute a pp-value for the best of two hundred backtests that respects both the correlation between them and the dependence over time?

Solution

Solution of Interview question 13.6.

Resample days jointly across all two hundred strategies (keeping their correlation) with a stationary bootstrap (keeping time dependence), centre each strategy’s resampled mean on its sample mean, compute the maximum tt-statistic per resample, and read the pp-value of the observed maximum from that distribution: White’s Reality Check, or Hansen’s SPA with studentisation and recentring, or Romano–Wolf for step-down.

What the interviewer is looking for: Joint resampling, recentring to impose the null, the maximum as the statistic.

Terms defined in this chapter

See all 2333 terms in the glossary