Quantitative Finance · Book 4 · Methods

Quantitative Methods

Quantitative Methods · Methods

1Probability at Speed

At ten to four the imbalance of the closing auction appears on the screen, and the desk’s running estimate of the closing price, a number it has been revising all day, jumps by 14 cents. A researcher asks a sharper question than whether the estimate was right: could its revisions have been predicted from the revisions before them? If they could, the estimate was not using what it knew, and someone trading against it would have made money. An honest forecast of a fixed quantity is a conditional expectation along a flow of information, and such a process has one defining property: its revisions cannot be forecast. This chapter sets out, at the speed of a reader who has met measure theory, the objects the series computes with (information, conditional expectation, martingales, stopping times, changes of measure) and closes by fixing the notation every later book uses.

1.1 Information: sigma-algebras, filtrations and conditional expectation

A probability space (Ω,F,P)(\Omega, \mathcal F, \P) is taken as known. What finance adds is time: at each instant some events are decided, and the decided ones grow.

Definition 1.1 (Filtration, adapted process)

A filtration is a family F=(Ft)t∈T\mathbb F = (\mathcal F_t)_{t \in \mathbb T} of sub-σ\sigma-algebras of F\mathcal F, increasing in tt: Fs⊆Ft\mathcal F_s \subseteq \mathcal F_t for s≤ts \le t. The index set T\mathbb T is {0,1,…,n}\{0, 1, \dots, n\} or [0,T][0, T]; in continuous time the filtration is assumed right-continuous and complete (the usual conditions). A process (Xt)(X_t) is an adapted process if each XtX_t is Ft\mathcal F_t-measurable: its value at tt is known at tt.

Ft\mathcal F_t is what is known at tt: the prices printed so far, the cards turned over. A trading rule must be adapted; a backtest that is not has looked into the future (chapter 3).

Definition 1.2 (Conditional expectation)

Let XX be integrable and G⊆F\mathcal G \subseteq \mathcal F a sub-σ\sigma-algebra. The conditional expectation E[X∣G]\E[X \mid \mathcal G] is the almost surely unique G\mathcal G-measurable, integrable random variable YY with E[X1A]=E[Y1A]\E[X\mathbf 1_A] = \E[Y\mathbf 1_A] for every A∈GA \in \mathcal G. We write Et[X]=E[X∣Ft]\E_t[X] = \E[X \mid \mathcal F_t].

Existence is the Radon–Nikodym theorem applied to the measure A↦E[X1A]A \mapsto \E[X\mathbf 1_A] on G\mathcal G. For square-integrable XX there is a more useful picture: E[X∣G]\E[X \mid \mathcal G] is the orthogonal projection of XX onto the closed subspace L2(G)L^2(\mathcal G), the best forecast of XX in mean square among all G\mathcal G-measurable random variables. Everything a forecaster can compute from the information G\mathcal G is in that subspace, and the forecast error X−E[X∣G]X - \E[X \mid \mathcal G] is orthogonal to all of it.

Proposition 1.3 (Rules of conditional expectation)

For integrable X,YX, Y and H⊆G\mathcal H \subseteq \mathcal G: (i) tower: E[E[X∣G]∣H]=E[X∣H]\E[\E[X \mid \mathcal G] \mid \mathcal H] = \E[X \mid \mathcal H]; (ii) taking out what is known: if YY is G\mathcal G-measurable and XYXY integrable, E[XY∣G]=YE[X∣G]\E[XY \mid \mathcal G] = Y\E[X \mid \mathcal G]; (iii) independence: if XX is independent of G\mathcal G, E[X∣G]=E[X]\E[X \mid \mathcal G] = \E[X]; (iv) Jensen: for convex φ\varphi with φ(X)\varphi(X) integrable, φ(E[X∣G])≤E[φ(X)∣G]\varphi(\E[X \mid \mathcal G]) \le \E[\varphi(X) \mid \mathcal G].

Proof. Check the defining integrals: (i) for A∈H⊆GA \in \mathcal H \subseteq \mathcal G, E[E[X∣G]1A]=E[X1A]\E[\E[X \mid \mathcal G]\mathbf 1_A] = \E[X\mathbf 1_A]; (ii) for Y=1BY = \mathbf 1_B, B∈GB \in \mathcal G, then limits; (iii) E[X1A]=E[X]P(A)\E[X\mathbf 1_A] = \E[X]\P(A); (iv) φ\varphi is a supremum of countably many affine functions. ∎

Figure 1.1 shows the rules at work on the smallest example that has them all. A price starts at 100 and moves one dollar up or down on each of two days with equal probability; X=(S2−100)+X = (S_2 - 100)^+. After one day the information is which half of the tree we are in, and E1[X]\E_1[X] averages XX over that half.

Conditional expectation as averaging over what is still unknown. F_1 splits the four outcomes into two atoms (dashed); _1[X] is the average of X on the atom reached, and _0[X] is the average of _1[X]: the tower rule. The three values 0.5 1 2 along the top path form a Doob martingale.
Figure 1.1. Conditional expectation as averaging over what is still unknown. F1\mathcal F_1 splits the four outcomes into two atoms (dashed); E1[X]\E_1[X] is the average of XX on the atom reached, and E0[X]\E_0[X] is the average of E1[X]\E_1[X]: the tower rule. The three values 0.5→1→20.5 \to 1 \to 2 along the top path form a Doob martingale.

Laws are handled through their transforms. The one the series uses most is defined here, so that later books (Lévy processes in chapter 6, Fourier pricing in chapter 28, stochastic volatility in One Quant Book 5) can point to it.

Definition 1.4 (Characteristic function)

The characteristic function of a random variable XX is φX(u)=E[eiuX]\varphi_X(u) = \E[e^{\iu uX}], u∈Ru \in \R; of a random vector, φX(u)=E[eiu⊤X]\varphi_X(u) = \E[e^{\iu u^\top X}], u∈Rdu \in \R^d. It exists for every law, determines it, and turns sums of independent variables into products.

Theorem 1.5 (Central limit theorem)

Let X1,X2,…X_1, X_2, \dots be independent with means μi\mu_i and variances σi2\sigma_i^2, sn2=∑i≤nσi2s_n^2 = \sum_{i \le n}\sigma_i^2, and suppose Lindeberg’s condition sn−2∑i≤nE[(Xi−μi)21{∣Xi−μi∣>εsn}]→0s_n^{-2}\sum_{i\le n}\E[(X_i - \mu_i)^2\mathbf 1_{\{|X_i - \mu_i| > \varepsilon s_n\}}] \to 0 holds for every ε>0\varepsilon > 0 (it does for identically distributed variables with finite variance). Then sn−1∑i≤n(Xi−μi)→dN(0,1)s_n^{-1}\sum_{i \le n}(X_i - \mu_i) \xrightarrow{d} \mathcal N(0, 1).

Proof. Admitted here. ∎

The proof expands the characteristic function to second order and uses Lévy’s continuity theorem (Williams, 1991). Every standard error in Part II rests on it.

1.2 Martingales and the Doob martingale of a forecast

Definition 1.6 (Martingale)

An adapted, integrable process (Mt)(M_t) is a martingale if Es[Mt]=Ms\E_s[M_t] = M_s for all s≤ts \le t; a submartingale if Es[Mt]≥Ms\E_s[M_t] \ge M_s; a supermartingale if Es[Mt]≤Ms\E_s[M_t] \le M_s.

A martingale is a fair game seen from any date. Four examples recur in the book. With ξi\xi_i independent, taking the values ±1\pm 1 with probabilities pp and q=1−pq = 1 - p, and Sn=∑i≤nξiS_n = \sum_{i \le n}\xi_i:

  • SnS_n is a martingale when p=12p = \tfrac12, a submartingale when p>12p > \tfrac12;
  • Sn2−nS_n^2 - n is a martingale when p=12p = \tfrac12: the variance grows by one per step;
  • (q/p)Sn(q/p)^{S_n} is a martingale for every pp, since E[(q/p)ξ]=p q/p+q p/q=1\E[(q/p)^{\xi}] = p\,q/p + q\,p/q = 1;
  • exp⁡(θSn−nln⁡cosh⁡θ)\exp(\theta S_n - n\ln\cosh\theta) is a martingale when p=12p = \tfrac12, for every real θ\theta: the exponential martingale, prototype of the density processes of Section 1.4.

The example that matters most for a trading desk is not a price but a forecast.

Definition 1.7 (Doob martingale)

For an integrable random variable XX and a filtration F\mathbb F, the Doob martingale of XX is Mt=Et[X]M_t = \E_t[X].

It is a martingale by the tower rule. Every honest forecast of a fixed quantity (a closing price, the sum of five cards, next month’s inflation print) is a Doob martingale along the forecaster’s filtration, and what makes it testable is that its revisions are orthogonal.

Proposition 1.8 (Orthogonal increments)

Let (Mk)k=0,…,n(M_k)_{k = 0,\dots,n} be a square-integrable martingale with increments Dk=Mk−Mk−1D_k = M_k - M_{k-1}. Then E[DkY]=0\E[D_k Y] = 0 for every square-integrable Fk−1\mathcal F_{k-1}-measurable YY; in particular E[DjDk]=0\E[D_jD_k] = 0 for j<kj < k, and

Var⁡(Mn)=Var⁡(M0)+∑k=1nE[Dk2].\Var(M_n) = \Var(M_0) + \sum_{k=1}^n \E[D_k^2].

Proof. E[DkY]=E[YEk−1[Dk]]=0\E[D_kY] = \E[Y\E_{k-1}[D_k]] = 0 by taking out what is known. With Y=DjY = D_j, j<kj < k, the cross terms of (Mn−M0)2=(∑kDk)2(M_n - M_0)^2 = (\sum_k D_k)^2 vanish, and Cov⁡(M0,Dk)=0\Cov(M_0, D_k) = 0 likewise. ∎

Two consequences drive this chapter’s tutorial and problem. First, a forecast whose revisions can be predicted from the past (a regression of DkD_k on Dk−1D_{k-1} with a nonzero slope) is not a conditional expectation: it is leaving information unused. Second, the variance of the final value splits into the variances of the revisions, so one can say how much of a closing price’s uncertainty is resolved at each time of day (Figure 1.2).

Left: one simulated day of M_t = _t[ close] in the weekend problem’s model (a random walk with a daily standard deviation of USD 1.80 on a USD 100 stock, and news on the closing imbalance with standard deviation USD 0.30 at 15:50, +0.48 on this day). Right: the share of the close’s variance resolved by each time of day, by : linear through the session, a step of 2.7% at the news, the last 2.5% in the final ten minutes. Data: the chapter’s tutorial, seeded.
Figure 1.2. Left: one simulated day of Mt=Et[close]M_t = \E_t[\text{close}] in the weekend problem’s model (a random walk with a daily standard deviation of USD 1.80 on a USD 100 stock, and news on the closing imbalance with standard deviation USD 0.30 at 15:50, +0.48+0.48 on this day). Right: the share of the close’s variance resolved by each time of day, by Proposition 1.8: linear through the session, a step of 2.7% at the news, the last 2.5% in the final ten minutes. Data: the chapter’s tutorial, seeded.

The maximal inequality bounds how far a martingale wanders over a period by where it ends; chapter 2 uses it to control Brownian paths.

Proposition 1.9 (Doob’s maximal inequality)

For a nonnegative submartingale (Xk)k≤n(X_k)_{k \le n} and c>0c > 0, c P(max⁡k≤nXk≥c)≤E[Xn1{max⁡kXk≥c}]≤E[Xn]c\,\P(\max_{k\le n}X_k \ge c) \le \E[X_n\mathbf 1_{\{\max_k X_k \ge c\}}] \le \E[X_n]; and for p>1p > 1, E[max⁡k≤nXkp]≤(p/(p−1))pE[Xnp]\E[\max_{k \le n}X_k^p] \le (p/(p-1))^p\E[X_n^p]. For a square-integrable martingale, applied to ∣Mk∣|M_k|: E[max⁡k≤nMk2]≤4E[Mn2]\E[\max_{k\le n}M_k^2] \le 4\E[M_n^2].

Proof. Admitted here. ∎

1.3 Stopping times and optional stopping

Definition 1.10 (Stopping time)

A random time τ\tau with values in T∪{∞}\mathbb T \cup \{\infty\} is a stopping time if {τ≤t}∈Ft\{\tau \le t\} \in \mathcal F_t for every tt: whether it has occurred by tt is known at tt.

“The first time the price touches 101” is a stopping time; “the time of the day’s high” is not, since it is known only at the close. When is a martingale’s expected value at a stopping time still its starting value?

Definition 1.11 (Uniform integrability)

A family (Xi)i∈I(X_i)_{i \in I} of random variables has uniform integrability if sup⁡iE[∣Xi∣1{∣Xi∣>c}]→0\sup_{i}\E[|X_i|\mathbf 1_{\{|X_i| > c\}}] \to 0 as c→∞c \to \infty.

Theorem 1.12 (Optional stopping)

Let (Mn)(M_n) be a martingale and τ\tau a stopping time. Then E[Mτ]=E[M0]\E[M_\tau] = \E[M_0] if either (i) τ\tau is bounded, or (ii) τ<∞\tau < \infty almost surely and the stopped process (Mn∧τ)(M_{n \wedge \tau}) is uniformly integrable; in particular if ∣Mn∧τ∣≤c|M_{n\wedge\tau}| \le c for all nn.

Proof. (i) If τ≤N\tau \le N, then Mτ=M0+∑k=1NDk1{τ≥k}M_\tau = M_0 + \sum_{k=1}^N D_k\mathbf 1_{\{\tau \ge k\}}, and {τ≥k}={τ≤k−1}c∈Fk−1\{\tau \ge k\} = \{\tau \le k-1\}^c \in \mathcal F_{k-1}, so each term has zero mean by Proposition 1.8. (ii) Apply (i) to τ∧n\tau \wedge n; Mn∧τ→MτM_{n \wedge \tau} \to M_\tau almost surely, and uniform integrability upgrades this to convergence in L1L^1 (Vitali), so the means converge. ∎

The theorem prices every take-profit and stop-loss rule on a martingale. A position whose value moves by ±1\pm 1 tick with equal probability is closed at +a+a or −b-b. The stopped value is bounded, so E[Sτ]=0\E[S_\tau] = 0: with P(up first)=π\P(\text{up first}) = \pi, πa−(1−π)b=0\pi a - (1 - \pi)b = 0 and

π=ba+b.\pi = \frac{b}{a+b}.

A stop close to entry and a distant profit target win rarely and lose often, and the expected profit is zero whatever aa and bb are: the rule reshapes the distribution of the result, not its mean. With a drift (p≠12p \ne \tfrac12) the martingale (q/p)Sn(q/p)^{S_n} gives π=(1−(q/p)b)/(1−(q/p)a+b)\pi = (1 - (q/p)^b)/(1 - (q/p)^{a+b}); Figure 1.3 compares both formulas with simulation.

Probability that a ±1-tick random walk reaches a take-profit 10 ticks away before a stop-loss b ticks away: the optional-stopping formulas (lines) and 20 000 seeded paths per point (marks). A 1-point drift against the position (p = 0.49) costs 10 points of probability at b = 10. Data: the chapter’s tutorial.
Figure 1.3. Probability that a ±1\pm1-tick random walk reaches a take-profit 10 ticks away before a stop-loss bb ticks away: the optional-stopping formulas (lines) and 20 000 seeded paths per point (marks). A 1-point drift against the position (p=0.49p = 0.49) costs 10 points of probability at b=10b = 10. Data: the chapter’s tutorial.

Remark 1.13 (The doubling strategy)

Doubling the stake after each loss of a fair even-money bet and stopping at the first win yields +1+1 almost surely, apparently contradicting the theorem. The stopping time is finite but unbounded, and the stopped process is not uniformly integrable: before the win it has lost 2k−12^k - 1 with probability 2−k2^{-k}. A finite credit line bounds the stakes, makes condition (ii) hold, and restores E[Mτ]=0\E[M_\tau] = 0: a small, frequent gain paid for by a rare, enormous loss.

1.4 Changing the measure

Much of quantitative finance computes under a measure other than the one that describes the world: one under which discounted prices are martingales (chapter 5), or one under which a rare loss is common enough to simulate (chapter 26).

Definition 1.14 (Change of measure)

Two probability measures P\P and Q\mathbb Q on (Ω,F)(\Omega, \mathcal F) are equivalent measures if they have the same null sets. A change of measure from P\P to an equivalent Q\mathbb Q is described by the Radon–Nikodym derivative Z=dQ/dPZ = d\mathbb Q/d\P, the almost surely unique positive random variable with Q(A)=E[Z1A]\mathbb Q(A) = \E[Z\mathbf 1_A] for all A∈FA \in \mathcal F; then EQ[X]=E[ZX]\E^{\mathbb Q}[X] = \E[ZX]. Along a filtration, the density process is Zt=Et[Z]Z_t = \E_t[Z], the Radon–Nikodym derivative of Q\mathbb Q restricted to Ft\mathcal F_t.

The density process is a positive P\P-martingale with Z0=1Z_0 = 1 (a Doob martingale), and conversely every such martingale defines a measure. Conditional expectations under the new measure follow from the old ones.

Proposition 1.15 (Bayes formula for a change of measure)

If Q∼P\mathbb Q \sim \P with density process (Zt)(Z_t) and XX is Q\mathbb Q-integrable, then for s≤ts \le t and XX Ft\mathcal F_t-measurable,

EsQ[X]=Es[ZtX]Zs.\E^{\mathbb Q}_s[X] = \frac{\E_s[Z_tX]}{Z_s}.

In particular, an adapted (Xt)(X_t) is a Q\mathbb Q-martingale if and only if (ZtXt)(Z_tX_t) is a P\P-martingale.

Proof. The right-hand side YY is Fs\mathcal F_s-measurable. For A∈FsA \in \mathcal F_s,

EQ[Y1A]=E[ZsY1A]=E[Es[ZtX]1A]=E[ZtX1A]=EQ[X1A],\E^{\mathbb Q}[Y\mathbf 1_A] = \E[Z_sY\mathbf 1_A] = \E\bigl[\E_s[Z_tX]\mathbf 1_A\bigr] = \E[Z_tX\mathbf 1_A] = \E^{\mathbb Q}[X\mathbf 1_A],

using E[ZW]=E[ZsW]\E[Z W] = \E[Z_s W] for Fs\mathcal F_s-measurable WW, and the tower rule. ∎

Example 1.16 (Shifting a Gaussian, and seeing a far tail)

Let X∼N(0,1)X \sim \mathcal N(0, 1) under P\P and Z=exp⁡(cX−c2/2)Z = \exp(cX - c^2/2). Then EQ[eiuX]=E[e(c+iu)X−c2/2]=eiuc−u2/2\E^{\mathbb Q}[e^{\iu uX}] = \E[e^{(c + \iu u)X - c^2/2}] = e^{\iu uc - u^2/2}: under Q\mathbb Q, X∼N(c,1)X \sim \mathcal N(c, 1). The change of measure has moved the mean without touching the shape. Run it backwards to estimate p=P(X>4)=3.17×10−5p = \P(X > 4) = 3.17 \times 10^{-5}: sample Y∼N(4,1)Y \sim \mathcal N(4, 1) and average e−4Y+81{Y>4}e^{-4Y + 8}\mathbf 1_{\{Y > 4\}}, the indicator reweighted by dP/dQd\P/d\mathbb Q. With 100 000 draws, plain sampling sees three exceedances and has a standard error of 1.7×10−51.7 \times 10^{-5}, half the answer; the reweighted estimate has a standard error of 2.1×10−72.1 \times 10^{-7}, eighty times smaller. Chapter 26 turns this into importance sampling, and chapter 5 does the same computation for whole Brownian paths.

1.5 The notation of the series

Every book of the series from this one on uses the symbols below. They extend the tables of One Quant Books 1 and 2, whose meanings are kept: sts_t is the bid–ask spread, S\mathcal S a credit spread, τ=T−t\tau = T - t a time to expiry, rr a rate, ℓ\ell a borrow fee, KK a strike, P(t,T)P(t,T) a discount factor, yy a yield, δ\delta an accrual fraction and RR a recovery rate. A symbol a chapter declares local may carry another meaning there.

Notation 1.17 (Probability, measures and processes)

(Ω,F,P)(\Omega,\mathcal F,\P), F=(Ft)\mathbb F = (\mathcal F_t)probability space; filtration (usual conditions)
Et[X]=E[X∣Ft]\E_t[X] = \E[X \mid \mathcal F_t], EtQ\E^{\mathbb Q}_tconditional expectation; under another measure
P\P, Q\mathbb Q, Bt=e∫0trs dsB_t = e^{\int_0^t r_s\,ds}real-world measure; risk-neutral measure; money-market account
Nt\mathcal N_t, QN\mathbb Q^{\mathcal N}; QT\mathbb Q^T, QA\mathbb Q^Ageneric numeraire and its measure; TT-forward and annuity measures
dQ/dPd\mathbb Q/d\P, ZtZ_tRadon–Nikodym derivative; density process
Φ\Phi, φ\varphi; N(m,s2)\mathcal N(m, s^2)standard normal cdf and pdf (never N(d1)N(d_1)); the normal law, always with arguments
φX(u)=E[eiuX]\varphi_X(u) = \E[e^{\iu uX}]characteristic function, always subscripted
1A\mathbf 1_A; =d\overset{d}{=}, →d\xrightarrow{d}, →P\xrightarrow{\P}indicator; equality and convergence in law, in probability
τ\taua stopping time; where it meets a time to expiry, write T−tT - t; default times are subscripted (τC\tau_C)
WtW_t; WtQW^{\mathbb Q}_tBrownian motion under the measure in force; decorated when two measures appear
[X]t[X]_t, [X,Y]t[X,Y]_t, E(X)t\mathcal E(X)_tquadratic variation, covariation, stochastic exponential
dX=μ dt+σ dWdX = \mu\,dt + \sigma\,dW; L\mathcal L, L∗\mathcal L^*a diffusion; its generator and the adjoint
dX=κ(xˉ−X) dt+σ dWdX = \kappa(\bar x - X)\,dt + \sigma\,dWOrnstein–Uhlenbeck: speed κ\kappa, level xˉ\bar x, half-life ln⁡2/κ\ln 2/\kappa
dv=κ(vˉ−v) dt+ηv dWdv = \kappa(\bar v - v)\,dt + \eta\sqrt v\,dWsquare-root process; Feller condition 2κvˉ≥η22\kappa\bar v \ge \eta^2; η\eta is the vol-of-vol throughout
NtN_t, t1<t2<…t_1 < t_2 < \dotscounting process and its event times
λt\lambda_t, Λt=∫0tλs ds\Lambda_t = \int_0^t\lambda_s\,dsintensity (Poisson rate, hazard rate, Hawkes intensity λt=μ+∑ti<tg(t−ti)\lambda_t = \mu + \sum_{t_i < t}g(t - t_i)); compensator
(σ2,ν,γ)(\sigma^2, \nu, \gamma), ψ\psiLévy triplet and characteristic exponent, φXt(u)=etψ(u)\varphi_{X_t}(u) = e^{t\psi(u)}
HH, WtHW^H_tHurst exponent; fractional Brownian motion

Notation 1.18 (Statistics, matrices and numerics)

nn, θ∈Θ\theta \in \Theta, θ0\theta_0sample size, parameter, true value; hat an estimate, tilde a shrunk estimate, bar a sample mean
ℓn(θ)\ell_n(\theta), I(θ)\mathcal I(\theta)log-likelihood (always with subscript and argument); Fisher information
se(θ^)\mathrm{se}(\hat\theta); SR\mathrm{SR}, SR^\widehat{\mathrm{SR}}standard error; Sharpe ratio and its estimate
Rt=ln⁡(St/St−1)R_t = \ln(S_t/S_{t-1})log return (RR alone stays the recovery rate)
LL, Δ=1−L\Delta = 1 - L; γ(h)\gamma(h), ρ(h)\rho(h)lag and difference operators; autocovariance, autocorrelation
ϕi\phi_i, ϑj\vartheta_j, εt\varepsilon_tautoregressive and moving-average coefficients (always indexed); innovations
RVt\mathrm{RV}_t, IVt\mathrm{IV}_t; σimp\sigma_{\mathrm{imp}}realised and integrated variance; implied volatility (never “IV” in a formula)
Σ\Sigma, CC, Σ^\hat\Sigma; 1\mathbf 1, InI_n, ⊤{}^\topcovariance, correlation, sample covariance; ones vector, identity, transpose; vectors are columns
λ1≥⋯≥λN\lambda_1 \ge \dots \ge \lambda_N, cond⁡(A)\operatorname{cond}(A)eigenvalues (always indexed); condition number
min⁡f(x)\min f(x) s.t. gi(x)≤0g_i(x) \le 0, hj(x)=0h_j(x) = 0optimisation; multipliers u≥0u \ge 0, vv; optimum x⋆x^\star, values p⋆p^\star, d⋆d^\star
fl(x)\mathrm{fl}(x), u=2−53u = 2^{-53}, εmach=2−52\varepsilon_{\mathrm{mach}} = 2^{-52}floating point in binary64, round to nearest; ulp(x)\mathrm{ulp}(x)
tk=kΔtt_k = k\Delta t, xjx_j, VjkV^k_j; MM, V^M\hat V_Mtime grid, space grid, numerical solution; Monte Carlo paths and estimator

The options, rates, credit and risk symbols (σimp(K,T)\sigma_{\mathrm{imp}}(K,T), k=ln⁡(K/F0,T)k = \ln(K/F_{0,T}), ξt(u)\xi_t(u), the Greeks with the rate Greek written Rho because ρ\rho is always a correlation, Sa,b(t)S_{a,b}(t), Aa,b(t)A_{a,b}(t), QC(t)Q_C(t), VaRα,h\mathrm{VaR}_{\alpha,h}, ESα,h\mathrm{ES}_{\alpha,h}) come from the same table and are introduced in One Quant Books 5 and 6.

1.6 Tutorial: the Doob martingale of a card game

Goal. Build the Doob martingale of the five-card game of One Quant Book 2, chapter 30 (the sum of five cards dealt from a 52-card deck, aces 1 to kings 13), check that its revisions behave as Proposition 1.8 says, and catch a forecaster that under-reacts. End state: the table below and Figures 1.2 and 1.3.

  1. The forecast. After kk cards with sum sks_k, the remaining 52−k52 - k cards have sum 364−sk364 - s_k, so Mk=sk+(5−k)(364−sk)/(52−k)M_k = s_k + (5 - k)(364 - s_k)/(52 - k), and M0=35M_0 = 35. The under-reacting forecaster moves only a fraction aa of the way from 35 until the last card.

    DECK = np.repeat(np.arange(1, 14), 4)          # aces 1 ... kings 13, four suits
    N_CARDS = 5
    
    
    # --- the card game -------------------------------------------------------------------------
    def card_forecasts(n_deals: int, seed: int = 1) -> np.ndarray:
        """Doob martingale M_k = E[sum of 5 cards | first k revealed], k = 0..5, one row per deal."""
        rng = np.random.default_rng(seed)
        total, size = DECK.sum(), DECK.size
        out = np.empty((n_deals, N_CARDS + 1))
        for i in range(n_deals):
            cards = rng.choice(DECK, N_CARDS, replace=False)
            s = np.concatenate([[0], np.cumsum(cards)])
            k = np.arange(N_CARDS + 1)
            out[i] = s + (N_CARDS - k) * (total - s) / (size - k)
        return out
    
    
    def underreacting(m: np.ndarray, a: float) -> np.ndarray:
        """A forecaster that moves only a fraction a of the way from the prior mean, until the end."""
        f = m[:, :1] + a * (m - m[:, :1])
        f[:, -1] = m[:, -1]
        return f
    Listing 1.1. The Doob martingale of the card game and an under-reacting forecaster. code/methods/01-probability-at-speed/python/qm_martingales.py
  2. The test. Regress the last revision on the one before it, pooled over deals; the running project’s function does it for any panel of forecast paths.

    def revision_regression(paths: np.ndarray, lag: int = 1, col: int | None = None) -> dict:
        """Pooled regression (no intercept) of each revision on the revision `lag` steps before it.
    
        With `col` given, regress only the revision at column `col` (of the revision array) on the one
        `lag` steps before it. For a martingale the slope is zero; the t-statistic uses the iid
        standard error of a no-intercept regression."""
        d = revisions(paths)
        if col is None:
            y = d[:, lag:].ravel()
            x = d[:, :-lag].ravel()
        else:
            y = d[:, col]
            x = d[:, col - lag]
        sxx = float(x @ x)
        slope = float(x @ y) / sxx
        resid = y - slope * x
        n = y.size
        se = float(np.sqrt(resid @ resid / (n - 1) / sxx))
        return {"slope": slope, "se": se, "t": slope / se, "n": n}
    Listing 1.2. Regression of a revision on an earlier one: zero slope for a martingale. code/firm/mgtest/firm_mgtest.py
  3. Run card_table() over 20 000 seeded deals, then fig_martingales.py for the figures.
honest forecastunder-reacting (a=0.6a = 0.6)
mean revision0.0010.0010.0010.001
slope of last revision on the fourth−0.005-0.0050.6610.661 (theory 0.6670.667)
tt-statistic of the slope—45.945.9

Both forecasts have revisions with mean zero; only the regression tells them apart. What to change next. Let the forecaster over-react (a>1a > 1) and predict the sign of the slope; replace the card game by the closing-price model of the weekend problem and find how many days of data the test needs.

1.7 Build: martingale diagnostics

Purpose. Every forecast the miniature firm publishes (a fair value, an expected closing price, a predicted fill rate) is checked for the martingale property before anyone trades on it or against it.

Interface. revisions(paths); revision_regression(paths, lag=1, col=None) returning slope, standard error, tt and nn; variance_ratio(increments, q); variance_shares(paths); martingale_report(paths). A panel has one row per day or deal and one column per revision time.

Rules. No intercept in the revision regression (the mean revision is reported separately); variance shares are those of the revisions, which for a martingale sum to the variance of the final value; everything seeded and vectorised.

Acceptance tests. code/firm/mgtest/tests/: random walks pass; a forecast that adds half of the news late is caught with the right slope; shares are uniform for iid revisions; the variance ratio is one for iid increments and below one for mean-reverting ones.

Stretch. Heteroskedasticity- and autocorrelation-robust standard errors (chapter 11); a test of the revisions against any Fk−1\mathcal F_{k-1}-measurable signal, not only past revisions; multiple-testing control across many forecasts (chapter 12).

Sources and further reading

  • A. N. Kolmogorov, Grundbegriffe der Wahrscheinlichkeitsrechnung, Springer, 1933: the measure-theoretic axioms.
  • J. Ville, Étude critique de la notion de collectif, Gauthier-Villars, 1939, where the word martingale enters probability.
  • J. L. Doob, Stochastic Processes, Wiley, 1953.
  • O. Nikodym, “Sur une généralisation des intégrales de M. J. Radon”, Fundamenta Mathematicae 15, 1930.
  • D. Williams, Probability with Martingales, Cambridge University Press, 1991: the proofs admitted here.

1.8 Exercises

Exercise 1.1 ★

In the five-card game, what is the forecast of the sum after the first card turns out to be a king?

Solution

Solution of Exercise 1.1.

M1=13+4×(364−13)/51=13+27.53=40.53M_1 = 13 + 4 \times (364 - 13)/51 = 13 + 27.53 = 40.53, against 35 before the deal.

Exercise 1.2 ★

Which of these are stopping times for the filtration of observed prices: (a) the first time the price is 1% above the open; (b) the time of the day’s low; (c) 15:50; (d) the first time after 15:50 that the forecast of the close moves by more than 10 cents?

Solution

Solution of Exercise 1.2.

(a) Yes: whether it has happened is known at each instant. (b) No: the low is known only at the close. (c) Yes: a deterministic time is a stopping time. (d) Yes, provided the forecast is itself adapted to the observed information.

Exercise 1.3 ★

A position on a fair ±1\pm1-tick walk is closed at +1+1 or −3-3. What is the probability of the take-profit, and why is the expected profit still zero?

Solution

Solution of Exercise 1.3.

b/(a+b)=3/4b/(a+b) = 3/4. The stopped walk is bounded, so optional stopping gives E[Sτ]=0\E[S_\tau] = 0: 34×1−14×3=0\tfrac34 \times 1 - \tfrac14 \times 3 = 0. Winning often is paid for by losing more when it loses.

Exercise 1.4 ★★

A forecast of the close is revised by independent news of equal variance in each of the 390 minutes of a session, and nothing else. What share of the close’s variance is resolved in the last half hour?

Solution

Solution of Exercise 1.4.

The revisions are orthogonal with equal variances, so the last 30 minutes resolve 30/390=7.7%30/390 = 7.7\% of the variance.

Exercise 1.5 ★★

Under P\P, X∼N(0,1)X \sim \mathcal N(0,1), and dQ/dP=exp⁡(cX−c2/2)d\mathbb Q/d\P = \exp(cX - c^2/2). Compute EQ[X]\E^{\mathbb Q}[X] and Q(X>c)\mathbb Q(X > c), and check that Q\mathbb Q is a probability measure.

Solution

Solution of Exercise 1.5.

E[Z]=e−c2/2E[ecX]=1\E[Z] = e^{-c^2/2}\E[e^{cX}] = 1 and Z>0Z > 0, so Q\mathbb Q is a probability measure equivalent to P\P. Its characteristic function is eiuc−u2/2e^{\iu uc - u^2/2} (Example 1.16): X∼N(c,1)X \sim \mathcal N(c,1) under Q\mathbb Q, so EQ[X]=c\E^{\mathbb Q}[X] = c and Q(X>c)=12\mathbb Q(X > c) = \tfrac12.

Exercise 1.6 ★★

The under-reacting card forecaster publishes Fk=35+a(Mk−35)F_k = 35 + a(M_k - 35) for k≤4k \le 4 and F5=M5F_5 = M_5. Compute E[F5−F4∣F4]\E[F_5 - F_4 \mid \mathcal F_4] and the slope of the regression of F5−F4F_5 - F_4 on F4−35F_4 - 35.

Solution

Solution of Exercise 1.6.

E4[F5]=E4[M5]=M4\E_4[F_5] = \E_4[M_5] = M_4, so E[F5−F4∣F4]=(1−a)(M4−35)=1−aa(F4−35)\E[F_5 - F_4 \mid \mathcal F_4] = (1 - a)(M_4 - 35) = \frac{1 - a}{a}(F_4 - 35): the revision is predictable, and the slope is (1−a)/a=0.667(1 - a)/a = 0.667 at a=0.6a = 0.6. The same slope appears on the previous revision F4−F3=a(M4−M3)F_4 - F_3 = a(M_4 - M_3), because M4−35M_4 - 35 is a sum of orthogonal revisions of which M4−M3M_4 - M_3 is one.

Exercise 1.7 ★★★

Coding. With revision_regression, estimate the slope of the last revision of the under-reacting card forecaster (a=0.6a = 0.6) on the previous revision over 20 000 seeded deals, and compare it with the answer to Exercise 1.6.

Solution

Solution of Exercise 1.7.

0.661 with a tt-statistic of 45.9, against 0.667 in theory; the honest forecaster’s slope is −0.005-0.005.

Exercise 1.8 ★★★

Find the flaw. “Over a year our fair-value model’s revisions averaged zero to four decimals, so the model is a martingale and nobody can trade against it.” Correct it.

Solution

Solution of Exercise 1.8.

A mean of zero is necessary, not sufficient: the under-reacting card forecaster also has revisions averaging zero. A martingale’s revisions must be unpredictable from everything known before them, starting with its own past revisions, and the test is a regression of revisions on earlier information, with standard errors that allow for the clustering of volatility (chapter 11) and a correction for the many signals one tries (chapter 12).

1.9 Problem: The Fair Value of the Close

Problem 1.1

Weekend problem — how much of the close is still unknown at ten to four

A desk publishes Mt=Et[C]M_t = \E_t[C], its forecast of a stock’s closing price CC. In its model the price starts at USD 100 and moves in each of the 390 minutes of the session by independent Gaussian increments with a total standard deviation of USD 1.80 over the day; at 15:50 the imbalance of the closing auction is published, which moves the forecast by an independent Gaussian amount JJ with standard deviation USD 0.30 and zero mean; the last ten minutes of the session follow.

Part I — The forecast.

  1. What are the variance and standard deviation of CC?
  2. Write MtM_t before 15:50 and after the news, in terms of the observed price.
  3. Why is (Mt)(M_t) a martingale for the desk’s filtration?
  4. What is the revision of MM at the 15:50 news?
  5. What is the probability that the news moves the forecast by more than 14 cents?

Part II — What is resolved when.

  1. Why do the variances of the revisions add up to Var⁡(C)\Var(C)?
  2. What share of Var⁡(C)\Var(C) is resolved by noon?
  3. What share is resolved at the news itself?
  4. What share is resolved from the news to the close, inclusive?
  5. Check this share against the simulation of 20 000 days.

Part III — A forecast that is too slow.

  1. A second desk adds only half of JJ at 15:50 and the other half at the close. Is its forecast a martingale?
  2. What is the correlation between its 15:50 revision and its revision over the last ten minutes?
  3. What slope does the regression of the second revision on the first give, in theory and in the simulation?
  4. How many days of data does the test need for the slope’s tt-statistic to reach 2?
  5. A trader buys one share from that desk at its forecast at 15:50 when its revision is positive, sells one when it is negative, and closes at the close. What does it earn on average per share?

Part IV — Judgement.

  1. If the last ten minutes carry three times the per-minute variance of the rest of the day (with the day’s total unchanged), what share is resolved from the news to the close?
  2. Does the answer depend on the price level?
  3. Why does the share matter to a trader who can only act in the closing auction?
  4. State the named result: the share of the close’s variance resolved from 15:50.
  5. In one sentence: what does the martingale property of a forecast promise, and what does it not?
Solution

Solution of Problem 1.1.

1. Var⁡(C)=1.802+0.302=3.33\Var(C) = 1.80^2 + 0.30^2 = 3.33; standard deviation USD 1.82. 2. Before the news, Mt=StM_t = S_t, the price observed at tt (the remaining increments and JJ have mean zero); after it, Mt=St+JM_t = S_t + J. 3. It is the Doob martingale Et[C]\E_t[C] of an integrable variable. 4. JJ. 5. P(∣J∣>0.14)=2(1−Φ(0.14/0.30))=64.1%\P(|J| > 0.14) = 2(1 - \Phi(0.14/0.30)) = 64.1\%. 6. The revisions are orthogonal (Proposition 1.8), so their variances add, and M0M_0 is a constant. 7. 150/390×3.24/3.33=37.4%150/390 \times 3.24/3.33 = 37.4\%. 8. 0.09/3.33=2.7%0.09/3.33 = 2.7\%. 9. (0.09+10/390×3.24)/3.33=5.2%(0.09 + 10/390 \times 3.24)/3.33 = 5.2\%: 2.7% at the news and 2.5% in the last ten minutes. 10. 5.23% in the 20 000 simulated days. 11. No: its revision over the last ten minutes contains J/2J/2, which its 15:50 revision J/2J/2 predicts. 12. The two revisions are J/2J/2 and J/2+J/2 + ten minutes of increments; the correlation is 0.152/(0.150.152+0.0831)=0.460.15^2/(0.15\sqrt{0.15^2 + 0.0831}) = 0.46. 13. Slope Cov⁡/Var⁡=0.152/0.152=1\Cov/\Var = 0.15^2/0.15^2 = 1; the simulation gives 0.98 with a tt-statistic of 72 over 20 000 days. 14. With t=rn−1/1−r2t = r\sqrt{n - 1}/\sqrt{1 - r^2}, t=2t = 2 needs n=4(1−r2)/r2+1=16n = 4(1 - r^2)/r^2 + 1 = 16 days. 15. The trade earns C−(S+J/2)C - (S + J/2) times the sign of JJ, whose mean is E∣J∣/2=0.5×0.302/π=0.120\E|J|/2 = 0.5 \times 0.30\sqrt{2/\pi} = 0.120: 12 cents a share. 16. The last ten minutes then carry 3.24×30/410=0.2373.24 \times 30/410 = 0.237, and the share is (0.09+0.237)/3.33=9.8%(0.09 + 0.237)/3.33 = 9.8\%. 17. No, if both standard deviations scale with the price: the shares are ratios of variances. 18. What is resolved after the last moment one can act is risk one carries, not information one can use; a trader deciding at 15:50 faces 5.2% of the day’s variance, not 100%. 19. Named result: the fair value of the close: in the model, 5.2% of the variance of the closing price is resolved from the 15:50 news to the close (2.7% by the news itself), and 9.8% if the last ten minutes are three times as volatile. 20. It promises that the revisions cannot be predicted from the forecaster’s own information, so nobody with that information can trade profitably against it; it promises nothing about accuracy.

1.10 Interview questions

Interview question 1.1 ★ trader

You play a fair coin game for a dollar a toss and stop as soon as you are one dollar up. Isn’t that a guaranteed win?

Solution

Solution of Interview question 1.1.

With unlimited time and credit you stop at +1+1 almost surely, but the stopping time is unbounded and the losses before it are not uniformly integrable; with any finite credit bb you win +1+1 with probability b/(b+1)b/(b+1) and lose bb otherwise: expected value zero.

What the interviewer is looking for: optional stopping and why its conditions matter.

Interview question 1.2 ★ researcher

What is a martingale? Is a stock price one?

Solution

Solution of Interview question 1.2.

An adapted integrable process whose conditional expectation of any future value is its current value. A stock price has a positive expected return under the real-world measure, so it is a submartingale there; discounted by the money-market account it is a martingale under the risk-neutral measure (chapter 5).

What the interviewer is looking for: the definition and the role of the measure.

Interview question 1.3 ★★ researcher, mle

Show that the increments of a martingale are uncorrelated. What does that imply for a forecast you publish?

Solution

Solution of Interview question 1.3.

For j<kj < k, E[DjDk]=E[DjEk−1[Dk]]=0\E[D_jD_k] = \E[D_j\E_{k-1}[D_k]] = 0. A forecast that is an honest conditional expectation has revisions no one can predict from its past revisions; if a regression of today’s revision on yesterday’s has a slope, the forecast is inefficient.

What the interviewer is looking for: taking out what is known, and the testable consequence.

Interview question 1.4 ★★ trader

At the imbalance publication your forecast of the close moves 14 cents; a competitor’s moves 7 cents, then another 7 at the close, every day. What do you do?

Solution

Solution of Interview question 1.4.

The competitor’s second revision is predictable from its first, so its forecast is not a martingale: it under-reacts. Trade against its stale price at 15:50 in the direction of the news; in the chapter’s model this earns half the mean absolute news, 12 cents a share.

What the interviewer is looking for: predictable revisions mean an exploitable forecast.

Interview question 1.5 ★★ researcher, risk

When is E[X∣Y]\E[X \mid Y] the linear regression of XX on YY?

Solution

Solution of Interview question 1.5.

E[X∣Y]\E[X \mid Y] is the best predictor among all functions of YY; the linear regression is the best among affine ones. They coincide when E[X∣Y]\E[X \mid Y] is affine, as it is when (X,Y)(X, Y) is jointly Gaussian.

What the interviewer is looking for: projections onto nested subspaces; the Gaussian case.

Interview question 1.6 ★★★ developer, researcher

Design a daily check that a real-time fair-value service behaves like a martingale. What can make the check lie?

Solution

Solution of Interview question 1.6.

Store every published value with its timestamp; each day regress revisions over fixed horizons on the previous revision and on other information available at the earlier time (order-book imbalance, recent trades), with robust standard errors, and report the variance shares by time of day. It lies if horizons overlap (serially correlated errors), if the volatility is heteroskedastic and the standard errors ignore it, if the timestamps are those of publication rather than of the information (look-ahead), if many signals are tested and the best one is reported, and if the final value is itself revised.

What the interviewer is looking for: regression of revisions on prior information, and the ways the test fails.

Terms defined in this chapter

See all 2333 terms in the glossary