Mathematics · Book 5 · Bachelor Year 3

University Mathematics — Year 3

University Mathematics — Year 3 · Bachelor Year 3

23Characteristic Functions and the Central Limit Theorem

The law of large numbers says averages converge; the central limit theorem says how they fluctuate: the error, magnified by n\sqrt n, is asymptotically Gaussian — whatever the law one started from. This universality is the deepest fact of elementary probability, and its natural proof is Fourier-analytic: the characteristic function (the Fourier transform of a law) converts independent sums into products, and Chapter 14’s machinery — injectivity, Gaussian fixed points — converts pointwise convergence of these products into convergence of laws (Lévy’s theorem, proved in full). The chapter ends with Gaussian vectors and the honest derivation of the confidence intervals used everywhere in statistics; the weekend problem gives Lindeberg’s second proof of the CLT, with an explicit error rate.

23.1 Characteristic functions

Definition 23.1

The characteristic function of a real random variable XX is

φX(ξ)=E[eiξX]=Reiξx ⁣dPX(x)(ξR)\varphi_X(\xi) = \E\bigl[\eu^{\iu\xi X}\bigr] = \int_\R \eu^{\iu\xi x}\,\dd\P_X(x) \qquad (\xi \in \R)

(the transfer theorem computes it from the law; for a density ff, φX(ξ)=f^(ξ)\varphi_X(\xi) = \hat f(-\xi) in Chapter 14’s convention).

Proposition 23.2

(a) φX(0)=1\varphi_X(0) = 1, φX1\abs{\varphi_X} \leq 1, and φX\varphi_X is uniformly continuous; φaX+b(ξ)=eibξφX(aξ)\varphi_{aX + b}(\xi) = \eu^{\iu b\xi}\varphi_X(a\xi). (b) If X,YX, Y are independent: φX+Y=φXφY\varphi_{X+Y} = \varphi_X\,\varphi_Y. (c) If EXk<\E\abs X^k < \infty, then φXCk\varphi_X \in \mathcal C^k with φX(j)(0)=ijE[Xj]\varphi_X^{(j)}(0) = \iu^j\,\E[X^j] for jkj \leq k; in particular, for centered XL2X \in L^2 with variance σ2\sigma^2:

φX(ξ)=1σ2ξ22+o(ξ2)(ξ0).\varphi_X(\xi) = 1 - \frac{\sigma^2\xi^2}{2} + o(\xi^2) \qquad (\xi \to 0).

(d) Gaussian: XN(m,σ2)X \sim \mathcal N(m, \sigma^2) has φX(ξ)=eimξσ2ξ2/2\varphi_X(\xi) = \eu^{\iu m\xi - \sigma^2\xi^2/2}.

Proof. (a) Bounds are immediate; continuity: φ(ξ+h)φ(ξ)EeihX10\abs{\varphi(\xi + h) - \varphi(\xi)} \leq \E\abs{\eu^{\iu hX} - 1} \to 0 as h0h \to 0 by dominated convergence, uniformly in ξ\xi. The affine rule is a substitution. (b) eiξ(X+Y)=eiξXeiξY\eu^{\iu\xi(X+Y)} = \eu^{\iu\xi X}\eu^{\iu\xi Y}, and expectations of products of independent variables factor (Theorem 22.5, applied to real and imaginary parts). (c) Differentiation under the expectation, dominated by EXj\E\abs X^j (Theorem 10.15); the Taylor expansion at 00 is then Taylor–Young for the C2\mathcal C^2 function φ\varphi. (d) For N(0,1)\mathcal N(0,1): the Gaussian transform (Example 14.2 with a=12a = \frac12) gives eiξxex2/22π ⁣dx=eξ2/2\int\eu^{\iu\xi x}\frac{\eu^{-x^2/2}}{\sqrt{2\pi}}\dd x = \eu^{-\xi^2/2}; the general case by the affine rule.

Theorem 23.3 (Injectivity)

If φX=φY\varphi_X = \varphi_Y, then XX and YY have the same law. More precisely, for NN(0,1)N \sim \mathcal N(0,1) independent of XX and ε>0\varepsilon > 0, the smoothed variable X+εNX + \varepsilon N has the density

pε(x)=12πRφX(ξ)eε2ξ2/2eiξx ⁣dξ,p_\varepsilon(x) = \frac1{2\pi}\int_\R \varphi_X(-\xi)\,\eu^{-\varepsilon^2\xi^2/2}\, \eu^{\iu\xi x}\,\dd\xi ,

determined by φX\varphi_X alone; letting ε0\varepsilon \to 0 recovers the law of XX.

Proof. X+εNX + \varepsilon N has the density pε(x)=E[gε(xX)]p_\varepsilon(x) = \E\bigl[g_\varepsilon(x - X)\bigr], where gεg_\varepsilon is the N(0,ε2)\mathcal N(0, \varepsilon^2) density: indeed for Borel BB, independence and Tonelli give P(X+εNB)= ⁣ ⁣1B(x+εn)g1(n) ⁣dn ⁣dPX(x)=BE[gε(tX)] ⁣dt\P(X + \varepsilon N \in B) = \int\!\!\int\mathbf 1_B(x + \varepsilon n)g_1(n)\,\dd n\,\dd\P_X(x) = \int_B\E[g_\varepsilon(t - X)]\dd t (substitute, then Tonelli again). Writing gεg_\varepsilon by Fourier inversion of its transform (Exercise 14.4, rescaled): gε(u)=12πeε2ξ2/2eiξu ⁣dξg_\varepsilon(u) = \frac1{2\pi}\int \eu^{-\varepsilon^2\xi^2/2}\eu^{\iu\xi u}\dd\xi, and Fubini (everything dominated by the Gaussian factor):

pε(x)=12πRφX(ξ)eε2ξ2/2eiξx ⁣dξ,p_\varepsilon(x) = \frac1{2\pi}\int_\R \varphi_X(-\xi)\,\eu^{-\varepsilon^2\xi^2/2}\, \eu^{\iu\xi x}\,\dd\xi ,

a functional of φX\varphi_X alone. If φX=φY\varphi_X = \varphi_Y: X+εNX + \varepsilon N and Y+εNY + \varepsilon N have equal laws for every ε\varepsilon; for bounded continuous ff, Ef(X+εN)Ef(X)\E f(X + \varepsilon N) \to \E f(X) as ε0\varepsilon \to 0 (dominated convergence, X+εNXX + \varepsilon N \to X pointwise on the product space), so Ef(X)=Ef(Y)\E f(X) = \E f(Y) for all such ff — and this determines the law: for each tt, squeeze 1(,t]\mathbf 1_{\intoc{-\infty}t} between the bounded continuous ramps fk±f_k^\pm (equal to 11 on (,t1k]\intoc{-\infty}{t \mp \frac1k}, to 00 beyond t±1kt \pm \frac1k, affine between); passing to the limit in Efk(X)FX(t)Efk+(X)\E f_k^-(X) \leq F_X(t) \leq \E f_k^+(X) gives FX(t)=FY(t)F_X(t) = F_Y(t) at every tt where both are continuous, hence everywhere by right-continuity and density of common continuity points (both FF’s have countably many jumps); equal distribution functions force equal laws (Exercise 9.3, resting on Theorem 9.7).

23.2 Convergence in distribution

Definition 23.4

XnX_n converges in distribution (or in law) to XX, written XnXX_n \Rightarrow X, if

E[f(Xn)]E[f(X)]for every bounded continuous f ⁣:RR.\E\bigl[f(X_n)\bigr] \longrightarrow \E\bigl[f(X)\bigr] \qquad\text{for every bounded continuous } f\colon\R\to\R .

Equivalently (Exercise 23.4): FXn(t)FX(t)F_{X_n}(t) \to F_X(t) at every continuity point tt of FXF_X. The XnX_n need not live on a common probability space: only laws matter.

Theorem 23.5 (Helly’s selection theorem)

Every sequence (Fn)(F_n) of distribution functions has a subsequence converging pointwise, at every continuity point of the limit, to a nondecreasing right-continuous G ⁣:R[0,1]G \colon \R \to \intcc01 — possibly with G(+)G()<1G(+\infty) - G(-\infty) < 1 (mass may escape to infinity).

Proof. Diagonal extraction gives Fnk(q)(q)F_{n_k}(q) \to \ell(q) for every rational qq (values in the compact [0,1]\intcc01). Define G(t)=inf{(q):qQ,q>t}G(t) = \inf\{\ell(q) : q \in \Q, q > t\}: nondecreasing; right-continuous (an infimum over shrinking rational neighborhoods from the right). At a continuity point tt of GG: for rationals q1<t<q2q_1 < t < q_2,

(q1)lim infFnk(t)lim supFnk(t)(q2),\ell(q_1) \leq \liminf F_{n_k}(t) \leq \limsup F_{n_k}(t) \leq \ell(q_2),

by monotonicity of each FnkF_{n_k}. From the definition of GG as an infimum and monotonicity of \ell on the rationals: G(s)(q)G(q)G(s) \leq \ell(q) \leq G(q) whenever s<qs < q. Taking s<q1<ts < q_1 < t gives (q1)G(s)\ell(q_1) \geq G(s), and (q2)G(q2)\ell(q_2) \leq G(q_2); letting sts \uparrow t and q2tq_2 \downarrow t, continuity of GG at tt squeezes both the lim inf\liminf and the lim sup\limsup to G(t)G(t).

Lemma 23.6 (Tightness from the characteristic function)

For any random variable XX and u>0u > 0:

P(X2u)    1uuu(1ReφX(ξ)) ⁣dξ.\P\Bigl(\abs X \geq \frac2u\Bigr) \;\leq\; \frac1u\int_{-u}^{u}\bigl(1 - \operatorname{Re}\varphi_X(\xi)\bigr)\,\dd\xi .

Proof. By Tonelli–Fubini (integrand bounded, region finite in ξ\xi):

1uuu(1ReφX(ξ)) ⁣dξ=E[1uuu(1cos(ξX)) ⁣dξ]=2E[1sin(uX)uX]\frac1u\int_{-u}^u\bigl(1 - \operatorname{Re}\varphi_X(\xi)\bigr)\dd\xi = \E\Bigl[\frac1u\int_{-u}^u(1 - \cos(\xi X))\,\dd\xi\Bigr] = 2\,\E\Bigl[1 - \frac{\sin(uX)}{uX}\Bigr]

(interpret the bracket as its limit 00 at X=0X = 0). The integrand is nonnegative (sintt\abs{\sin t} \leq \abs t), and for uX2\abs{uX} \geq 2: 1sin(uX)uX11uX121 - \frac{\sin(uX)}{uX} \geq 1 - \frac1{\abs{uX}} \geq \frac12. Keeping only the event {uX2}\{\abs{uX} \geq 2\} inside the expectation therefore leaves at least 212P(X2u)2 \cdot \frac12\,\P(\abs X \geq \frac2u), which is the claim.

Theorem 23.7 (Lévy’s continuity theorem)

Let (Xn)(X_n) be random variables whose characteristic functions converge pointwise: φXn(ξ)φ(ξ)\varphi_{X_n}(\xi) \to \varphi(\xi) for every ξ\xi, where φ=φX\varphi = \varphi_X is the characteristic function of some random variable XX. Then XnXX_n \Rightarrow X.

Proof. Tightness. Fix ε>0\varepsilon > 0. Since φ\varphi is continuous at 00 with φ(0)=1\varphi(0) = 1, choose u>0u > 0 with 1uuu(1Reφ)<ε\frac1u\int_{-u}^u(1 - \operatorname{Re}\varphi) < \varepsilon; by dominated convergence (integrand bounded by 22 on the fixed [u,u][-u,u]), the same integral for φXn\varphi_{X_n} is <2ε< 2\varepsilon for nn large: Lemma 23.6 gives P(Xn2u)2ε\P(\abs{X_n} \geq \frac2u) \leq 2\varepsilon for large nn, and enlarging the constant handles the finitely many others: the laws are tight — no mass escapes.

Subsequences. Let (Fnk)(F_{n_k}) be any subsequence; by Helly (Theorem 23.5) extract FnkjGF_{n_{k_j}} \to G at continuity points. Tightness forces G()=0G(-\infty) = 0, G(+)=1G(+\infty) = 1 (G(2u)G(2u)12εG(\frac2u) - G(-\frac2u) \geq 1 - 2\varepsilon at continuity points): GG is a genuine distribution function, of some random variable YY. Then XnkjYX_{n_{k_j}} \Rightarrow Y (Exercise 23.4, distributional convergence from FF’s), so φXnkjφY\varphi_{X_{n_{k_j}}} \to \varphi_Y pointwise (xeiξxx \mapsto \eu^{\iu\xi x} is bounded continuous, real and imaginary parts separately); comparing with the hypothesis: φY=φ=φX\varphi_Y = \varphi = \varphi_X, and injectivity (Theorem 23.3) gives YXY \sim X, i.e. G=FXG = F_X.

Conclusion. Every subsequence of (Fn)(F_n) has a sub-subsequence converging to the same FXF_X (at its continuity points); hence Fn(t)FX(t)F_n(t) \to F_X(t) at every continuity point tt (a real sequence all of whose subsequences have subsubsequences with the same limit converges): XnXX_n \Rightarrow X.

23.3 The central limit theorem

Theorem 23.8 (Central limit theorem)

Let (Xn)(X_n) be i.i.d. with EX1=m\E X_1 = m and V(X1)=σ2(0,)\V(X_1) = \sigma^2 \in \intoo0\infty. Then

Snnmσn    N(0,1):P(aSnnmσnb)12πabex2/2 ⁣dx\frac{S_n - nm}{\sigma\sqrt n} \;\Longrightarrow\; \mathcal N(0, 1) : \qquad \P\Bigl(a \leq \frac{S_n - nm}{\sigma\sqrt n} \leq b\Bigr) \longrightarrow \frac{1}{\sqrt{2\pi}}\int_a^b\eu^{-x^2/2}\,\dd x

for all a<ba < b.

Proof. Center and normalize: Zi=XimσZ_i = \frac{X_i - m}{\sigma} (i.i.d., mean 00, variance 11) and Tn=1ninZiT_n = \frac1{\sqrt n}\sum_{i\leq n}Z_i. By independence and the affine rule (Proposition 23.2):

φTn(ξ)=φZ(ξn)n,φZ(η)=1η22+η2ρ(η),ρ(η)0.\varphi_{T_n}(\xi) = \varphi_{Z}\Bigl(\frac{\xi}{\sqrt n}\Bigr)^{n}, \qquad \varphi_Z(\eta) = 1 - \frac{\eta^2}2 + \eta^2\rho(\eta),\quad \rho(\eta)\to0 .

Fix ξ\xi and let an=φZ(ξ/n)a_n = \varphi_Z(\xi/\sqrt n), bn=1ξ22nb_n = 1 - \frac{\xi^2}{2n}: both have modulus 1\leq 1 for nn large (bn1\abs{b_n} \leq 1 once ξ24n\xi^2 \leq 4n; an1\abs{a_n} \leq 1 always). The elementary inequality anbnnab\abs{a^n - b^n} \leq n\abs{a - b} for a,b1\abs a, \abs b \leq 1 (telescoping anbn=ak(ab)bn1ka^n - b^n = \sum a^k(a - b)b^{n-1-k}) gives

φTn(ξ)(1ξ22n)nnφZ(ξn)1+ξ22n=ξ2ρ(ξn)0,\Bigl|\varphi_{T_n}(\xi) - \Bigl(1 - \frac{\xi^2}{2n}\Bigr)^{n}\Bigr| \leq n\,\Bigl|\varphi_Z\Bigl(\frac\xi{\sqrt n}\Bigr) - 1 + \frac{\xi^2}{2n}\Bigr| = \xi^2\,\Bigl|\rho\Bigl(\frac{\xi}{\sqrt n}\Bigr)\Bigr| \longrightarrow 0,

while (1ξ22n)neξ2/2\bigl(1 - \frac{\xi^2}{2n}\bigr)^n \to \eu^{-\xi^2/2} (real logarithm). So φTn(ξ)eξ2/2=φN(0,1)(ξ)\varphi_{T_n}(\xi) \to \eu^{-\xi^2/2} = \varphi_{\mathcal N(0,1)}(\xi) (Proposition 23.2(d)) for every ξ\xi: Lévy (Theorem 23.7) concludes TnN(0,1)T_n \Rightarrow \mathcal N(0,1). The interval probabilities follow since FNF_{\mathcal N} is continuous everywhere.

Example 23.9 (Confidence intervals, honestly derived)

Poll nn independent voters; p^n=Sn/n\hat p_n = S_n/n estimates the true pp, with σ2=p(1p)14\sigma^2 = p(1-p) \leq \frac14. The CLT gives, for large nn,

P(p^npz2n)    P(Snnpσnz)Φ(z)Φ(z),\P\Bigl(\abs{\hat p_n - p} \leq \frac{z}{2\sqrt n}\Bigr) \;\geq\; \P\Bigl(\Bigl|\frac{S_n - np}{\sigma\sqrt n}\Bigr| \leq z\Bigr) \longrightarrow \Phi(z) - \Phi(-z),

where Φ\Phi is the standard Gaussian distribution function. With z=1.96z = 1.96: asymptotic confidence 95%95\%, and margin 1.962n3%\frac{1.96}{2\sqrt n} \leq 3\% requires n(1.960.06)21068n \geq \bigl(\frac{1.96}{0.06}\bigr)^2 \approx 1068 — the number behind every “±3\pm3 points, 95%95\%” one reads; compare Chebyshev’s 55565556 (Exercise 22.7). The n\sqrt n is universal: to halve the error, quadruple the sample — the same law that fixes Monte Carlo’s cost (Exercise 23.7).

23.4 Gaussian vectors

Definition 23.10

A random vector X=(X1,,Xd)X = (X_1, \dots, X_d) is Gaussian if every linear combination t,X=tiXi\langle t, X\rangle = \sum t_iX_i is a (possibly degenerate) real Gaussian variable. Its law is determined by the mean vector m=(EXi)m = (\E X_i) and the covariance matrix Σ=(Cov(Xi,Xj))\Sigma = \bigl(\operatorname{Cov} (X_i, X_j)\bigr): indeed the characteristic function of the vector, φX(t)=Eeit,X\varphi_X(t) = \E\eu^{\iu\langle t, X\rangle}, is the value at 11 of the cf of t,X\langle t, X\rangle:

φX(t)=exp(it,m12tTΣt),\varphi_X(t) = \exp\Bigl(\iu\langle t, m\rangle - \tfrac12\,t^{\mathsf T}\Sigma\,t\Bigr),

and dd-dimensional characteristic functions are injective (same smoothing proof as Theorem 23.3, coordinatewise Gaussians).

Theorem 23.11

Let XX be a Gaussian vector.

  1. Every affine image AX+bAX + b is a Gaussian vector.
  2. The components XiX_i are independent if and only if Σ\Sigma is diagonal: for jointly Gaussian variables, uncorrelated == independent.
  3. If Σ\Sigma is invertible, XX has the density 1(2π)d/2detΣexp(12(xm)TΣ1(xm))\frac{1}{(2\pi)^{d/2}\sqrt{\det\Sigma}} \exp\bigl(-\frac12(x - m)^{\mathsf T}\Sigma^{-1}(x - m)\bigr).

Proof. (1) Linear combinations of components of AX+bAX + b are affine functions of linear combinations of XX: Gaussian (an affine image of a Gaussian variable is Gaussian). (2) If Σ\Sigma is diagonal, the characteristic function factorizes: φX(t)=iexp(itimi12Σiiti2)=φXi(ti)\varphi_X(t) = \prod_i\exp(\iu t_im_i - \frac12\Sigma_{ii}t_i^2) = \prod\varphi_{X_i}(t_i), which is the characteristic function of the product law (Theorem 22.5 read through dd-dimensional injectivity): the components are independent. The converse is the vanishing of covariances of independent L2L^2 variables. (3) Diagonalize Σ=PDPT\Sigma = P D P^{\mathsf T} (PP orthogonal, D>0D > 0 diagonal — Exercise 20.8); the vector Y=PT(Xm)Y = P^{\mathsf T}(X - m) is Gaussian with covariance DD: by (2) its components are independent N(0,di)\mathcal N(0, d_i), so YY has the product density; push forward by the volume-preserving x=m+PYx = m + PY (Theorem 11.10, detP=1\abs{\det P} = 1) and rewrite the exponent invariantly.

Theorem 23.12 (Multidimensional CLT)

Let (Xn)(X_n) be i.i.d. square-integrable random vectors of Rd\R^d with mean mm and covariance matrix Σ\Sigma. Then Snnmn\frac{S_n - nm}{\sqrt n} converges in distribution to the Gaussian vector N(0,Σ)\mathcal N(0, \Sigma).

Proof. Admitted at this level.

Remark 23.13

Almost everything is already in our hands. For each direction tRdt \in \R^d, the real variable t,Snnmn\langle t, \frac{S_n - nm}{\sqrt n}\rangle is a normalized sum of i.i.d. real variables of variance tTΣtt^{\mathsf T}\Sigma t, so Theorem 23.8’s computation gives pointwise convergence of the dd-dimensional characteristic functions to etTΣt/2\eu^{-t^{\mathsf T}\Sigma t/2}, the characteristic function of N(0,Σ)\mathcal N(0, \Sigma) (Definition 23.10). What we have not re-proved is Lévy’s continuity theorem in Rd\R^d: Helly’s selection and the tightness estimate generalize routinely (coordinatewise), and this Cramér–Wold reduction is carried out honestly in any graduate probability course; nothing beyond this chapter’s methods is needed.

Method 23.14

To identify a limit law: compute characteristic functions, take the pointwise limit, recognize it (Gaussian eσ2ξ2/2\eu^{-\sigma^2\xi^2/2}, Poisson eλ(eiξ1)\eu^{\lambda(\eu^{\iu\xi}-1)}, exponential λλiξ\frac{\lambda} {\lambda - \iu\xi}, …) and invoke Lévy. The three-step ritual (independence \to product; Taylor at 00 \to exponential limit; Lévy \to convergence in law) proves the CLT, the Poisson law of rare events (Exercise 23.5), and every classical limit theorem of this course. For a.s. statements, go back to Chapter 22’s toolkit: the two chapters answer different questions about the same SnS_n.

23.5 Exercises

Exercise 23.1

Compute the characteristic functions: uniform on [1,1]\intcc{-1}1; exponential E(λ)\mathcal E(\lambda); Poisson P(λ)\mathcal P(\lambda); binomial B(n,p)\mathcal B(n, p). Deduce via Theorem 23.3 that the sum of independent Poisson variables (λ,μ\lambda, \mu) is Poisson (λ+μ)(\lambda + \mu).

Solution

Solution of Exercise 23.1.

Uniform on [1,1]\intcc{-1}1: φ(ξ)=1211eiξx ⁣dx=sinξξ\varphi(\xi) = \frac12\int_{-1}^1\eu^{\iu\xi x}\dd x = \frac{\sin\xi}{\xi} (equal to 11 at ξ=0\xi = 0). Exponential E(λ)\mathcal E(\lambda): φ(ξ)=λ0e(iξλ)x ⁣dx=λλiξ\varphi(\xi) = \lambda\int_0^\infty\eu^{(\iu\xi - \lambda)x}\dd x = \frac{\lambda}{\lambda - \iu\xi} (the antiderivative vanishes at ++\infty since Re(iξλ)<0\operatorname{Re}(\iu\xi - \lambda) < 0). Poisson P(λ)\mathcal P(\lambda): by the transfer theorem for discrete laws,

φ(ξ)=k0eiξkeλλkk!=eλexp(λeiξ)=exp(λ(eiξ1)).\varphi(\xi) = \sum_{k\geq0}\eu^{\iu\xi k}\,\eu^{-\lambda}\frac{\lambda^k}{k!} = \eu^{-\lambda}\exp\bigl(\lambda\eu^{\iu\xi}\bigr) = \exp\bigl(\lambda(\eu^{\iu\xi} - 1)\bigr).

Binomial B(n,p)\mathcal B(n, p): a sum of nn independent Bernoulli variables, each with cf 1p+peiξ1 - p + p\eu^{\iu\xi}, so φ(ξ)=(1p+peiξ)n\varphi(\xi) = \bigl(1 - p + p\eu^{\iu\xi}\bigr)^n (Proposition 23.2(b)). Poisson additivity: if XP(λ)X \sim \mathcal P(\lambda), YP(μ)Y \sim \mathcal P(\mu) are independent,

φX+Y(ξ)=eλ(eiξ1)eμ(eiξ1)=e(λ+μ)(eiξ1),\varphi_{X+Y}(\xi) = \eu^{\lambda(\eu^{\iu\xi}-1)} \eu^{\mu(\eu^{\iu\xi}-1)} = \eu^{(\lambda+\mu)(\eu^{\iu\xi}-1)},

the cf of P(λ+μ)\mathcal P(\lambda + \mu); injectivity (Theorem 23.3) identifies the law.

Exercise 23.2 ★★

(a) Show that φX\varphi_X is real-valued if and only if XX and X-X have the same law (a symmetric variable). (b) Suppose φX(ξ0)=1\abs{\varphi_X(\xi_0)} = 1 for some ξ00\xi_0 \neq 0. Show that XX is almost surely supported on an arithmetic progression a+2πξ0Za + \frac{2\pi}{\xi_0}\Z (write φX(ξ0)=eiθ\varphi_X(\xi_0) = \eu^{\iu\theta} and compute E[1cos(ξ0Xθ)]\E[1 - \cos(\xi_0X - \theta)]). Deduce that if XX has a density, then φX(ξ)<1\abs{\varphi_X(\xi)} < 1 for all ξ0\xi \neq 0.

Solution

Solution of Exercise 23.2.

(a) φX(ξ)=EeiξX=φX(ξ)\overline{\varphi_X(\xi)} = \E\eu^{-\iu\xi X} = \varphi_{-X}(\xi). So φX\varphi_X is real if and only if φX=φX\varphi_X = \varphi_{-X}, if and only if (injectivity, Theorem 23.3) XX and X-X have the same law. (b) Write φX(ξ0)=eiθ\varphi_X(\xi_0) = \eu^{\iu\theta}. Then

E[1cos(ξ0Xθ)]=1Re(eiθφX(ξ0))=11=0.\E\bigl[1 - \cos(\xi_0X - \theta)\bigr] = 1 - \operatorname{Re}\bigl(\eu^{-\iu\theta} \varphi_X(\xi_0)\bigr) = 1 - 1 = 0 .

The integrand is nonnegative, so cos(ξ0Xθ)=1\cos(\xi_0X - \theta) = 1 almost surely (a nonnegative variable with zero expectation vanishes a.s.), i.e. ξ0Xθ2πZ\xi_0X - \theta \in 2\pi\Z a.s.: XX takes its values in the arithmetic progression θξ0+2πξ0Z\frac{\theta}{\xi_0} + \frac{2\pi}{\xi_0}\Z almost surely. If XX has a density, this countable set is Lebesgue-null, so it carries probability 00 — contradiction; therefore φX(ξ)<1\abs{\varphi_X(\xi)} < 1 for every ξ0\xi \neq 0.

Exercise 23.3 ★★

Let XN(m1,σ12)X \sim \mathcal N(m_1, \sigma_1^2) and YN(m2,σ22)Y \sim \mathcal N(m_2, \sigma_2^2) be independent. Show X+YN(m1+m2,σ12+σ22)X + Y \sim \mathcal N(m_1 + m_2, \sigma_1^2 + \sigma_2^2), and more generally that the Gaussian family is stable under independent sums and affine maps. Contrast: is the sum of two dependent Gaussians always Gaussian? (Exercise 23.9.)

Solution

Solution of Exercise 23.3.

By independence and Proposition 23.2:

φX+Y(ξ)=eim1ξσ12ξ2/2eim2ξσ22ξ2/2=ei(m1+m2)ξ(σ12+σ22)ξ2/2,\varphi_{X+Y}(\xi) = \eu^{\iu m_1\xi - \sigma_1^2\xi^2/2}\, \eu^{\iu m_2\xi - \sigma_2^2\xi^2/2} = \eu^{\iu(m_1+m_2)\xi - (\sigma_1^2+\sigma_2^2)\xi^2/2},

the cf of N(m1+m2,σ12+σ22)\mathcal N(m_1 + m_2, \sigma_1^2 + \sigma_2^2); injectivity concludes. Stability under affine maps is the affine rule (aX+bN(am1+b,a2σ12)aX + b \sim \mathcal N(am_1 + b, a^2\sigma_1^2), allowing the degenerate case a=0a = 0), and stability under independent sums follows by induction on the computation above. For dependent Gaussians the sum need not be Gaussian: in Exercise 23.9, XX and Y=εXY = \varepsilon X are each standard Gaussian but X+YX + Y vanishes with probability 12\frac12 without being a.s. zero, so it is not Gaussian.

Exercise 23.4 ★★

(a) Prove the equivalence in Definition 23.4: if Ef(Xn)Ef(X)\E f(X_n) \to \E f(X) for all bounded continuous ff, then FXn(t)FX(t)F_{X_n}(t) \to F_X(t) at continuity points (squeeze 1(,t]\mathbf 1_{\intoc{-\infty}t} between two continuous staircase-ramps); and conversely (approximate a bounded continuous ff by sums of ramp functions, or condition on a fine grid of continuity points) — the converse may be treated for ff uniformly continuous first, then in general. (b) Show that XncX_n \Rightarrow c (a constant) implies XncX_n \to c in probability.

Solution

Solution of Exercise 23.4.

(a) Direct implication. Let tt be a continuity point of FXF_X and δ>0\delta > 0. Take the continuous ramps ff^- (=1= 1 on (,tδ]\intoc{-\infty}{t-\delta}, 00 from tt on, affine between) and f+f^+ (=1= 1 on (,t]\intoc{-\infty}t, 00 from t+δt + \delta on, affine between); then f1(,t]f+f^- \leq \mathbf 1_{\intoc{-\infty}t} \leq f^+, so

Ef(Xn)FXn(t)Ef+(Xn),\E f^-(X_n) \leq F_{X_n}(t) \leq \E f^+(X_n),

and the outer terms converge to Ef±(X)\E f^\pm(X), themselves squeezed between FX(tδ)F_X(t - \delta) and FX(t+δ)F_X(t + \delta). Letting nn \to \infty then δ0\delta \to 0 and using continuity of FXF_X at tt: FXn(t)FX(t)F_{X_n}(t) \to F_X(t).

Converse. Let ff be bounded continuous, M=supfM = \sup\abs f, ε>0\varepsilon > 0. The continuity points of FXF_X are dense (FXF_X has at most countably many jumps), so choose continuity points a<ba < b with FX(a)<εF_X(a) < \varepsilon and 1FX(b)<ε1 - F_X(b) < \varepsilon. On the compact [a,b]\intcc ab the function ff is uniformly continuous: choose continuity points a=t0<t1<<tm=ba = t_0 < t_1 < \dots < t_m = b of FXF_X with oscillation of ff at most ε\varepsilon on each (tj1,tj]\intoc{t_{j-1}}{t_j}, and set g=jf(tj)1(tj1,tj]g = \sum_j f(t_j)\,\mathbf 1_{\intoc{t_{j-1}}{t_j}}. Then fgε\abs{f - g} \leq \varepsilon on (a,b]\intoc ab, gM\abs g \leq M, and for T=XnT = X_n or XX:

Ef(T)Eg(T)ε+2M(FT(a)+1FT(b)).\bigl|\E f(T) - \E g(T)\bigr| \leq \varepsilon + 2M\bigl(F_T(a) + 1 - F_T(b)\bigr).

Moreover Eg(Xn)=jf(tj)(FXn(tj)FXn(tj1))Eg(X)\E g(X_n) = \sum_j f(t_j)\bigl(F_{X_n}(t_j) - F_{X_n}(t_{j-1})\bigr) \to \E g(X) (a finite sum of converging terms, all tjt_j being continuity points), and FXn(a)FX(a)<εF_{X_n}(a) \to F_X(a) < \varepsilon, 1FXn(b)1FX(b)<ε1 - F_{X_n}(b) \to 1 - F_X(b) < \varepsilon. Assembling: lim supnEf(Xn)Ef(X)2ε+8Mε\limsup_n\abs{\E f(X_n) - \E f(X)} \leq 2\varepsilon + 8M\varepsilon; let ε0\varepsilon \to 0.

(b) The distribution function of the constant cc is 1[c,)\mathbf 1_{\intco c\infty}, continuous except at cc. For ε>0\varepsilon > 0, the points cεc - \varepsilon and c+ε2c + \frac\varepsilon2 are continuity points, so

P(Xnc>ε)FXn(cε)+1FXn(c+ε2)0+11=0.\P(\abs{X_n - c} > \varepsilon) \leq F_{X_n}(c - \varepsilon) + 1 - F_{X_n}\Bigl(c + \frac\varepsilon2\Bigr) \longrightarrow 0 + 1 - 1 = 0 .

Exercise 23.5 ★★

(Law of rare events) Let XnB(n,pn)X_n \sim \mathcal B(n, p_n) with npnλ>0np_n \to \lambda > 0. Show, via characteristic functions and Theorem 23.7, that XnP(λ)X_n \Rightarrow \mathcal P(\lambda). Numerical sanity check: compare P(X=0)\P(X = 0) for B(100,0.02)\mathcal B(100, 0.02) and P(2)\mathcal P(2).

Solution

Solution of Exercise 23.5.

Let zn=pn(eiξ1)z_n = p_n(\eu^{\iu\xi} - 1), so φXn(ξ)=(1+zn)n\varphi_{X_n}(\xi) = (1 + z_n)^n (Exercise 23.1) and zn2pn0\abs{z_n} \leq 2p_n \to 0 (note pn=npnn0p_n = \frac{np_n}n \to 0). Both 1+zn1 + z_n and ezn\eu^{z_n} have modulus at most 11: 1+zn=(1pn)+pneiξ1\abs{1 + z_n} = \abs{(1 - p_n) + p_n\eu^{\iu\xi}} \leq 1 by the triangle inequality, and ezn=epn(cosξ1)1\abs{\eu^{z_n}} = \eu^{p_n(\cos\xi - 1)} \leq 1. The telescoping inequality anbnnab\abs{a^n - b^n} \leq n\abs{a - b} (proof of Theorem 23.8) and the power series bound ez1zz2ez\abs{\eu^z - 1 - z} \leq \abs z^2\eu^{\abs z} give

(1+zn)nenznn1+zneznnzn2ezn4e2npn2=4e2(npn)pn0.\bigl|(1 + z_n)^n - \eu^{nz_n}\bigr| \leq n\bigl|1 + z_n - \eu^{z_n}\bigr| \leq n\,\abs{z_n}^2\,\eu^{\abs{z_n}} \leq 4\eu^2\,np_n^2 = 4\eu^2\,(np_n)\,p_n \longrightarrow 0 .

Since nzn=npn(eiξ1)λ(eiξ1)nz_n = np_n(\eu^{\iu\xi} - 1) \to \lambda(\eu^{\iu\xi} - 1), we conclude φXn(ξ)exp(λ(eiξ1))\varphi_{X_n}(\xi) \to \exp\bigl(\lambda(\eu^{\iu\xi} - 1)\bigr) for every ξ\xi: the cf of P(λ)\mathcal P(\lambda), and Lévy (Theorem 23.7) gives XnP(λ)X_n \Rightarrow \mathcal P(\lambda). Numerically: P(B(100,0.02)=0)=0.98100=e100ln0.98e2.0200.1326\P\bigl(\mathcal B(100, 0.02) = 0\bigr) = 0.98^{100} = \eu^{100\ln 0.98} \approx \eu^{-2.020} \approx 0.1326, while P(P(2)=0)=e20.1353\P\bigl(\mathcal P(2) = 0\bigr) = \eu^{-2} \approx 0.1353: two percent apart already at this coarse nn.

Exercise 23.6 ★★

(a) A fair die is rolled n=1000n = 1000 times; approximate the probability that the total exceeds 36003600 (mean 35003500, variance per roll 3512\frac{35}{12}). (b) For SB(100,12)S \sim \mathcal B(100, \frac12), approximate P(45S55)\P(45 \leq S \leq 55) by the CLT with the continuity correction (±12\pm\frac12), and comment on the correction’s effect.

Solution

Solution of Exercise 23.6.

(a) One roll has mean 72\frac72 and variance 3512\frac{35}{12}, so SS has mean 35003500, variance 35000122916.7\frac{35000}{12} \approx 2916.7 and standard deviation 54.0\approx 54.0. By the CLT,

P(S>3600)=P(S350054.0>1.85)1Φ(1.85)0.032:\P(S > 3600) = \P\Bigl(\frac{S - 3500}{54.0} > 1.85\Bigr) \approx 1 - \Phi(1.85) \approx 0.032 :

about a 3%3\% chance. (b) SB(100,12)S \sim \mathcal B(100, \frac12): mean 5050, standard deviation 55. With the continuity correction,

P(45S55)Φ(55.5505)Φ(44.5505)=2Φ(1.1)10.729,\P(45 \leq S \leq 55) \approx \Phi\Bigl(\frac{55.5 - 50}{5}\Bigr) - \Phi\Bigl(\frac{44.5 - 50}{5}\Bigr) = 2\Phi(1.1) - 1 \approx 0.729,

against the exact value 0.72870.7287; without the correction, 2Φ(1)10.6832\Phi(1) - 1 \approx 0.683, off by almost five points. The correction matters because SS is a lattice variable: the atom P(S=k)\P(S = k) is well approximated by the Gaussian mass of [k12,k+12]\intcc{k - \frac12}{k + \frac12}, and clipping the interval at the integers 4545 and 5555 discards half an atom at each end.

Exercise 23.7 ★★

(Monte Carlo error) In the setting of Problem 22.1, question 11, with gL2([0,1]d)g \in L^2(\intcc01^d), let σ2=V(g(U1))\sigma^2 = \V(g(U_1)) and I=gI = \int g. Show

n(1nkng(Uk)I)N(0,σ2),\sqrt n\,\Bigl(\frac1n\sum_{k\leq n}g(U_k) - I\Bigr) \Longrightarrow \mathcal N(0, \sigma^2),

and deduce the asymptotic 95%95\% error bar ±1.96σ/n\pm 1.96\,\sigma/\sqrt nindependent of the dimension dd. Compare with the deterministic midpoint rule in dimension dd (error n2/d\sim n^{-2/d} for C2\mathcal C^2 integrands): from which dimension on does random sampling win?

Solution

Solution of Exercise 23.7.

The variables g(Uk)g(U_k) are i.i.d. (measurable images of i.i.d. variables), square-integrable, with mean II (transfer theorem, Exercise 11.9) and variance σ2\sigma^2. If σ>0\sigma > 0, Theorem 23.8 applied to them is exactly the stated convergence

n(1nkng(Uk)I)=kn(g(Uk)I)nN(0,σ2)\sqrt n\,\Bigl(\frac1n\sum_{k\leq n}g(U_k) - I\Bigr) = \frac{\sum_{k\leq n}\bigl(g(U_k) - I\bigr)}{\sqrt n} \Longrightarrow \mathcal N(0, \sigma^2)

(if σ=0\sigma = 0, gg is a.s. constant and the left-hand side vanishes identically). Hence P(1ng(Uk)I1.96σ/n)0.95\P\bigl(\abs{\frac1n\sum g(U_k) - I} \leq 1.96\,\sigma/\sqrt n\bigr) \to 0.95: the error bar ±1.96σ/n\pm 1.96\,\sigma/\sqrt n sees the dimension dd only through the constant σ\sigma, never through the rate in nn. The midpoint rule with nn nodes in dimension dd has mesh n1/dn^{-1/d} and error of order n2/dn^{-2/d} for C2\mathcal C^2 integrands. Monte Carlo’s n1/2n^{-1/2} decays faster than n2/dn^{-2/d} exactly when 12>2d\frac12 > \frac2d, i.e. d>4d > 4: from dimension 55 on, random sampling asymptotically beats the grid — the curse of dimensionality spares probabilistic methods, which is why Monte Carlo rules high-dimensional integration.

Exercise 23.8 ★★★

(Slutsky) Suppose XnXX_n \Rightarrow X and YncY_n \to c in probability (cc constant). Show Xn+YnX+cX_n + Y_n \Rightarrow X + c and YnXncXY_nX_n \Rightarrow cX. (Work with characteristic functions and the bound Eeiξ(Xn+Yn)eiξcEeiξXnEeiξ(Ync)1\abs{\E\eu^{\iu\xi (X_n+Y_n)} - \eu^{\iu\xi c}\E\eu^{\iu\xi X_n}} \leq \E\abs{\eu^{\iu\xi(Y_n - c)} - 1}, split on Yncδ\abs{Y_n - c} \leq \delta.) Application: in Example 23.9, justify replacing the unknown σ=p(1p)\sigma = \sqrt{p(1-p)} by p^n(1p^n)\sqrt{\hat p_n(1 - \hat p_n)}.

Solution

Solution of Exercise 23.8.

Sum. For fixed ξ\xi:

Eeiξ(Xn+Yn)eiξcEeiξXn=E[eiξXn(eiξYneiξc)]Eeiξ(Ync)1.\bigl|\E\eu^{\iu\xi(X_n+Y_n)} - \eu^{\iu\xi c}\,\E\eu^{\iu\xi X_n}\bigr| = \bigl|\E\bigl[\eu^{\iu\xi X_n}\bigl(\eu^{\iu\xi Y_n} - \eu^{\iu\xi c}\bigr)\bigr]\bigr| \leq \E\bigl|\eu^{\iu\xi(Y_n - c)} - 1\bigr| .

Split on the event {Yncδ}\{\abs{Y_n - c} \leq \delta\}: there, eiξ(Ync)1ξδ\abs{\eu^{\iu\xi(Y_n-c)} - 1} \leq \abs\xi\,\delta (the chord is shorter than the arc); the complement contributes at most 2P(Ync>δ)02\,\P(\abs{Y_n - c} > \delta) \to 0. Hence the lim sup\limsup is ξδ\leq \abs\xi\,\delta for every δ>0\delta > 0: the difference tends to 00. Since EeiξXnφX(ξ)\E\eu^{\iu\xi X_n} \to \varphi_X(\xi), we get φXn+Yn(ξ)eiξcφX(ξ)=φX+c(ξ)\varphi_{X_n+Y_n}(\xi) \to \eu^{\iu\xi c}\varphi_X(\xi) = \varphi_{X+c}(\xi), and Lévy (Theorem 23.7) yields Xn+YnX+cX_n + Y_n \Rightarrow X + c.

Product. First, cXncXcX_n \Rightarrow cX: φcXn(ξ)=φXn(cξ)φX(cξ)=φcX(ξ)\varphi_{cX_n}(\xi) = \varphi_{X_n}(c\xi) \to \varphi_X(c\xi) = \varphi_{cX}(\xi). Next, (Ync)Xn0(Y_n - c)X_n \to 0 in probability: the laws of the XnX_n are tight (their cfs converge to a cf; see the tightness step of Theorem 23.7), so given ε>0\varepsilon > 0 pick MM with P(Xn>M)ε\P(\abs{X_n} > M) \leq \varepsilon for all nn; then

P((Ync)Xn>ε)P(Xn>M)+P(Ync>εM)ε+o(1).\P\bigl(\abs{(Y_n - c)X_n} > \varepsilon\bigr) \leq \P(\abs{X_n} > M) + \P\Bigl(\abs{Y_n - c} > \frac{\varepsilon}{M}\Bigr) \leq \varepsilon + o(1) .

Writing YnXn=cXn+(Ync)XnY_nX_n = cX_n + (Y_n - c)X_n and applying the sum part (whose proof only used Yn:=(Ync)Xn0Y_n' := (Y_n - c)X_n \to 0 in probability, with constant 00): YnXncXY_nX_n \Rightarrow cX.

Application. By the strong law of large numbers (Theorem 22.13), p^np\hat p_n \to p a.s., so by continuity σ^n=p^n(1p^n)σ=p(1p)>0\hat\sigma_n = \sqrt{\hat p_n(1 - \hat p_n)} \to \sigma = \sqrt{p(1 - p)} > 0 a.s., hence σσ^n1\frac{\sigma}{\hat\sigma_n} \to 1 in probability. Slutsky’s product rule upgrades SnnpσnN(0,1)\frac{S_n - np}{\sigma\sqrt n} \Rightarrow \mathcal N(0,1) to Snnpσ^nn=σσ^nSnnpσnN(0,1)\frac{S_n - np}{\hat\sigma_n\sqrt n} = \frac{\sigma}{\hat\sigma_n}\cdot \frac{S_n - np}{\sigma\sqrt n} \Rightarrow \mathcal N(0,1): the usable confidence interval p^n±1.96σ^n/n\hat p_n \pm 1.96\,\hat\sigma_n/\sqrt n, built from the data alone, keeps its asymptotic 95%95\% level.

Exercise 23.9 ★★★

Let XN(0,1)X \sim \mathcal N(0,1) and ε\varepsilon independent with P(ε=±1)=12\P(\varepsilon = \pm1) = \frac12; set Y=εXY = \varepsilon X. (a) Show YN(0,1)Y \sim \mathcal N(0,1) and Cov(X,Y)=0\operatorname{Cov}(X, Y) = 0. (b) Show XX and YY are not independent, and that (X,Y)(X, Y) is not a Gaussian vector (compute P(X+Y=0)\P(X + Y = 0)). (c) Moral: Theorem 23.11(2) requires joint Gaussianity — “uncorrelated Gaussians” alone proves nothing.

Solution

Solution of Exercise 23.9.

(a) Splitting the expectation over the two values of ε\varepsilon (independence): for Borel BB, P(YB)=12P(XB)+12P(XB)=P(XB)\P(Y \in B) = \frac12\P(X \in B) + \frac12\P(-X \in B) = \P(X \in B), since XX-X \sim X (N(0,1)\mathcal N(0,1) is symmetric): YN(0,1)Y \sim \mathcal N(0,1). And Cov(X,Y)=E[εX2]=E[ε]E[X2]=01=0\operatorname{Cov}(X, Y) = \E[\varepsilon X^2] = \E[\varepsilon]\,\E[X^2] = 0 \cdot 1 = 0. (b) Y=X\abs Y = \abs X, so P(X1, Y2)=0\P(\abs X \leq 1,\ \abs Y \geq 2) = 0 while P(X1)P(Y2)>0\P(\abs X \leq 1)\,\P(\abs Y \geq 2) > 0: not independent. If (X,Y)(X, Y) were a Gaussian vector, X+Y=(1+ε)XX + Y = (1 + \varepsilon)X would be a real Gaussian variable (Definition 23.10 with t=(1,1)t = (1,1)); but P(X+Y=0)=P(ε=1)=12\P(X + Y = 0) = \P(\varepsilon = -1) = \frac12, whereas a Gaussian variable has an atom only if it is a.s. constant — and X+YX + Y equals 2X02X \neq 0 a.s. on {ε=1}\{\varepsilon = 1\}. Contradiction: (X,Y)(X, Y) is not Gaussian. (c) Each marginal is Gaussian and the covariance vanishes, yet independence fails — because the pair is not jointly Gaussian. Theorem 23.11(2) cannot be weakened to “Gaussian marginals”.

Exercise 23.10 ★★

The Cauchy law has density 1π(1+x2)\frac1{\pi(1 + x^2)}. (a) Show its characteristic function is eξ\eu^{-\abs\xi} (Exercise 14.1 and inversion). (b) Show that if X1,,XnX_1, \dots, X_n are i.i.d. Cauchy, then Snn\frac{S_n}n is again Cauchy — the same law: the average never concentrates. (c) Reconcile with the laws of large numbers and the CLT: which hypotheses fail? (Compute EX1\E\abs{X_1}.)

Solution

Solution of Exercise 23.10.

(a) Exercise 14.1 computes e^(ξ)=21+ξ2\widehat{\eu^{-\abs\cdot}}(\xi) = \frac{2}{1 + \xi^2}; both sides being integrable, Fourier inversion (Theorem 14.5) turns this around:

Reiξx ⁣dxπ(1+x2)=eξ,\int_\R\eu^{\iu\xi x}\,\frac{\dd x}{\pi(1 + x^2)} = \eu^{-\abs\xi},

which is exactly φX(ξ)\varphi_X(\xi) for a Cauchy variable XX. (b) By independence, φSn(ξ)=(eξ)n=enξ\varphi_{S_n}(\xi) = \bigl(\eu^{-\abs\xi}\bigr)^n = \eu^{-n\abs\xi}, so φSn/n(ξ)=φSn(ξ/n)=eξ\varphi_{S_n/n}(\xi) = \varphi_{S_n}(\xi/n) = \eu^{-\abs\xi}: the empirical mean Snn\frac{S_n}n is again standard Cauchy for every nn (injectivity). The average never concentrates: its fluctuations at time 10610^6 are those of a single observation. (c) EX1=2π0x ⁣dx1+x2=+\E\abs{X_1} = \frac2\pi\int_0^\infty\frac{x\,\dd x}{1 + x^2} = +\infty: the Cauchy law is not integrable, so the strong law of large numbers (Theorem 22.13) does not apply, and the CLT (which needs a finite variance) even less. Here their conclusions genuinely fail, not merely their proofs. Consistency check: φ(ξ)=eξ\varphi(\xi) = \eu^{-\abs\xi} is not differentiable at 00, as Proposition 23.2(c) read contrapositively predicts for a non-integrable variable.

Exercise 23.11 ★★

(Stable laws in embryo) Let (Xn)(X_n) be i.i.d. standard Cauchy (Exercise 23.10). (a) Show that for any a,b>0a, b > 0, aX1+bX2aX_1 + bX_2 has the law of (a+b)X1(a + b)X_1: the Cauchy family is strictly stable of index 11. (b) Show that the Gaussian family is strictly stable of index 22: aX1+bX2a2+b2X1aX_1 + bX_2 \sim \sqrt{a^2 + b^2}\,X_1 for XiX_i i.i.d. N(0,1)\mathcal N(0,1). (c) Explain, via characteristic functions of the form ecξα\eu^{-c\abs\xi^\alpha}, why index-α\alpha stability forces the normalization n1/αn^{1/\alpha} for sums, and what this says about the basins of attraction of the CLT: which i.i.d. sums can converge, after affine normalization, to a Cauchy law rather than a Gaussian?

Solution

Solution of Exercise 23.11.

(a) φaX1+bX2(ξ)=eaξebξ=e(a+b)ξ=φ(a+b)X1(ξ)\varphi_{aX_1 + bX_2}(\xi) = \eu^{-a\abs\xi}\eu^{-b\abs\xi} = \eu^{-(a+b)\abs\xi} = \varphi_{(a+b)X_1}(\xi) (independence and Exercise 23.10); injectivity identifies the laws.

(b) φaX1+bX2(ξ)=ea2ξ2/2eb2ξ2/2=e(a2+b2)ξ2/2\varphi_{aX_1+bX_2}(\xi) = \eu^{-a^2\xi^2/2} \eu^{-b^2\xi^2/2} = \eu^{-(a^2+b^2)\xi^2/2}: the law of a2+b2X1\sqrt{a^2+b^2}\,X_1.

(c) If φX(ξ)=ecξα\varphi_X(\xi) = \eu^{-c\abs\xi^\alpha}, then Sn=X1++XnS_n = X_1 + \dots + X_n has φSn=ecnξα\varphi_{S_n} = \eu^{-cn\abs\xi^\alpha}, and Sn/n1/αS_n/n^{1/\alpha} has φ(ξ)=ecξα\varphi(\xi) = \eu^{-c\abs\xi^\alpha} again: exact self-reproduction under the n1/αn^{1/\alpha} scaling — n\sqrt n for the Gaussian (α=2\alpha = 2), nn itself for Cauchy (α=1\alpha = 1, Exercise 23.10(b)). A sum of i.i.d. variables can only converge (after affine normalization) to a law that is stable under such convolutions; the CLT says finite variance forces the Gaussian basin, and the Cauchy basin is reserved for laws with tails so heavy that EX2=\E X^2 = \infty and even EX=\E\abs X = \infty — e.g. sums of Cauchy variables themselves. Universality has several islands, indexed by the tail exponent α(0,2]\alpha \in \intoc02.

Exercise 23.12 ★★

(The empirical distribution function) Let (Xn)(X_n) be i.i.d. with distribution function FF, and Fn(t)=1n#{kn:Xkt}F_n(t) = \frac1n\#\{k \leq n : X_k \leq t\}. (a) Fix tt. Show that nFn(t)B(n,F(t))n F_n(t) \sim \mathcal B(n, F(t)), that Fn(t)F(t)F_n(t) \to F(t) a.s. (Theorem 22.13), and that

n(Fn(t)F(t))N(0, F(t)(1F(t))).\sqrt n\,\bigl(F_n(t) - F(t)\bigr) \Longrightarrow \mathcal N\bigl(0,\ F(t)(1 - F(t))\bigr) .

(b) At which tt is the asymptotic variance maximal? Interpret: the median is where an empirical distribution is hardest to pin down. (c) For FF continuous, show that the law of suptFn(t)F(t)\sup_t\abs{F_n(t) - F(t)} does not depend on FF (reduce to uniform variables via Exercise 22.1) — the distribution-free miracle behind the Kolmogorov–Smirnov test; no computation of that law is asked.

Solution

Solution of Exercise 23.12.

(a) The indicators 1Xkt\mathbf 1_{X_k \leq t} are i.i.d. Bernoulli of parameter p=F(t)p = F(t): their sum nFn(t)nF_n(t) is binomial B(n,p)\mathcal B(n, p); the strong law gives Fn(t)pF_n(t) \to p a.s., and the CLT (Theorem 23.8) applied to the same indicators (variance p(1p)p(1-p)) gives the stated Gaussian limit.

(b) p(1p)p(1 - p) is maximal at p=12p = \frac12, i.e. where F(t)=12F(t) = \frac12: at the median. Estimating tail probabilities is asymptotically easy (variance 0\to 0 as p0,1p \to 0, 1); the median region carries the largest statistical noise — the empirical curve wobbles most in its middle.

(c) For continuous FF, the variables Uk=F(Xk)U_k = F(X_k) are i.i.d. uniform on (0,1)\intoo01 (Exercise 22.1), and monotonicity of FF gives, writing GnG_n for the empirical distribution function of the UkU_k:

suptRFn(t)F(t)=supuimFGn(u)u=supu[0,1]Gn(u)u:\sup_{t\in\R}\,\abs{F_n(t) - F(t)} = \sup_{u \in \operatorname{im}F}\,\abs{G_n(u) - u} = \sup_{u\in\intcc01}\abs{G_n(u) - u} :

the first equality because {Xkt}={UkF(t)}\{X_k \leq t\} = \{U_k \leq F(t)\} up to null events (monotonicity; strict inequality can fail only on the flat parts of FF, where both sides are unchanged), and the second because a continuous FF, running from 00 to 11, attains every value of (0,1)\intoo01 (intermediate value theorem), and the endpoints add nothing (Gn(0)0=0G_n(0) - 0 = 0 and Gn(1)1=0G_n(1) - 1 = 0). The right side involves only uniforms: one law for all FF — so a single table of critical values (that of Kolmogorov’s distribution) tests any continuous model against data.

23.6 Problem: Lindeberg’s proof of the CLT, with a rate

Problem 23.1

Weekend problem — the replacement method

Lindeberg (1922) proved the central limit theorem by an idea of disarming simplicity: swap the summands one at a time for Gaussians and control each swap by a Taylor expansion. The method needs no Fourier analysis, produces an explicit error rate, and today powers universality proofs across probability theory. Let (Xi)(X_i) be i.i.d., centered, V(X1)=1\V(X_1) = 1, with β=EX13<\beta = \E\abs{X_1}^3 < \infty; let (Ni)(N_i) be i.i.d. N(0,1)\mathcal N(0,1), independent of the XiX_i (existence: Theorem 22.6). Set

Tn=X1++Xnn,Gn=N1++NnnN(0,1).T_n = \frac{X_1 + \dots + X_n}{\sqrt n}, \qquad G_n = \frac{N_1 + \dots + N_n}{\sqrt n} \sim \mathcal N(0,1).

Part I — The swapping identity. Fix fCb3(R)f \in \mathcal C^3_b(\R) (three bounded continuous derivatives; M3=supfM_3 = \sup\abs{f'''}). For 0in0 \leq i \leq n define the hybrid sums

Hi=X1++Xi+Ni+1++Nnn,H_i = \frac{X_1 + \dots + X_i + N_{i+1} + \dots + N_n}{\sqrt n},

so Hn=TnH_n = T_n and H0=GnH_0 = G_n.

  1. Write Hi=Wi+XinH_i = W_i + \frac{X_i}{\sqrt n} and Hi1=Wi+NinH_{i-1} = W_i + \frac{N_i}{\sqrt n} with Wi=1n(j<iXj+j>iNj)W_i = \frac{1}{\sqrt n}\bigl(\sum_{j<i}X_j + \sum_{j>i}N_j\bigr), and note that WiW_i is independent of the pair (Xi,Ni)(X_i, N_i). Justify.
  2. Taylor with integral or Lagrange remainder: for any real w,hw, h:

    f(w+h)f(w)f(w)h12f(w)h2M3h36.\Bigl|f(w + h) - f(w) - f'(w)h - \tfrac12f''(w)h^2\Bigr| \leq \frac{M_3\,\abs h^3}{6} .
  3. Apply question 2 twice (h=Xinh = \frac{X_i}{\sqrt n} and h=Ninh = \frac{N_i}{\sqrt n} at w=Wiw = W_i), take expectations, and use independence plus the matching of the first two moments of XiX_i and NiN_i to show

    Ef(Hi)Ef(Hi1)M36β+γn3/2,γ=EN13=22π.\bigl|\E f(H_i) - \E f(H_{i-1})\bigr| \leq \frac{M_3}{6}\cdot \frac{\beta + \gamma}{n^{3/2}}, \qquad \gamma = \E\abs{N_1}^3 = \frac{2\sqrt2}{\sqrt\pi} .
  4. Telescope over ii and conclude the Lindeberg bound:

    Ef(Tn)Ef(Gn)M3(β+γ)6n.\bigl|\E f(T_n) - \E f(G_n)\bigr| \leq \frac{M_3\,(\beta + \gamma)}{6\,\sqrt n} .

Part II — From smooth ff to the CLT.

  1. Show that Ef(Tn)Ef(N)\E f(T_n) \to \E f(N) for every fCb3f \in \mathcal C_b^3, and upgrade to all bounded continuous ff: given such ff and ε\varepsilon, construct fεCb3f_\varepsilon \in \mathcal C^3_b with ffεε\norm{f - f_\varepsilon}_\infty \leq \varepsilon on a large interval — e.g. convolve ff with a C\mathcal C^\infty bump (Theorem 12.9) — and handle the tails by tightness (V(Tn)=1\V(T_n) = 1 and Chebyshev). Conclude TnN(0,1)T_n \Rightarrow \mathcal N(0, 1): the central limit theorem, re-proved.
  2. Where did the proof use that the XiX_i are identically distributed? Show that it barely did: state and prove the version for independent, centered, non-identical XiX_i with iV(Xi)=sn2\sum_i\V(X_i) = s_n^2 and third moments, obtaining the error M36sn3i(EXi3+V(Xi)3/2γ)\frac{M_3}{6s_n^3}\sum_i\bigl(\E\abs{X_i}^3 + \V(X_i)^{3/2}\gamma\bigr) — Lindeberg’s true theorem in its Lyapunov form.

Part III — Quantitative dividends.

  1. (Distribution functions) Let tRt \in \R and approximate 1(,t]\mathbf 1_{\intoc{-\infty}t} above and below by Cb3\mathcal C^3_b ramps of width δ\delta (construct them, with M3=O(δ3)M_3 = O(\delta^{-3})). Combining with Part I, derive the two-term bound

    suptRP(Tnt)Φ(t)    C1(β+γ)δ3n+C2δ(every δ>0),\sup_{t\in\R}\,\bigl|\P(T_n \leq t) - \Phi(t)\bigr| \;\leq\; \frac{C_1(\beta + \gamma)}{\delta^3\sqrt n} + C_2\,\delta \qquad (\text{every } \delta > 0),

    with explicit constants (the C2δC_2\delta term uses that Φ\Phi has density bounded by 12π\frac1{\sqrt{2\pi}}), and optimize δn1/8\delta \sim n^{-1/8} to obtain a uniform rate of order n1/8n^{-1/8}. (The optimal n1/2n^{-1/2} — Berry–Esseen — needs finer tools; the point is an explicit rate from elementary swapping.)

  2. (De Moivre–Laplace, quantified) Specialize to Xi=2Bi1X_i = 2B_i - 1 (signs of fair coins): compare the conclusion with the local estimate of Problem 11.1, question 7 — what does each method give that the other does not?
  3. (Universality) Explain in a paragraph why the replacement method shows more than the CLT: any statistic of the form Ef(sum)\E f(\text{sum}) with ff smooth is insensitive, at order n1/2n^{-1/2}, to the entire law of the summands beyond its first two moments — the “invariance principle” that underlies modern universality results (random matrices, random polynomials), of which the CLT is the first instance.

Part IV — Smoothing, pushed: better rates. The loss from n1/2n^{-1/2} (smooth ff) to n1/8n^{-1/8} (distribution functions) came from charging ff''' in sup norm. The hybrids can repair part of it: they contain Gaussian summands, and Gaussians smooth.

  1. (A hidden Gaussian) For 1in11 \leq i \leq n - 1, h=Xinh = \frac{X_i}{\sqrt n} or Nin\frac{N_i}{\sqrt n}, and θ[0,1]\theta \in \intcc01, write Wi+θh=A+ZW_i + \theta h = A + Z with Z=Ni+1++NnnZ = \frac{N_{i+1} + \dots + N_n}{\sqrt n}. Show that ZN(0,nin)Z \sim \mathcal N\bigl(0, \frac{n-i}n\bigr) is independent of the pair (A,h)(A, h), and deduce, for every continuous gL1(R)g \in L^1(\R),

    E[h3g(Wi+θh)]    n2π(ni)  gL1  Eh3.\E\bigl[\abs h^3\,\abs{g(W_i + \theta h)}\bigr] \;\leq\; \sqrt{\frac{n}{2\pi(n - i)}}\; \norm{g}_{L^1}\;\E\abs h^3 .
  2. Combine question 10 with the integral form of the Taylor remainder,

    f(w+h)=f(w)+f(w)h+12f(w)h2+01(1θ)22f(w+θh)h3 ⁣dθ,f(w + h) = f(w) + f'(w)h + \tfrac12f''(w)h^2 + \int_0^1\frac{(1 - \theta)^2}2\,f'''(w + \theta h)\,h^3\,\dd\theta,

    to redo questions 3–4: for fCb3f \in \mathcal C^3_b with moreover fL1(R)f''' \in L^1(\R),

    Ef(Tn)Ef(Gn)β+γ32πfL1n+M3(β+γ)6n3/2\bigl|\E f(T_n) - \E f(G_n)\bigr| \leq \frac{\beta + \gamma}{3\sqrt{2\pi}}\cdot \frac{\norm{f'''}_{L^1}}{\sqrt n} + \frac{M_3(\beta + \gamma)}{6\,n^{3/2}}

    (question 10 handles the swaps in1i \leq n - 1 — use m=1n1m1/22n\sum_{m=1}^{n-1}m^{-1/2} \leq 2\sqrt n — and question 3’s crude bound handles the last one). Check that question 7’s ramps satisfy ψδL1=K1δ2\norm{\psi_\delta'''}_{L^1} = K_1\delta^{-2} while M3=Kδ3M_3 = K\delta^{-3}, feed them in, and optimize δ\delta: the uniform distribution-function rate improves to O(n1/6)O(n^{-1/6}).

  3. (Matching one more moment) Assume in addition EX13=0\E X_1^3 = 0 and β4=EX14<\beta_4 = \E X_1^4 < \infty. Compute EN13\E N_1^3 and EN14\E N_1^4, expand to fourth order, and prove along the same lines that the distribution-function rate becomes O(n1/4)O(n^{-1/4}) (now ψδ(4)L1=K2δ3\norm{\psi_\delta^{(4)}}_{L^1} = K_2\delta^{-3} and M4=Kδ4M_4 = K'\delta^{-4}; choose δ=n1/4\delta = n^{-1/4}).
  4. (The obstruction) Suppose the first kk moments of X1X_1 agree with the Gaussian ones (k=2k = 2 always; k=3k = 3 exactly when EX13=0\E X_1^3 = 0; k4k \geq 4 essentially never, as EN14=3\E N_1^4 = 3). Verify that the scheme of questions 10–12 delivers the distribution-function rate n(k1)/(2k+2)n^{-(k-1)/(2k+2)}, by balancing δkn(k1)/2\delta^{-k}n^{-(k-1)/2} against δ\delta, and observe that the exponent approaches the Berry–Esseen value 12\frac12 only as kk \to \infty. Explain in a few sentences why the swapping method saturates: each swap is charged in absolute value, whereas the Fourier route (Esseen’s smoothing inequality) exploits the oscillation of the characteristic-function difference and reaches Cβn1/2C\beta n^{-1/2} with three moments only.

Part V — Two dimensions: the multidimensional CLT, by swapping. Now let the XiX_i be i.i.d. centered random vectors of R2\R^2 with covariance matrix Σ\Sigma and β=EX13<\beta' = \E\norm{X_1}^3 < \infty (Euclidean norm).

  1. (Gaussian vectors, to order) Diagonalize Σ=PDPT\Sigma = PDP^{\mathsf T} (Exercise 20.8) and set C=PDPTC = P\sqrt DP^{\mathsf T}. For Z=(Z1,Z2)Z = (Z^1, Z^2) a pair of independent standard Gaussians (Theorem 22.6), show that N=CZN = CZ is a Gaussian vector (Definition 23.10) of mean 00, covariance Σ\Sigma, with γ=EN3<\gamma' = \E\norm N^3 < \infty; and that Gn=N1++NnnG_n = \frac{N_1 + \dots + N_n}{\sqrt n} has law N(0,Σ)\mathcal N(0, \Sigma) exactly for i.i.d. copies NiN_i.
  2. (Taylor in two variables) For f ⁣:R2Rf \colon \R^2 \to \R of class C3\mathcal C^3 with M3=maxα=3supαf<M_3 = \max_{\abs\alpha = 3}\sup\abs{\partial^\alpha f} < \infty, prove

    f(w+h)f(w)f(w),h12h,D2f(w)hM36(h1+h2)32M33h3\Bigl|f(w + h) - f(w) - \langle\nabla f(w), h\rangle - \tfrac12\langle h, D^2f(w)\,h\rangle \Bigr| \leq \frac{M_3}6\,\bigl(\abs{h_1} + \abs{h_2}\bigr)^3 \leq \frac{\sqrt2\,M_3}3\, \norm h^3

    (study tf(w+th)t \mapsto f(w + th) on [0,1]\intcc01).

  3. (The CLT in R2\R^2) Run the replacement scheme on the vector hybrids HiH_i: show that the first- and second-order terms cancel (means and covariances match), telescope, and upgrade as in question 5 (tightness from ETn2=trΣ\E\norm{T_n}^2 = \operatorname{tr}\Sigma; mollification now in R2\R^2, Theorem 12.9) to conclude: for every bounded continuous f ⁣:R2Rf \colon \R^2 \to \R,

    Ef(X1++Xnn)Ef(N),NN(0,Σ):\E\,f\Bigl(\frac{X_1 + \dots + X_n}{\sqrt n}\Bigr) \longrightarrow \E\,f(N), \qquad N \sim \mathcal N(0, \Sigma) :

    Theorem 23.12 in dimension 22, with a rate for smooth ff and no Fourier analysis.

  4. (Cramér–Wold, and a joint fluctuation) Deduce that t,SnnN(0,tTΣt)\langle t, \frac{S_n}{\sqrt n}\rangle \Rightarrow \mathcal N(0, t^{\mathsf T}\Sigma t) for every fixed tR2t \in \R^2. Application: for i.i.d. real (ξi)(\xi_i), centered, Eξ12=1\E\xi_1^2 = 1, Eξ16<\E\xi_1^6 < \infty (so that Part V applies to Vi=(ξi,ξi21)V_i = (\xi_i, \xi_i^2 - 1)), show

    1n(inξi, in(ξi21))N(0,(1Eξ13Eξ13Eξ141)):\frac1{\sqrt n}\Bigl(\sum_{i\leq n}\xi_i,\ \sum_{i\leq n}(\xi_i^2 - 1)\Bigr) \Longrightarrow \mathcal N\Bigl(0, \begin{pmatrix} 1 & \E\xi_1^3\\ \E\xi_1^3 & \E\xi_1^4 - 1\end{pmatrix}\Bigr) :

    empirical mean and empirical second moment fluctuate jointly Gaussianly — independently in the limit if and only if Eξ13=0\E\xi_1^3 = 0 (Theorem 23.11).

Part VI — The delta method.

  1. Let (θ^n)(\hat\theta_n) be random variables with n(θ^nθ)N(0,σ2)\sqrt n(\hat\theta_n - \theta) \Rightarrow \mathcal N(0, \sigma^2) for a real parameter θ\theta, and let gg be differentiable at θ\theta. Prove the delta method:

    n(g(θ^n)g(θ))N(0,g(θ)2σ2)\sqrt n\bigl(g(\hat\theta_n) - g(\theta)\bigr) \Longrightarrow \mathcal N\bigl(0, g'(\theta)^2\sigma^2\bigr)

    (write g(x)g(θ)=(g(θ)+η(x))(xθ)g(x) - g(\theta) = (g'(\theta) + \eta(x))(x - \theta) with η0\eta \to 0 at θ\theta; show θ^nθ\hat\theta_n \to \theta, then η(θ^n)0\eta(\hat\theta_n) \to 0, in probability; finish with Slutsky, Exercise 23.8, and Exercise 23.4(b)).

  2. Applications. (a) For i.i.d. real (ξi)(\xi_i) with mean μ\mu and variance σ2\sigma^2, and Xˉn=1ninξi\bar X_n = \frac1n\sum_{i\leq n}\xi_i: show n(Xˉn2μ2)N(0,4μ2σ2)\sqrt n(\bar X_n^2 - \mu^2) \Rightarrow \mathcal N(0, 4\mu^2\sigma^2) when μ0\mu \neq 0, and that for μ=0\mu = 0 the correct statement lives at another scale: nXˉn2σ2N2n\bar X_n^2 \Rightarrow \sigma^2N^2 with NN(0,1)N \sim \mathcal N(0,1) (identify the limit’s distribution function). (b) (Variance stabilization) For p^n\hat p_n the success frequency of a B(1,p)\mathcal B(1, p) sample, p(0,1)p \in \intoo01: show that g(p)=arcsinpg(p) = \arcsin\sqrt p satisfies

    n(g(p^n)g(p))N(0,14)\sqrt n\,\bigl(g(\hat p_n) - g(p)\bigr) \Longrightarrow \mathcal N\Bigl(0, \frac14\Bigr)

    whatever pp — an asymptotic error bar free of the unknown parameter; compare with Example 23.9.

Part VII — Poisson, by the same method: Le Cam’s theorem. Replacement knows a second universality class: sums of many independent rare events. For laws on N\N the right distance is total variation,

dTV(μ,ν)=supANμ(A)ν(A).d_{\mathrm{TV}}(\mu, \nu) = \sup_{A\subseteq\N}\, \abs{\mu(A) - \nu(A)} .
  1. Show that dTV(μ,ν)=12k0μ({k})ν({k})d_{\mathrm{TV}}(\mu, \nu) = \frac12\sum_{k\geq0}\abs{\mu(\{k\}) - \nu(\{k\})}, and prove the coupling bound: for any pair (X,Y)(X, Y) of random variables with laws μ\mu and ν\nu on the same space, dTV(μ,ν)P(XY)d_{\mathrm{TV}}(\mu, \nu) \leq \P(X \neq Y).
  2. Compute exactly, for p(0,1)p \in \intoo01:

    dTV(B(1,p),P(p))=p(1ep)p2.d_{\mathrm{TV}}\bigl(\mathcal B(1, p), \mathcal P(p)\bigr) = p\bigl(1 - \eu^{-p}\bigr) \leq p^2 .
  3. (Le Cam, by swapping) Let XiB(1,pi)X_i \sim \mathcal B(1, p_i) and YiP(pi)Y_i \sim \mathcal P(p_i), the 2n2n variables independent; S=X1++XnS = X_1 + \dots + X_n, and recall Y1++YnP(λ)Y_1 + \dots + Y_n \sim \mathcal P(\lambda) with λ=ipi\lambda = \sum_ip_i (Exercise 23.1). Swap one coordinate at a time in the integer hybrids Hi=Y1++Yi+Xi+1++XnH_i = Y_1 + \dots + Y_i + X_{i+1} + \dots + X_n: show, for every ANA \subseteq \N,

    P(Hi1A)P(HiA)dTV(B(1,pi),P(pi)),\abs{\P(H_{i-1} \in A) - \P(H_i \in A)} \leq d_{\mathrm{TV}}\bigl(\mathcal B(1, p_i), \mathcal P(p_i)\bigr),

    and conclude Le Cam’s inequality:

    dTV(law of S, P(λ))i=1npi2.d_{\mathrm{TV}}\bigl(\text{law of } S,\ \mathcal P(\lambda)\bigr) \leq \sum_{i=1}^np_i^2 .
  4. Dividends. (a) For pi=λnp_i = \frac\lambda n: the bound is λ2n\frac{\lambda^2}n — the law of rare events (Exercise 23.5) upgraded to an explicit rate, uniform over all events, and valid for unequal pip_i as well. (b) 500500 letters are delivered, each going astray independently with probability 1500\frac1{500}: bound the error of the Poisson model of parameter 11, and estimate the probability that no letter goes astray. (c) Close the problem: compare the two universality classes met here — Gaussian (many small spread-out contributions; two moments matched; Taylor) and Poisson (many rare contributions; one mean matched; an exact total-variation coupling) — and the single replacement method behind both.
  5. (Relative error and the log transform) Let (Xn)(X_n) be i.i.d., positive, mean μ>0\mu > 0, variance σ2\sigma^2, and Xˉn\bar X_n the empirical mean. Show by the delta method that

    n(lnXˉnlnμ)N(0, σ2μ2):\sqrt n\,\bigl(\ln\bar X_n - \ln\mu\bigr) \Longrightarrow \mathcal N\Bigl(0,\ \frac{\sigma^2}{\mu^2}\Bigr) :

    the asymptotic parameter of lnXˉn\ln\bar X_n is the coefficient of variation σ/μ\sigma/\mu — relative, scale-free error. Deduce a 95%95\% confidence interval for μ\mu of the multiplicative form Xˉne±1.96σ/(μn)\bar X_n\cdot\eu^{\pm1.96\,\sigma/(\mu\sqrt n)}, and explain when it is preferable to the additive one.

  6. (The third moment steers the error) For XX \sim Bernoulli(pp) centered, compute E[(Xp)3]=p(1p)(12p)\E\bigl[(X - p)^3\bigr] = p(1-p)(1-2p). Using Part IV’s analysis (the swap error is driven by third moments), explain why the normal approximation of B(n,p)\mathcal B(n, p) is asymmetric for p12p \neq \frac12 — overshooting on one side, undershooting on the other — and why p=12p = \frac12 enjoys the faster matched-moment rate. Verify the sign of the skew numerically on B(20,0.1)\mathcal B(20, 0.1) against N(2,1.8)\mathcal N(2, 1.8): compare P(S=0)=0.920\P(S = 0) = 0.9^{20} with the Gaussian mass of (,0.5)\intoo{-\infty}{0.5}.
Solution

Solution of Problem 23.1.

1. The family (X1,,Xn,N1,,Nn)(X_1, \dots, X_n, N_1, \dots, N_n) is independent: the two blocks are independent of each other by construction and each block is i.i.d. WiW_i is a measurable function of the variables (Xj)j<i(X_j)_{j<i} and (Nj)j>i(N_j)_{j>i} only, all distinct from XiX_i and NiN_i: by the coalition principle (Theorem 22.5), WiW_i is independent of the pair (Xi,Ni)(X_i, N_i). The decompositions Hi=Wi+XinH_i = W_i + \frac{X_i}{\sqrt n} and Hi1=Wi+NinH_{i-1} = W_i + \frac{N_i}{\sqrt n} are immediate from the definitions: passing from HiH_i to Hi1H_{i-1} swaps the single summand XiX_i for NiN_i.

2. Taylor–Lagrange at order 33: there is cc between ww and w+hw + h with f(w+h)=f(w)+f(w)h+12f(w)h2+16f(c)h3f(w + h) = f(w) + f'(w)h + \frac12f''(w)h^2 + \frac16f'''(c)h^3, and f(c)M3\abs{f'''(c)} \leq M_3 gives the bound.

3. Subtracting the two expansions at the common base point w=Wiw = W_i:

f(Hi)f(Hi1)=f(Wi)XiNin+f(Wi)2Xi2Ni2n+Ri,RiM36Xi3+Ni3n3/2.f(H_i) - f(H_{i-1}) = f'(W_i)\,\frac{X_i - N_i}{\sqrt n} + \frac{f''(W_i)}{2}\,\frac{X_i^2 - N_i^2}{n} + R_i, \qquad \abs{R_i} \leq \frac{M_3}{6}\cdot \frac{\abs{X_i}^3 + \abs{N_i}^3}{n^{3/2}} .

Take expectations. By question 1, f(Wi)f'(W_i) and f(Wi)f''(W_i) are independent of (Xi,Ni)(X_i, N_i), so the mixed expectations factor:

E[f(Wi)XiNin]=E[f(Wi)]EXiENin=0,E[f(Wi)Xi2Ni2n]=E[f(Wi)]11n=0:\begin{align*} \E\Bigl[f'(W_i)\,\frac{X_i - N_i}{\sqrt n}\Bigr] &= \E\bigl[f'(W_i)\bigr]\,\frac{\E X_i - \E N_i}{\sqrt n} = 0, \\ \E\Bigl[f''(W_i)\,\frac{X_i^2 - N_i^2}{n}\Bigr] &= \E\bigl[f''(W_i)\bigr]\,\frac{1 - 1}{n} = 0 : \end{align*}

the first two moments of XiX_i and NiN_i match, and only the remainder survives:

Ef(Hi)Ef(Hi1)ERiM36β+γn3/2.\bigl|\E f(H_i) - \E f(H_{i-1})\bigr| \leq \E\abs{R_i} \leq \frac{M_3}{6}\cdot\frac{\beta + \gamma}{n^{3/2}} .

The Gaussian third moment: γ=EN13=20x3ex2/22π ⁣dx=22π02ueu ⁣du=42π=22π\gamma = \E\abs{N_1}^3 = 2\int_0^\infty x^3\,\frac{\eu^{-x^2/2}}{\sqrt{2\pi}}\,\dd x = \frac{2}{\sqrt{2\pi}}\int_0^\infty 2u\,\eu^{-u}\dd u = \frac{4}{\sqrt{2\pi}} = \frac{2\sqrt2}{\sqrt\pi} (substitution u=x2/2u = x^2/2, then Γ(2)=1\Gamma(2) = 1).

4. Telescoping Ef(Tn)Ef(Gn)=i=1n(Ef(Hi)Ef(Hi1))\E f(T_n) - \E f(G_n) = \sum_{i=1}^n\bigl(\E f(H_i) - \E f(H_{i-1})\bigr) and applying question 3 to each of the nn terms:

Ef(Tn)Ef(Gn)nM3(β+γ)6n3/2=M3(β+γ)6n.\bigl|\E f(T_n) - \E f(G_n)\bigr| \leq n \cdot \frac{M_3(\beta + \gamma)}{6\,n^{3/2}} = \frac{M_3\,(\beta + \gamma)}{6\,\sqrt n} .

5. GnG_n is N(0,1)\mathcal N(0,1) exactly for every nn (a normalized sum of independent standard Gaussians, Exercise 23.3), so Ef(Gn)=Ef(N)\E f(G_n) = \E f(N) and question 4 reads Ef(Tn)Ef(N)M3(β+γ)6n0\abs{\E f(T_n) - \E f(N)} \leq \frac{M_3(\beta+\gamma)}{6\sqrt n} \to 0 for fCb3f \in \mathcal C^3_b. Upgrade. Let ff be bounded continuous, M=supfM = \sup\abs f, ε>0\varepsilon > 0. Choose A1A \geq 1 with 1A2ε\frac1{A^2} \leq \varepsilon: Chebyshev with V(Tn)=1\V(T_n) = 1 gives P(Tn>A)ε\P(\abs{T_n} > A) \leq \varepsilon for all nn, and likewise P(N>A)ε\P(\abs N > A) \leq \varepsilon. Let χ\chi be C\mathcal C^\infty with 1[A,A]χ1[A1,A+1]\mathbf 1_{\intcc{-A}A} \leq \chi \leq \mathbf 1_{\intcc{-A-1}{A+1}} (a smooth plateau, built by mollifying 1[A12,A+12]\mathbf 1_{\intcc{-A-\frac12}{A+\frac12}}, Theorem 12.9); g=fχg = f\chi is continuous with compact support, hence uniformly continuous, so its mollification gη=gρηg_\eta = g * \rho_\eta is C\mathcal C^\infty with bounded derivatives of all orders and ggηε\norm{g - g_\eta}_\infty \leq \varepsilon for η\eta small enough. For T=TnT = T_n or NN, since f=gf = g on [A,A]\intcc{-A}A and fg2M\abs{f - g} \leq 2M everywhere:

Ef(T)Egη(T)E(fg)(T)+ggη2MP(T>A)+ε(2M+1)ε.\bigl|\E f(T) - \E g_\eta(T)\bigr| \leq \E\abs{(f - g)(T)} + \norm{g - g_\eta}_\infty \leq 2M\,\P(\abs T > A) + \varepsilon \leq (2M + 1)\,\varepsilon .

Combining with Egη(Tn)Egη(N)\E g_\eta(T_n) \to \E g_\eta(N) (question 4 applies: gηCb3g_\eta \in \mathcal C^3_b):

lim supn  Ef(Tn)Ef(N)2(2M+1)ε,\limsup_n\;\bigl|\E f(T_n) - \E f(N)\bigr| \leq 2(2M + 1)\,\varepsilon ,

and ε\varepsilon was arbitrary: Ef(Tn)Ef(N)\E f(T_n) \to \E f(N) for every bounded continuous ff, i.e. TnN(0,1)T_n \Rightarrow \mathcal N(0,1).

6. Identical distribution entered only through one sentence: “XiX_i and NiN_i have the same first two moments”. So let X1,,XnX_1, \dots, X_n be independent, centered, with variances σi2\sigma_i^2 and finite third moments, sn2=iσi2>0s_n^2 = \sum_i\sigma_i^2 > 0, and take NiN(0,σi2)N_i \sim \mathcal N(0, \sigma_i^2) independent of everything. Define the hybrids with normalization sns_n: Hi=1sn(jiXj+j>iNj)H_i = \frac1{s_n}(\sum_{j\leq i}X_j + \sum_{j>i}N_j). In the ii-th swap, EXi=ENi=0\E X_i = \E N_i = 0 and EXi2=ENi2=σi2\E X_i^2 = \E N_i^2 = \sigma_i^2 again kill the ff' and ff'' terms, and the remainder gives (using ENi3=σi3γ\E\abs{N_i}^3 = \sigma_i^3\gamma by scaling):

Ef(Hi)Ef(Hi1)M36sn3(EXi3+σi3γ).\bigl|\E f(H_i) - \E f(H_{i-1})\bigr| \leq \frac{M_3}{6\,s_n^3}\bigl(\E\abs{X_i}^3 + \sigma_i^3\gamma\bigr) .

Telescoping:

Ef(X1++Xnsn)Ef(N)M36sn3i=1n(EXi3+V(Xi)3/2γ).\Bigl|\E f\Bigl(\frac{X_1 + \dots + X_n}{s_n}\Bigr) - \E f(N)\Bigr| \leq \frac{M_3}{6\,s_n^3}\sum_{i=1}^n \Bigl(\E\abs{X_i}^3 + \V(X_i)^{3/2}\,\gamma\Bigr) .

Since σi3=(EXi2)3/2EXi3\sigma_i^3 = (\E X_i^2)^{3/2} \leq \E\abs{X_i}^3 (the power–mean inequality, i.e. Jensen for tt3/2t \mapsto t^{3/2} applied to Xi2X_i^2), the right side is at most M3(1+γ)6iEXi3sn3\frac{M_3(1 + \gamma)}{6}\cdot \frac{\sum_i\E\abs{X_i}^3}{s_n^3}: under Lyapunov’s condition 1sn3iEXi30\frac1{s_n^3}\sum_i\E\abs{X_i}^3 \to 0, the normalized sums converge in law to N(0,1)\mathcal N(0,1) — the CLT without identical distribution.

7. Let ρCc((0,1))\rho \in \mathcal C^\infty_c(\intoo01) with ρ=1\int\rho = 1 and set ψ(x)=x1ρ(s) ⁣ds\psi(x) = \int_x^1\rho(s)\dd s: ψ\psi is C\mathcal C^\infty, nonincreasing, ψ=1\psi = 1 on R\R_-, ψ=0\psi = 0 on [1,)\intco1\infty; let K=ψK = \norm{\psi'''}_\infty. For tRt \in \R and δ>0\delta > 0 define ψδ(x)=ψ(xtδ)\psi_\delta(x) = \psi\bigl(\frac{x - t}\delta\bigr) and ψ~δ(x)=ψ(xtδ+1)\tilde\psi_\delta(x) = \psi\bigl(\frac{x - t}\delta + 1\bigr): these are Cb3\mathcal C^3_b with third derivative bounded by K/δ3K/\delta^3, and

1(,tδ]ψ~δ1(,t]ψδ1(,t+δ].\mathbf 1_{\intoc{-\infty}{t-\delta}} \leq \tilde\psi_\delta \leq \mathbf 1_{\intoc{-\infty}t} \leq \psi_\delta \leq \mathbf 1_{\intoc{-\infty}{t+\delta}} .

Upper bound: by question 4 applied to ψδ\psi_\delta (with M3=K/δ3M_3 = K/\delta^3),

P(Tnt)Eψδ(Tn)Eψδ(N)+K(β+γ)6δ3nΦ(t+δ)+K(β+γ)6δ3nΦ(t)+δ2π+K(β+γ)6δ3n,\P(T_n \leq t) \leq \E\psi_\delta(T_n) \leq \E\psi_\delta(N) + \frac{K(\beta + \gamma)}{6\,\delta^3\sqrt n} \leq \Phi(t + \delta) + \frac{K(\beta + \gamma)}{6\,\delta^3\sqrt n} \leq \Phi(t) + \frac{\delta}{\sqrt{2\pi}} + \frac{K(\beta + \gamma)}{6\,\delta^3\sqrt n},

because Φ\Phi is Lipschitz with constant 12π\frac1{\sqrt{2\pi}} (its density is bounded by 12π\frac1{\sqrt{2\pi}}). The symmetric lower bound via ψ~δ\tilde\psi_\delta gives the two-term estimate

suptRP(Tnt)Φ(t)K(β+γ)61δ3n+δ2π(δ>0 arbitrary).\sup_{t\in\R}\,\bigl|\P(T_n \leq t) - \Phi(t)\bigr| \leq \frac{K(\beta + \gamma)}{6}\cdot\frac{1}{\delta^3\sqrt n} + \frac{\delta}{\sqrt{2\pi}} \qquad(\delta > 0\ \text{arbitrary}).

The two terms balance when δ3n1/2δ\delta^{-3}n^{-1/2} \asymp \delta, i.e. δ=n1/8\delta = n^{-1/8}: both are then O(n1/8)O(n^{-1/8}), an explicit uniform rate valid for every nn. (The optimal Berry–Esseen rate Cβ/nC\beta/\sqrt n requires the Fourier smoothing inequality; swapping trades sharpness for complete elementarity.)

8. For Xi=2Bi1X_i = 2B_i - 1 (fair signs): centered, variance 11, and Xi=1\abs{X_i} = 1 so β=1\beta = 1. Question 7 then bounds suptP(Snnt)Φ(t)\sup_t\abs{\P(\frac{S_n}{\sqrt n} \leq t) - \Phi(t)} explicitly and uniformly for every finite nn — a global, non-asymptotic statement about the distribution function. The local estimate of Problem 11.1, question 7, gives instead the exact asymptotics of an individual atom, P(S2n=2k)ek2/nπn\P(S_{2n} = 2k) \sim \frac{\eu^{-k^2/n}}{\sqrt{\pi n}}: it resolves probabilities of size n1/2n^{-1/2}, far below question 7’s n1/8n^{-1/8} resolution, but it is pointwise, asymptotic (no explicit error at fixed nn) and tied to this particular lattice law. Local precision versus global uniformity: the two methods are complementary, and summing the local estimate over k[ ⁣[an,bn] ⁣]k \in \intint{a\sqrt n}{b\sqrt n} recovers de Moivre–Laplace on intervals — with a sharper rate, but only for this law.

9. The swapping argument used nothing about the law of the XiX_i beyond EXi=0\E X_i = 0, EXi2=1\E X_i^2 = 1 and the finiteness of EXi3\E\abs{X_i}^3: had we replaced the Gaussians NiN_i by any other i.i.d. family with the same first two moments and finite third moment, the same telescoping would bound Ef(sumX)Ef(sumY)\abs{\E f(\text{sum}_X) - \E f(\text{sum}_Y)} by O(n1/2)O(n^{-1/2}) for every smooth ff. Smooth statistics of large independent sums are therefore universal: up to a quantified error, they depend on the summands’ law only through two numbers. This is the invariance principle: prove a limit theorem for the most computable law (the Gaussian, where everything is exact), then transfer it to all laws by swapping. The same scheme — with sums replaced by more elaborate functionals — drives Wigner’s semicircle law for random matrices, the universality of roots of random polynomials, and much of modern probability; the central limit theorem is its first and simplest instance.

10. ZZ is a Borel function of (Ni+1,,Nn)(N_{i+1}, \dots, N_n) only, while A=Wi+θhZA = W_i + \theta h - Z and hh are functions of the remaining variables of the independent family (X1,,Xn,N1,,Nn)(X_1, \dots, X_n, N_1, \dots, N_n): by the coalition principle (Theorem 22.5), ZZ is independent of (A,h)(A, h). As a sum of the independent Nj/nN(0,1n)N_j/\sqrt n \sim \mathcal N(0, \frac1n), ZN(0,s2)Z \sim \mathcal N(0, s^2) with s2=nins^2 = \frac{n-i}n (Exercise 23.3), with density bounded by 1s2π\frac1{s\sqrt{2\pi}}. The law of ((A,h),Z)((A, h), Z) is the product of the two marginal laws, so Tonelli (transfer) freezes the first block: with G(a)=Eg(a+Z)=g(a+z)φs(z) ⁣dzgL1s2πG(a) = \E\abs{g(a + Z)} = \int\abs{g(a + z)}\,\varphi_s(z)\,\dd z \leq \frac{\norm g_{L^1}}{s\sqrt{2\pi}} for every aa,

E[h3g(A+Z)]=E[h3G(A)]gL1s2πEh3=n2π(ni)gL1Eh3.\E\bigl[\abs h^3\abs{g(A + Z)}\bigr] = \E\bigl[\abs h^3\,G(A)\bigr] \leq \frac{\norm g_{L^1}}{s\sqrt{2\pi}}\,\E\abs h^3 = \sqrt{\frac{n}{2\pi(n-i)}}\,\norm g_{L^1}\,\E\abs h^3 .

11. The integral form of Taylor’s formula follows by integrating f(w+h)f(w)=h01f(w+θh) ⁣dθf(w + h) - f(w) = h\int_0^1f'(w + \theta h)\,\dd\theta by parts twice in θ\theta. Taking expectations in the ii-th swap, the orders 0,1,20, 1, 2 cancel exactly as in question 3, and the two remainders (for h=Xi/nh = X_i/\sqrt n and Ni/nN_i/\sqrt n) are bounded, for in1i \leq n - 1, by question 10 with g=fg = f''':

Ef(Hi)Ef(Hi1)01(1θ)22 ⁣dθ  n2π(ni)fL1β+γn3/2=β+γ6n3/2n2π(ni)fL1.\bigl|\E f(H_i) - \E f(H_{i-1})\bigr| \leq \int_0^1\frac{(1-\theta)^2}2\,\dd\theta\; \sqrt{\frac{n}{2\pi(n-i)}}\,\norm{f'''}_{L^1} \frac{\beta + \gamma}{n^{3/2}} = \frac{\beta + \gamma}{6\,n^{3/2}} \sqrt{\frac{n}{2\pi(n-i)}}\,\norm{f'''}_{L^1} .

Summing, with i=1n1nni=nm=1n1m1/22n\sum_{i=1}^{n-1}\sqrt{\frac n{n-i}} = \sqrt n\sum_{m=1}^{n-1}m^{-1/2} \leq 2n, and adding question 3’s bound for the last swap (i=ni = n, no Gaussian left):

Ef(Tn)Ef(Gn)β+γ32πfL1n+M3(β+γ)6n3/2.\bigl|\E f(T_n) - \E f(G_n)\bigr| \leq \frac{\beta + \gamma}{3\sqrt{2\pi}}\cdot \frac{\norm{f'''}_{L^1}}{\sqrt n} + \frac{M_3(\beta + \gamma)}{6\,n^{3/2}} .

Ramps: ψδ(x)=δ3ψ(xtδ)\psi_\delta'''(x) = \delta^{-3}\psi'''\bigl(\frac{x - t}\delta\bigr), so M3=Kδ3M_3 = K\delta^{-3} with K=ψK = \norm{\psi'''}_\infty and ψδL1=δ2ψL1=K1δ2\norm{\psi_\delta'''}_{L^1} = \delta^{-2} \norm{\psi'''}_{L^1} = K_1\delta^{-2} (substitution). The sandwich of question 7 then gives

suptP(Tnt)Φ(t)K1(β+γ)32π1δ2n+K(β+γ)61δ3n3/2+δ2π.\sup_t\,\bigl|\P(T_n \leq t) - \Phi(t)\bigr| \leq \frac{K_1(\beta + \gamma)}{3\sqrt{2\pi}}\cdot \frac1{\delta^2\sqrt n} + \frac{K(\beta + \gamma)}{6}\cdot \frac1{\delta^3n^{3/2}} + \frac\delta{\sqrt{2\pi}} .

At δ=n1/6\delta = n^{-1/6} the first and third terms are O(n1/6)O(n^{-1/6}) and the middle one O(n1)O(n^{-1}): a uniform rate O(n1/6)O(n^{-1/6}), strictly better than question 7’s n1/8n^{-1/8} — the Gaussian half of the hybrid did the extra smoothing.

12. EN13=0\E N_1^3 = 0 (odd integrand), and integration by parts gives EN14=3EN12=3\E N_1^4 = 3\,\E N_1^2 = 3 (x3xφ(x) ⁣dx=3x2φ\int x^3\cdot x\varphi(x)\dd x = 3\int x^2\varphi). For ff of class C4\mathcal C^4 with bounded derivatives, expand each swap to fourth order: the third-order terms carry the factor EXi3ENi3=0\E X_i^3 - \E N_i^3 = 0 (independence factors them as in question 3), so only the fourth-order remainder 01(1θ)36f(4)(w+θh)h4 ⁣dθ\int_0^1\frac{(1-\theta)^3}6f^{(4)}(w + \theta h)h^4\dd\theta survives, with 01(1θ)36 ⁣dθ=124\int_0^1 \frac{(1-\theta)^3}6\dd\theta = \frac1{24} and Eh4=β4n2\E h^4 = \beta_4n^{-2} or 3n23n^{-2}. Question 10 (with g=f(4)g = f^{(4)}) bounds the swaps in1i \leq n - 1, and summing as in question 11:

Ef(Tn)Ef(Gn)β4+3122πf(4)L1n+M4(β4+3)24n2.\bigl|\E f(T_n) - \E f(G_n)\bigr| \leq \frac{\beta_4 + 3}{12\sqrt{2\pi}}\cdot \frac{\norm{f^{(4)}}_{L^1}}{n} + \frac{M_4(\beta_4 + 3)}{24\,n^2} .

With ψδ(4)L1=K2δ3\norm{\psi_\delta^{(4)}}_{L^1} = K_2\delta^{-3} and M4=Kδ4M_4 = K'\delta^{-4}, the distribution-function bound becomes Cδ3n1+Cδ4n2+δ2πC\delta^{-3}n^{-1} + C'\delta^{-4}n^{-2} + \frac\delta{\sqrt{2\pi}}; at δ=n1/4\delta = n^{-1/4} the outer terms are O(n1/4)O(n^{-1/4}) and the middle O(n1)O(n^{-1}): rate O(n1/4)O(n^{-1/4}).

13. With kk matched moments the surviving remainder per swap is of order Ehk+1n(k+1)/2\E\abs h^{k+1} \asymp n^{-(k+1)/2}; the hidden-Gaussian bound charges f(k+1)L1\norm{f^{(k+1)}}_{L^1} and the sum over swaps contributes the factor 2n2n, giving f(k+1)L1n(k1)/2\asymp\norm{f^{(k+1)}}_{L^1}\, n^{-(k-1)/2} for smooth ff. Ramps cost ψδ(k+1)L1δk\norm{\psi_\delta^{(k+1)}}_{L^1} \asymp \delta^{-k}, so the distribution-function error is δkn(k1)/2+δ\asymp \delta^{-k}n^{-(k-1)/2} + \delta, balanced at δ=n(k1)/(2k+2)\delta = n^{-(k-1)/(2k+2)}: rate n(k1)/(2k+2)n^{-(k-1)/(2k+2)}, which is n1/6n^{-1/6} for k=2k = 2, n1/4n^{-1/4} for k=3k = 3, and tends to n1/2n^{-1/2} only as kk \to \infty — but k4k \geq 4 would force EX14=3\E X_1^4 = 3 and beyond, i.e. a law that already imitates the Gaussian. The saturation is structural: swapping adds nn swap errors in absolute value, renouncing all cancellation between swaps. The Fourier proof compares characteristic functions, where the errors appear with their oscillating phases; Esseen’s smoothing inequality converts φTnφN\abs{\varphi_{T_n} - \varphi_N}, integrated against  ⁣dξξ\frac{\dd\xi}{\abs\xi}, into a distribution-function bound at only logarithmic cost, and delivers Berry–Esseen’s Cβn1/2C\beta n^{-1/2} from three moments. Replacement trades optimality for robustness — and, as Part VII shows, for portability.

14. Σ\Sigma is symmetric positive semidefinite; with Σ=PDPT\Sigma = PDP^{\mathsf T} (PP orthogonal, D0D \geq 0 diagonal, Exercise 20.8), the symmetric C=PDPTC = P\sqrt DP^{\mathsf T} satisfies C2=ΣC^2 = \Sigma. For any tR2t \in \R^2, t,CZ=Ct,Z=(Ct)1Z1+(Ct)2Z2\langle t, CZ\rangle = \langle Ct, Z\rangle = (Ct)_1Z^1 + (Ct)_2Z^2 is a linear combination of independent Gaussians, hence Gaussian (Exercise 23.3): N=CZN = CZ is a Gaussian vector; its mean is 00 and its covariance E[NNT]=CE[ZZT]CT=CCT=Σ\E[NN^{\mathsf T}] = C\,\E[ZZ^{\mathsf T}]\,C^{\mathsf T} = CC^{\mathsf T} = \Sigma. Moments: N3(N1+N2)34(N13+N23)\norm N^3 \leq (\abs{N_1} + \abs{N_2})^3 \leq 4(\abs{N_1}^3 + \abs{N_2}^3) (convexity of x3x^3 on R+\R_+), and each coordinate is a real Gaussian with moments of all orders (Exercise 11.10): γ<\gamma' < \infty. Finally each t,Gn=1nit,Ni\langle t, G_n\rangle = \frac1{\sqrt n}\sum_i\langle t, N_i\rangle is a normalized sum of i.i.d. N(0,tTΣt)\mathcal N(0, t^{\mathsf T}\Sigma t), hence exactly N(0,tTΣt)\mathcal N(0, t^{\mathsf T}\Sigma t): GnG_n is a Gaussian vector with mean 00 and covariance Σ\Sigma, and its law is N(0,Σ)\mathcal N(0, \Sigma) (Definition 23.10: the law is determined by these data).

15. Let ϕ(t)=f(w+th)\phi(t) = f(w + th), t[0,1]t \in \intcc01: ϕ\phi is C3\mathcal C^3 with

ϕ(t)=j,k,l{1,2}jklf(w+th)hjhkhl,ϕ(t)M3(jhj)3=M3(h1+h2)3.\phi'''(t) = \sum_{j,k,l\in\{1,2\}}\partial_{jkl}f(w + th)\,h_jh_kh_l, \qquad \abs{\phi'''(t)} \leq M_3\Bigl(\sum_j\abs{h_j}\Bigr)^3 = M_3(\abs{h_1} + \abs{h_2})^3 .

Taylor–Lagrange at order 33 for ϕ\phi between 00 and 11 gives the first inequality; Cauchy–Schwarz gives h1+h22h\abs{h_1} + \abs{h_2} \leq \sqrt2\norm h, whence the constant 22M36=2M33\frac{2\sqrt2M_3}6 = \frac{\sqrt2M_3}3.

16. Define HiH_i and WiW_i as in question 1, now in R2\R^2; the coalition argument is unchanged. In the ii-th swap, the first-order terms give jE[jf(Wi)](EXi,jENi,j)/n=0\sum_j\E[\partial_jf(W_i)]\,(\E X_{i,j} - \E N_{i,j})/ \sqrt n = 0 and the second-order terms give 12nj,kE[jkf(Wi)](ΣjkΣjk)=0\frac1{2n}\sum_{j,k}\E[\partial_{jk}f(W_i)]\,(\Sigma_{jk} - \Sigma_{jk}) = 0: means and covariances match. Question 15 bounds the two remainders:

Ef(Hi)Ef(Hi1)2M33EXi3+ENi3n3/2=2M3(β+γ)3n3/2,\bigl|\E f(H_i) - \E f(H_{i-1})\bigr| \leq \frac{\sqrt2M_3}3\cdot\frac{\E\norm{X_i}^3 + \E\norm{N_i}^3}{n^{3/2}} = \frac{\sqrt2M_3(\beta' + \gamma')}{3\,n^{3/2}},

and telescoping over the nn swaps:

Ef(Snn)Ef(Gn)2M3(β+γ)3n,GnN(0,Σ) exactly.\Bigl|\E f\Bigl(\frac{S_n}{\sqrt n}\Bigr) - \E f(G_n)\Bigr| \leq \frac{\sqrt2\,M_3(\beta' + \gamma')}{3\sqrt n}, \qquad G_n \sim \mathcal N(0, \Sigma)\ \text{exactly} .

Upgrade: ETn2=EX12=trΣ\E\norm{T_n}^2 = \E\norm{X_1}^2 = \operatorname{tr}\Sigma (cross terms vanish by independence and centering), so P(Tn>A)trΣ/A2\P(\norm{T_n} > A) \leq \operatorname{tr}\Sigma/A^2, and likewise for NN: tightness. Given a bounded continuous ff and ε>0\varepsilon > 0, multiply by a smooth plateau χ\chi equal to 11 on the ball of radius AA and supported in radius A+1A + 1 (mollify an indicator in R2\R^2, Theorem 12.9); g=fχg = f\chi is uniformly continuous with compact support, so its two-dimensional mollification gηg_\eta is C\mathcal C^\infty with bounded derivatives of all orders and ggηε\norm{g - g_\eta}_\infty \leq \varepsilon for small η\eta. The three-ε\varepsilon chain of question 5 then transfers verbatim: Ef(Tn)Ef(N)\E f(T_n) \to \E f(N) for every bounded continuous f ⁣:R2Rf \colon \R^2 \to \R. This is Theorem 23.12 for d=2d = 2, now proved — swapping sidesteps the two-dimensional Lévy theorem that the chapter had left admitted.

17. For bounded continuous g ⁣:RRg \colon \R \to \R, the map xg(t,x)x \mapsto g(\langle t, x\rangle) is bounded continuous on R2\R^2, so question 16 gives Eg(t,Tn)Eg(t,N)\E g(\langle t, T_n\rangle) \to \E g(\langle t, N\rangle): every projection converges in distribution, and t,NN(0,tTΣt)\langle t, N\rangle \sim \mathcal N(0, t^{\mathsf T}\Sigma t). (This is the easy direction of Cramér–Wold: joint convergence implies convergence of all linear images.) Application: Vi=(ξi,ξi21)V_i = (\xi_i, \xi_i^2 - 1) are i.i.d. centered vectors (Eξ12=1\E\xi_1^2 = 1), with covariance entries V(ξ1)=1\V(\xi_1) = 1, Cov(ξ1,ξ121)=Eξ13\operatorname{Cov}(\xi_1, \xi_1^2 - 1) = \E\xi_1^3 and V(ξ121)=Eξ141\V(\xi_1^2 - 1) = \E\xi_1^4 - 1; the third moment EV134(Eξ13+Eξ1213)\E\norm{V_1}^3 \leq 4\bigl(\E\abs{\xi_1}^3 + \E\abs{\xi_1^2 - 1}^3\bigr) is finite when ξ1L6\xi_1 \in L^6. Question 16 yields the displayed joint Gaussian limit, and Theorem 23.11(2): the two limit coordinates are independent exactly when the covariance Eξ13\E\xi_1^3 vanishes — for symmetric laws, empirical mean and empirical variance decouple asymptotically.

18. Write g(x)g(θ)=(g(θ)+η(x))(xθ)g(x) - g(\theta) = (g'(\theta) + \eta(x))(x - \theta) where η(x)=g(x)g(θ)xθg(θ)\eta(x) = \frac{g(x) - g(\theta)}{x - \theta} - g'(\theta) for xθx \neq \theta and η(θ)=0\eta(\theta) = 0: differentiability at θ\theta means precisely η(x)0\eta(x) \to 0 as xθx \to \theta. Step 1: θ^nθ\hat\theta_n \to \theta in probability: for ε>0\varepsilon > 0 and any A>0A > 0, eventually εnA\varepsilon\sqrt n \geq A, so P(θ^nθ>ε)P(n(θ^nθ)>A)P(σN>A)\P(\abs{\hat\theta_n - \theta} > \varepsilon) \leq \P(\abs{\sqrt n(\hat\theta_n - \theta)} > A) \to \P(\sigma\abs N > A) (distribution functions converge at the continuity points ±A\pm A), and the right side tends to 00 as AA \to \infty. Step 2: η(θ^n)0\eta(\hat\theta_n) \to 0 in probability: given ε>0\varepsilon' > 0, pick δ\delta with ηε\abs\eta \leq \varepsilon' on xθδ\abs{x - \theta} \leq \delta; then P(η(θ^n)>ε)P(θ^nθ>δ)0\P(\abs{\eta(\hat\theta_n)} > \varepsilon') \leq \P(\abs{\hat\theta_n - \theta} > \delta) \to 0. Step 3:

n(g(θ^n)g(θ))=g(θ)n(θ^nθ)+η(θ^n)n(θ^nθ).\sqrt n\bigl(g(\hat\theta_n) - g(\theta)\bigr) = g'(\theta)\,\sqrt n(\hat\theta_n - \theta) + \eta(\hat\theta_n)\cdot\sqrt n(\hat\theta_n - \theta) .

By Slutsky’s product rule (Exercise 23.8, with the sequence η(θ^n)0\eta(\hat\theta_n) \to 0 in probability and the convergent-in-law n(θ^nθ)\sqrt n(\hat\theta_n - \theta)), the second term converges in law to 0N(0,σ2)=00\cdot\mathcal N(0, \sigma^2) = 0, hence to 00 in probability (Exercise 23.4(b)); the first converges in law to g(θ)N(0,σ2)g'(\theta)\mathcal N(0, \sigma^2) (Slutsky again, or the affine rule for characteristic functions); Slutsky’s sum rule assembles them: the limit is N(0,g(θ)2σ2)\mathcal N(0, g'(\theta)^2\sigma^2).

19. (a) The CLT gives n(Xˉnμ)N(0,σ2)\sqrt n(\bar X_n - \mu) \Rightarrow \mathcal N(0, \sigma^2); the delta method with g(x)=x2g(x) = x^2, g(μ)=2μg'(\mu) = 2\mu, gives n(Xˉn2μ2)N(0,4μ2σ2)\sqrt n(\bar X_n^2 - \mu^2) \Rightarrow \mathcal N(0, 4\mu^2\sigma^2) — degenerate (limit 00) when μ=0\mu = 0. In that case the fluctuation lives one scale up: nXˉn2=(nXˉn)2n\bar X_n^2 = (\sqrt n\,\bar X_n)^2, and for t>0t > 0

P(nXˉn2t)=P(tnXˉnt)Φ(tσ)Φ(tσ)=P(σ2N2t):\P\bigl(n\bar X_n^2 \leq t\bigr) = \P\bigl(-\sqrt t \leq \sqrt n\,\bar X_n \leq \sqrt t\bigr) \longrightarrow \Phi\Bigl(\frac{\sqrt t}\sigma\Bigr) - \Phi\Bigl(-\frac{\sqrt t}\sigma\Bigr) = \P(\sigma^2N^2 \leq t) :

nXˉn2σ2N2n\bar X_n^2 \Rightarrow \sigma^2N^2, the square of a Gaussian (a “chi-squared” law) — when the first derivative dies, the second-order term of Taylor dictates a non-Gaussian limit. (b) Here n(p^np)N(0,p(1p))\sqrt n(\hat p_n - p) \Rightarrow \mathcal N(0, p(1 - p)) and g(p)=arcsinpg(p) = \arcsin\sqrt p has g(p)=12p(1p)g'(p) = \frac1{2\sqrt{p(1 - p)}}, so g(p)2p(1p)=14g'(p)^2\,p(1 - p) = \frac14: the limit is N(0,14)\mathcal N(0, \frac14) for every p(0,1)p \in \intoo01. On the arcsin\arcsin scale the asymptotic 95%95\% error bar is ±0.98n\pm \frac{0.98}{\sqrt n}, known in advance — whereas in Example 23.9 the width involved the unknown σ=p(1p)\sigma = \sqrt{p(1-p)}, to be worst-cased by 12\frac12 or estimated: the transformation stabilizes the variance.

20. Let A={k:μ({k})>ν({k})}A^* = \{k : \mu(\{k\}) > \nu(\{k\})\} and Δk=μ({k})ν({k})\Delta_k = \mu(\{k\}) - \nu(\{k\}), so kΔk=0\sum_k\Delta_k = 0. For any ANA \subseteq \N: μ(A)ν(A)=kAΔkkAΔk\mu(A) - \nu(A) = \sum_{k\in A}\Delta_k \leq \sum_{k\in A^*}\Delta_k, with equality at A=AA = A^*; and since the positive and negative parts of (Δk)(\Delta_k) have equal total mass, AΔk=12kΔk\sum_{A^*}\Delta_k = \frac12\sum_k\abs{\Delta_k}. Exchanging μ,ν\mu, \nu handles the sign: dTV(μ,ν)=12kΔkd_{\mathrm{TV}}(\mu, \nu) = \frac12\sum_k\abs{\Delta_k}. Coupling: for any AA,

μ(A)ν(A)=E[1A(X)1A(Y)]=E[(1A(X)1A(Y))1XY]P(XY),\mu(A) - \nu(A) = \E\bigl[\mathbf 1_A(X) - \mathbf 1_A(Y)\bigr] = \E\bigl[(\mathbf 1_A(X) - \mathbf 1_A(Y))\,\mathbf 1_{X\neq Y}\bigr] \leq \P(X \neq Y),

and take the supremum over AA.

21. The two laws charge: k=0k = 0: 1p1 - p versus ep\eu^{-p}, with ep>1p\eu^{-p} > 1 - p; k=1k = 1: pp versus pep<pp\,\eu^{-p} < p; k2k \geq 2: 00 versus the Poisson remainder 1eppep01 - \eu^{-p} - p\eu^{-p} \geq 0. Hence

dTV=12[(ep1+p)+(ppep)+(1eppep)]=12(2p2pep)=p(1ep),d_{\mathrm{TV}} = \tfrac12\bigl[(\eu^{-p} - 1 + p) + (p - p\eu^{-p}) + (1 - \eu^{-p} - p\eu^{-p})\bigr] = \tfrac12\bigl(2p - 2p\eu^{-p}\bigr) = p(1 - \eu^{-p}),

and 1epp1 - \eu^{-p} \leq p gives the bound p2p^2.

22. Write Hi1=Wi+XiH_{i-1} = W_i + X_i and Hi=Wi+YiH_i = W_i + Y_i with Wi=j<iYj+j>iXjW_i = \sum_{j<i}Y_j + \sum_{j>i}X_j, independent of the pair (Xi,Yi)(X_i, Y_i) (coalitions). For ANA \subseteq \N, conditioning on the countably many values by independence,

P(Hi1A)=k0P(Xi=k)P(Wi+kA),\P(H_{i-1} \in A) = \sum_{k\geq0}\P(X_i = k)\,\P(W_i + k \in A),

and likewise for HiH_i with YiY_i. Subtracting, with ck=P(Wi+kA)[0,1]c_k = \P(W_i + k \in A) \in \intcc01 and Δk=P(Xi=k)P(Yi=k)\Delta_k = \P(X_i = k) - \P(Y_i = k) of zero sum:

P(Hi1A)P(HiA)=kΔk(ck12)12kΔk=dTV(B(1,pi),P(pi)).\abs{\P(H_{i-1} \in A) - \P(H_i \in A)} = \Bigl|\sum_k\Delta_k\bigl(c_k - \tfrac12\bigr)\Bigr| \leq \tfrac12\sum_k\abs{\Delta_k} = d_{\mathrm{TV}}\bigl(\mathcal B(1, p_i), \mathcal P(p_i)\bigr) .

Telescoping from H0=SH_0 = S to Hn=iYiP(λ)H_n = \sum_iY_i \sim \mathcal P(\lambda) (Exercise 23.1, iterated) and using question 21:

P(SA)P(P(λ)A)i=1npi(1epi)i=1npi2for every A:\abs{\P(S \in A) - \P(\mathcal P(\lambda) \in A)} \leq \sum_{i=1}^np_i\bigl(1 - \eu^{-p_i}\bigr) \leq \sum_{i=1}^np_i^2 \qquad\text{for every } A :

Le Cam’s inequality. (Question 20’s coupling bound gives an alternative route: couple each pair on one uniform variable so that P(XiYi)pi2\P(X_i \neq Y_i) \leq p_i^2 and bound P(SYi)\P(S \neq \sum Y_i); swapping needs no construction at all.)

23. (a) With pi=λnp_i = \frac\lambda n: dTV(law of S,P(λ))λ2nd_{\mathrm{TV}}(\text{law of }S, \mathcal P(\lambda)) \leq \frac{\lambda^2}n. This sharpens Exercise 23.5 three times over: an explicit error at every finite nn, uniformity over all events AA at once (not one interval at a time), and no need for equal pip_i — only ipi2\sum_ip_i^2 small, e.g. pi2λmaxipi\sum p_i^2 \leq \lambda\max_ip_i: many rare events, none dominant. (b) Here n=500n = 500, pi=1500p_i = \frac1{500}, λ=1\lambda = 1: the Poisson model errs by at most 50015002=0.002500\cdot\frac1{500^2} = 0.002 on every event; in particular, taking A={0}A = \{0\},

P(no letter astray)=(11500)500,P(no letter astray)e10.002,\P(\text{no letter astray}) = \Bigl(1 - \frac1{500}\Bigr)^{500}, \qquad \Bigl|\P(\text{no letter astray}) - \eu^{-1}\Bigr| \leq 0.002,

so the answer is e10.368\eu^{-1} \approx 0.368 up to a guaranteed 0.0020.002 (the true discrepancy is about 41044\cdot10^{-4}). (c) The problem closes on one method with two regimes. When nn comparable contributions each carry variance 1n\frac1n, matching two moments against the Gaussian makes the swap errors o(1n)o(\frac1n) each: sums go Gaussian — with Taylor as the local comparison tool. When nn contributions are indicators of probability pip_i, matching the mean against a Poisson atom makes each swap cost pi2p_i^2: counts of rare events go Poisson — with total variation as the exact local comparison. Same hybrids, same telescope, different local estimate: replacement is a strategy, not a theorem, and the Gaussian and Poisson limits are its two oldest dividends.

24. The CLT gives n(Xˉnμ)N(0,σ2)\sqrt n(\bar X_n - \mu) \Rightarrow \mathcal N(0, \sigma^2), and g(x)=lnxg(x) = \ln x is differentiable at μ>0\mu > 0 with g(μ)=1μg'(\mu) = \frac1\mu: the delta method (Part VI) yields n(lnXˉnlnμ)N(0,σ2/μ2)\sqrt n(\ln\bar X_n - \ln\mu) \Rightarrow \mathcal N(0, \sigma^2/\mu^2). Unwinding the interval lnXˉnlnμ1.96σμn\abs{\ln\bar X_n - \ln\mu} \leq \frac{1.96\,\sigma}{\mu\sqrt n} by exponentiation:

μXˉne±1.96σ/(μn)with asymptotic probability 95%\mu \in \bar X_n\cdot \eu^{\pm1.96\,\sigma/(\mu\sqrt n)} \qquad\text{with asymptotic probability } 95\%

(in practice σ/μ\sigma/\mu is replaced by its empirical version, Slutsky as in Exercise 23.8). The multiplicative interval is the natural one when the data are positive with errors proportional to their size — incomes, concentrations, half-lives: quantities that live on a log scale, where symmetric additive intervals could even cross zero.

25. E[(Xp)3]=(1p)3p+(p)3(1p)=p(1p)[(1p)2p2]=p(1p)(12p)\E[(X - p)^3] = (1-p)^3p + (-p)^3(1 - p) = p(1-p)\bigl[(1-p)^2 - p^2\bigr] = p(1-p)(1 - 2p). In the swapping analysis (Part IV), the leading error term after matching two moments carries the signed third moment: for p<12p < \frac12 it is positive (the law leans right: rare large excursions above the mean), and the normal approximation systematically misplaces mass — underestimating the short left tail and overestimating the right — with an error of order n1/2n^{-1/2}; at p=12p = \frac12 the third moment vanishes, the Bernoulli matches the Gaussian to third order, and the rate improves (Part IV’s matched-moment question). Numerically: P(S=0)=0.920=0.1216\P(S = 0) = 0.9^{20} = 0.1216, while the Gaussian N(2,1.8)\mathcal N(2, 1.8) gives Φ(0.521.8)=Φ(1.118)0.132\Phi\bigl(\frac{0.5 - 2}{\sqrt{1.8}}\bigr) = \Phi(-1.118) \approx 0.132: the normal curve, ignorant of the wall at 00 and of the rightward skew, puts too much mass at the bottom — the predicted sign of the error, visible at n=20n = 20.