University Mathematics — Year 3 · Bachelor Year 3
23Characteristic Functions and the Central Limit Theorem
The law of large numbers says averages converge; the central limit theorem says how they fluctuate: the error, magnified by , is asymptotically Gaussian — whatever the law one started from. This universality is the deepest fact of elementary probability, and its natural proof is Fourier-analytic: the characteristic function (the Fourier transform of a law) converts independent sums into products, and Chapter 14’s machinery — injectivity, Gaussian fixed points — converts pointwise convergence of these products into convergence of laws (Lévy’s theorem, proved in full). The chapter ends with Gaussian vectors and the honest derivation of the confidence intervals used everywhere in statistics; the weekend problem gives Lindeberg’s second proof of the CLT, with an explicit error rate.
23.1 Characteristic functions
Definition 23.1
The characteristic function of a real random variable is
(the transfer theorem computes it from the law; for a density , in Chapter 14’s convention).
Proposition 23.2
(a) , , and is uniformly continuous; . (b) If are independent: . (c) If , then with for ; in particular, for centered with variance :
(d) Gaussian: has .
Proof. (a) Bounds are immediate; continuity: as by dominated convergence, uniformly in . The affine rule is a substitution. (b) , and expectations of products of independent variables factor (Theorem 22.5, applied to real and imaginary parts). (c) Differentiation under the expectation, dominated by (Theorem 10.15); the Taylor expansion at is then Taylor–Young for the function . (d) For : the Gaussian transform (Example 14.2 with ) gives ; the general case by the affine rule. ∎
Theorem 23.3 (Injectivity)
If , then and have the same law. More precisely, for independent of and , the smoothed variable has the density
determined by alone; letting recovers the law of .
Proof. has the density , where is the density: indeed for Borel , independence and Tonelli give (substitute, then Tonelli again). Writing by Fourier inversion of its transform (Exercise 14.4, rescaled): , and Fubini (everything dominated by the Gaussian factor):
a functional of alone. If : and have equal laws for every ; for bounded continuous , as (dominated convergence, pointwise on the product space), so for all such — and this determines the law: for each , squeeze between the bounded continuous ramps (equal to on , to beyond , affine between); passing to the limit in gives at every where both are continuous, hence everywhere by right-continuity and density of common continuity points (both ’s have countably many jumps); equal distribution functions force equal laws (Exercise 9.3, resting on Theorem 9.7). ∎
23.2 Convergence in distribution
Definition 23.4
converges in distribution (or in law) to , written , if
Equivalently (Exercise 23.4): at every continuity point of . The need not live on a common probability space: only laws matter.
Theorem 23.5 (Helly’s selection theorem)
Every sequence of distribution functions has a subsequence converging pointwise, at every continuity point of the limit, to a nondecreasing right-continuous — possibly with (mass may escape to infinity).
Proof. Diagonal extraction gives for every rational (values in the compact ). Define : nondecreasing; right-continuous (an infimum over shrinking rational neighborhoods from the right). At a continuity point of : for rationals ,
by monotonicity of each . From the definition of as an infimum and monotonicity of on the rationals: whenever . Taking gives , and ; letting and , continuity of at squeezes both the and the to . ∎
Lemma 23.6 (Tightness from the characteristic function)
For any random variable and :
Proof. By Tonelli–Fubini (integrand bounded, region finite in ):
(interpret the bracket as its limit at ). The integrand is nonnegative (), and for : . Keeping only the event inside the expectation therefore leaves at least , which is the claim. ∎
Theorem 23.7 (Lévy’s continuity theorem)
Let be random variables whose characteristic functions converge pointwise: for every , where is the characteristic function of some random variable . Then .
Proof. Tightness. Fix . Since is continuous at with , choose with ; by dominated convergence (integrand bounded by on the fixed ), the same integral for is for large: Lemma 23.6 gives for large , and enlarging the constant handles the finitely many others: the laws are tight — no mass escapes.
Subsequences. Let be any subsequence; by Helly (Theorem 23.5) extract at continuity points. Tightness forces , ( at continuity points): is a genuine distribution function, of some random variable . Then (Exercise 23.4, distributional convergence from ’s), so pointwise ( is bounded continuous, real and imaginary parts separately); comparing with the hypothesis: , and injectivity (Theorem 23.3) gives , i.e. .
Conclusion. Every subsequence of has a sub-subsequence converging to the same (at its continuity points); hence at every continuity point (a real sequence all of whose subsequences have subsubsequences with the same limit converges): . ∎
23.3 The central limit theorem
Theorem 23.8 (Central limit theorem)
Let be i.i.d. with and . Then
for all .
Proof. Center and normalize: (i.i.d., mean , variance ) and . By independence and the affine rule (Proposition 23.2):
Fix and let , : both have modulus for large ( once ; always). The elementary inequality for (telescoping ) gives
while (real logarithm). So (Proposition 23.2(d)) for every : Lévy (Theorem 23.7) concludes . The interval probabilities follow since is continuous everywhere. ∎
Example 23.9 (Confidence intervals, honestly derived)
Poll independent voters; estimates the true , with . The CLT gives, for large ,
where is the standard Gaussian distribution function. With : asymptotic confidence , and margin requires — the number behind every “ points, ” one reads; compare Chebyshev’s (Exercise 22.7). The is universal: to halve the error, quadruple the sample — the same law that fixes Monte Carlo’s cost (Exercise 23.7).
23.4 Gaussian vectors
Definition 23.10
A random vector is Gaussian if every linear combination is a (possibly degenerate) real Gaussian variable. Its law is determined by the mean vector and the covariance matrix : indeed the characteristic function of the vector, , is the value at of the cf of :
and -dimensional characteristic functions are injective (same smoothing proof as Theorem 23.3, coordinatewise Gaussians).
Theorem 23.11
Let be a Gaussian vector.
- Every affine image is a Gaussian vector.
- The components are independent if and only if is diagonal: for jointly Gaussian variables, uncorrelated independent.
- If is invertible, has the density .
Proof. (1) Linear combinations of components of are affine functions of linear combinations of : Gaussian (an affine image of a Gaussian variable is Gaussian). (2) If is diagonal, the characteristic function factorizes: , which is the characteristic function of the product law (Theorem 22.5 read through -dimensional injectivity): the components are independent. The converse is the vanishing of covariances of independent variables. (3) Diagonalize ( orthogonal, diagonal — Exercise 20.8); the vector is Gaussian with covariance : by (2) its components are independent , so has the product density; push forward by the volume-preserving (Theorem 11.10, ) and rewrite the exponent invariantly. ∎
Theorem 23.12 (Multidimensional CLT)
Let be i.i.d. square-integrable random vectors of with mean and covariance matrix . Then converges in distribution to the Gaussian vector .
Proof. Admitted at this level. ∎
Remark 23.13
Almost everything is already in our hands. For each direction , the real variable is a normalized sum of i.i.d. real variables of variance , so Theorem 23.8’s computation gives pointwise convergence of the -dimensional characteristic functions to , the characteristic function of (Definition 23.10). What we have not re-proved is Lévy’s continuity theorem in : Helly’s selection and the tightness estimate generalize routinely (coordinatewise), and this Cramér–Wold reduction is carried out honestly in any graduate probability course; nothing beyond this chapter’s methods is needed.
Method 23.14
To identify a limit law: compute characteristic functions, take the pointwise limit, recognize it (Gaussian , Poisson , exponential , …) and invoke Lévy. The three-step ritual (independence product; Taylor at exponential limit; Lévy convergence in law) proves the CLT, the Poisson law of rare events (Exercise 23.5), and every classical limit theorem of this course. For a.s. statements, go back to Chapter 22’s toolkit: the two chapters answer different questions about the same .
23.5 Exercises
Exercise 23.1 ★
Compute the characteristic functions: uniform on ; exponential ; Poisson ; binomial . Deduce via Theorem 23.3 that the sum of independent Poisson variables () is Poisson .
Solution
Solution of Exercise 23.1.
Uniform on : (equal to at ). Exponential : (the antiderivative vanishes at since ). Poisson : by the transfer theorem for discrete laws,
Binomial : a sum of independent Bernoulli variables, each with cf , so (Proposition 23.2(b)). Poisson additivity: if , are independent,
the cf of ; injectivity (Theorem 23.3) identifies the law.
Exercise 23.2 ★★
(a) Show that is real-valued if and only if and have the same law (a symmetric variable). (b) Suppose for some . Show that is almost surely supported on an arithmetic progression (write and compute ). Deduce that if has a density, then for all .
Solution
Solution of Exercise 23.2.
(a) . So is real if and only if , if and only if (injectivity, Theorem 23.3) and have the same law. (b) Write . Then
The integrand is nonnegative, so almost surely (a nonnegative variable with zero expectation vanishes a.s.), i.e. a.s.: takes its values in the arithmetic progression almost surely. If has a density, this countable set is Lebesgue-null, so it carries probability — contradiction; therefore for every .
Exercise 23.3 ★★
Let and be independent. Show , and more generally that the Gaussian family is stable under independent sums and affine maps. Contrast: is the sum of two dependent Gaussians always Gaussian? (Exercise 23.9.)
Solution
Solution of Exercise 23.3.
By independence and Proposition 23.2:
the cf of ; injectivity concludes. Stability under affine maps is the affine rule (, allowing the degenerate case ), and stability under independent sums follows by induction on the computation above. For dependent Gaussians the sum need not be Gaussian: in Exercise 23.9, and are each standard Gaussian but vanishes with probability without being a.s. zero, so it is not Gaussian.
Exercise 23.4 ★★
(a) Prove the equivalence in Definition 23.4: if for all bounded continuous , then at continuity points (squeeze between two continuous staircase-ramps); and conversely (approximate a bounded continuous by sums of ramp functions, or condition on a fine grid of continuity points) — the converse may be treated for uniformly continuous first, then in general. (b) Show that (a constant) implies in probability.
Solution
Solution of Exercise 23.4.
(a) Direct implication. Let be a continuity point of and . Take the continuous ramps ( on , from on, affine between) and ( on , from on, affine between); then , so
and the outer terms converge to , themselves squeezed between and . Letting then and using continuity of at : .
Converse. Let be bounded continuous, , . The continuity points of are dense ( has at most countably many jumps), so choose continuity points with and . On the compact the function is uniformly continuous: choose continuity points of with oscillation of at most on each , and set . Then on , , and for or :
Moreover (a finite sum of converging terms, all being continuity points), and , . Assembling: ; let .
(b) The distribution function of the constant is , continuous except at . For , the points and are continuity points, so
Exercise 23.5 ★★
(Law of rare events) Let with . Show, via characteristic functions and Theorem 23.7, that . Numerical sanity check: compare for and .
Solution
Solution of Exercise 23.5.
Let , so (Exercise 23.1) and (note ). Both and have modulus at most : by the triangle inequality, and . The telescoping inequality (proof of Theorem 23.8) and the power series bound give
Since , we conclude for every : the cf of , and Lévy (Theorem 23.7) gives . Numerically: , while : two percent apart already at this coarse .
Exercise 23.6 ★★
(a) A fair die is rolled times; approximate the probability that the total exceeds (mean , variance per roll ). (b) For , approximate by the CLT with the continuity correction (), and comment on the correction’s effect.
Solution
Solution of Exercise 23.6.
(a) One roll has mean and variance , so has mean , variance and standard deviation . By the CLT,
about a chance. (b) : mean , standard deviation . With the continuity correction,
against the exact value ; without the correction, , off by almost five points. The correction matters because is a lattice variable: the atom is well approximated by the Gaussian mass of , and clipping the interval at the integers and discards half an atom at each end.
Exercise 23.7 ★★
(Monte Carlo error) In the setting of Problem 22.1, question 11, with , let and . Show
and deduce the asymptotic error bar — independent of the dimension . Compare with the deterministic midpoint rule in dimension (error for integrands): from which dimension on does random sampling win?
Solution
Solution of Exercise 23.7.
The variables are i.i.d. (measurable images of i.i.d. variables), square-integrable, with mean (transfer theorem, Exercise 11.9) and variance . If , Theorem 23.8 applied to them is exactly the stated convergence
(if , is a.s. constant and the left-hand side vanishes identically). Hence : the error bar sees the dimension only through the constant , never through the rate in . The midpoint rule with nodes in dimension has mesh and error of order for integrands. Monte Carlo’s decays faster than exactly when , i.e. : from dimension on, random sampling asymptotically beats the grid — the curse of dimensionality spares probabilistic methods, which is why Monte Carlo rules high-dimensional integration.
Exercise 23.8 ★★★
(Slutsky) Suppose and in probability ( constant). Show and . (Work with characteristic functions and the bound , split on .) Application: in Example 23.9, justify replacing the unknown by .
Solution
Solution of Exercise 23.8.
Sum. For fixed :
Split on the event : there, (the chord is shorter than the arc); the complement contributes at most . Hence the is for every : the difference tends to . Since , we get , and Lévy (Theorem 23.7) yields .
Product. First, : . Next, in probability: the laws of the are tight (their cfs converge to a cf; see the tightness step of Theorem 23.7), so given pick with for all ; then
Writing and applying the sum part (whose proof only used in probability, with constant ): .
Application. By the strong law of large numbers (Theorem 22.13), a.s., so by continuity a.s., hence in probability. Slutsky’s product rule upgrades to : the usable confidence interval , built from the data alone, keeps its asymptotic level.
Exercise 23.9 ★★★
Let and independent with ; set . (a) Show and . (b) Show and are not independent, and that is not a Gaussian vector (compute ). (c) Moral: Theorem 23.11(2) requires joint Gaussianity — “uncorrelated Gaussians” alone proves nothing.
Solution
Solution of Exercise 23.9.
(a) Splitting the expectation over the two values of (independence): for Borel , , since ( is symmetric): . And . (b) , so while : not independent. If were a Gaussian vector, would be a real Gaussian variable (Definition 23.10 with ); but , whereas a Gaussian variable has an atom only if it is a.s. constant — and equals a.s. on . Contradiction: is not Gaussian. (c) Each marginal is Gaussian and the covariance vanishes, yet independence fails — because the pair is not jointly Gaussian. Theorem 23.11(2) cannot be weakened to “Gaussian marginals”.
Exercise 23.10 ★★
The Cauchy law has density . (a) Show its characteristic function is (Exercise 14.1 and inversion). (b) Show that if are i.i.d. Cauchy, then is again Cauchy — the same law: the average never concentrates. (c) Reconcile with the laws of large numbers and the CLT: which hypotheses fail? (Compute .)
Solution
Solution of Exercise 23.10.
(a) Exercise 14.1 computes ; both sides being integrable, Fourier inversion (Theorem 14.5) turns this around:
which is exactly for a Cauchy variable . (b) By independence, , so : the empirical mean is again standard Cauchy for every (injectivity). The average never concentrates: its fluctuations at time are those of a single observation. (c) : the Cauchy law is not integrable, so the strong law of large numbers (Theorem 22.13) does not apply, and the CLT (which needs a finite variance) even less. Here their conclusions genuinely fail, not merely their proofs. Consistency check: is not differentiable at , as Proposition 23.2(c) read contrapositively predicts for a non-integrable variable.
Exercise 23.11 ★★
(Stable laws in embryo) Let be i.i.d. standard Cauchy (Exercise 23.10). (a) Show that for any , has the law of : the Cauchy family is strictly stable of index . (b) Show that the Gaussian family is strictly stable of index : for i.i.d. . (c) Explain, via characteristic functions of the form , why index- stability forces the normalization for sums, and what this says about the basins of attraction of the CLT: which i.i.d. sums can converge, after affine normalization, to a Cauchy law rather than a Gaussian?
Solution
Solution of Exercise 23.11.
(a) (independence and Exercise 23.10); injectivity identifies the laws.
(b) : the law of .
(c) If , then has , and has again: exact self-reproduction under the scaling — for the Gaussian (), itself for Cauchy (, Exercise 23.10(b)). A sum of i.i.d. variables can only converge (after affine normalization) to a law that is stable under such convolutions; the CLT says finite variance forces the Gaussian basin, and the Cauchy basin is reserved for laws with tails so heavy that and even — e.g. sums of Cauchy variables themselves. Universality has several islands, indexed by the tail exponent .
Exercise 23.12 ★★
(The empirical distribution function) Let be i.i.d. with distribution function , and . (a) Fix . Show that , that a.s. (Theorem 22.13), and that
(b) At which is the asymptotic variance maximal? Interpret: the median is where an empirical distribution is hardest to pin down. (c) For continuous, show that the law of does not depend on (reduce to uniform variables via Exercise 22.1) — the distribution-free miracle behind the Kolmogorov–Smirnov test; no computation of that law is asked.
Solution
Solution of Exercise 23.12.
(a) The indicators are i.i.d. Bernoulli of parameter : their sum is binomial ; the strong law gives a.s., and the CLT (Theorem 23.8) applied to the same indicators (variance ) gives the stated Gaussian limit.
(b) is maximal at , i.e. where : at the median. Estimating tail probabilities is asymptotically easy (variance as ); the median region carries the largest statistical noise — the empirical curve wobbles most in its middle.
(c) For continuous , the variables are i.i.d. uniform on (Exercise 22.1), and monotonicity of gives, writing for the empirical distribution function of the :
the first equality because up to null events (monotonicity; strict inequality can fail only on the flat parts of , where both sides are unchanged), and the second because a continuous , running from to , attains every value of (intermediate value theorem), and the endpoints add nothing ( and ). The right side involves only uniforms: one law for all — so a single table of critical values (that of Kolmogorov’s distribution) tests any continuous model against data.
23.6 Problem: Lindeberg’s proof of the CLT, with a rate
Problem 23.1
Weekend problem — the replacement method
Lindeberg (1922) proved the central limit theorem by an idea of disarming simplicity: swap the summands one at a time for Gaussians and control each swap by a Taylor expansion. The method needs no Fourier analysis, produces an explicit error rate, and today powers universality proofs across probability theory. Let be i.i.d., centered, , with ; let be i.i.d. , independent of the (existence: Theorem 22.6). Set
Part I — The swapping identity. Fix (three bounded continuous derivatives; ). For define the hybrid sums
so and .
- Write and with , and note that is independent of the pair . Justify.
Taylor with integral or Lagrange remainder: for any real :
Apply question 2 twice ( and at ), take expectations, and use independence plus the matching of the first two moments of and to show
Telescope over and conclude the Lindeberg bound:
Part II — From smooth to the CLT.
- Show that for every , and upgrade to all bounded continuous : given such and , construct with on a large interval — e.g. convolve with a bump (Theorem 12.9) — and handle the tails by tightness ( and Chebyshev). Conclude : the central limit theorem, re-proved.
- Where did the proof use that the are identically distributed? Show that it barely did: state and prove the version for independent, centered, non-identical with and third moments, obtaining the error — Lindeberg’s true theorem in its Lyapunov form.
Part III — Quantitative dividends.
(Distribution functions) Let and approximate above and below by ramps of width (construct them, with ). Combining with Part I, derive the two-term bound
with explicit constants (the term uses that has density bounded by ), and optimize to obtain a uniform rate of order . (The optimal — Berry–Esseen — needs finer tools; the point is an explicit rate from elementary swapping.)
- (De Moivre–Laplace, quantified) Specialize to (signs of fair coins): compare the conclusion with the local estimate of Problem 11.1, question 7 — what does each method give that the other does not?
- (Universality) Explain in a paragraph why the replacement method shows more than the CLT: any statistic of the form with smooth is insensitive, at order , to the entire law of the summands beyond its first two moments — the “invariance principle” that underlies modern universality results (random matrices, random polynomials), of which the CLT is the first instance.
Part IV — Smoothing, pushed: better rates. The loss from (smooth ) to (distribution functions) came from charging in sup norm. The hybrids can repair part of it: they contain Gaussian summands, and Gaussians smooth.
(A hidden Gaussian) For , or , and , write with . Show that is independent of the pair , and deduce, for every continuous ,
Combine question 10 with the integral form of the Taylor remainder,
to redo questions 3–4: for with moreover ,
(question 10 handles the swaps — use — and question 3’s crude bound handles the last one). Check that question 7’s ramps satisfy while , feed them in, and optimize : the uniform distribution-function rate improves to .
- (Matching one more moment) Assume in addition and . Compute and , expand to fourth order, and prove along the same lines that the distribution-function rate becomes (now and ; choose ).
- (The obstruction) Suppose the first moments of agree with the Gaussian ones ( always; exactly when ; essentially never, as ). Verify that the scheme of questions 10–12 delivers the distribution-function rate , by balancing against , and observe that the exponent approaches the Berry–Esseen value only as . Explain in a few sentences why the swapping method saturates: each swap is charged in absolute value, whereas the Fourier route (Esseen’s smoothing inequality) exploits the oscillation of the characteristic-function difference and reaches with three moments only.
Part V — Two dimensions: the multidimensional CLT, by swapping. Now let the be i.i.d. centered random vectors of with covariance matrix and (Euclidean norm).
- (Gaussian vectors, to order) Diagonalize (Exercise 20.8) and set . For a pair of independent standard Gaussians (Theorem 22.6), show that is a Gaussian vector (Definition 23.10) of mean , covariance , with ; and that has law exactly for i.i.d. copies .
(Taylor in two variables) For of class with , prove
(study on ).
(The CLT in ) Run the replacement scheme on the vector hybrids : show that the first- and second-order terms cancel (means and covariances match), telescope, and upgrade as in question 5 (tightness from ; mollification now in , Theorem 12.9) to conclude: for every bounded continuous ,
Theorem 23.12 in dimension , with a rate for smooth and no Fourier analysis.
(Cramér–Wold, and a joint fluctuation) Deduce that for every fixed . Application: for i.i.d. real , centered, , (so that Part V applies to ), show
empirical mean and empirical second moment fluctuate jointly Gaussianly — independently in the limit if and only if (Theorem 23.11).
Part VI — The delta method.
Let be random variables with for a real parameter , and let be differentiable at . Prove the delta method:
(write with at ; show , then , in probability; finish with Slutsky, Exercise 23.8, and Exercise 23.4(b)).
Applications. (a) For i.i.d. real with mean and variance , and : show when , and that for the correct statement lives at another scale: with (identify the limit’s distribution function). (b) (Variance stabilization) For the success frequency of a sample, : show that satisfies
whatever — an asymptotic error bar free of the unknown parameter; compare with Example 23.9.
Part VII — Poisson, by the same method: Le Cam’s theorem. Replacement knows a second universality class: sums of many independent rare events. For laws on the right distance is total variation,
- Show that , and prove the coupling bound: for any pair of random variables with laws and on the same space, .
Compute exactly, for :
(Le Cam, by swapping) Let and , the variables independent; , and recall with (Exercise 23.1). Swap one coordinate at a time in the integer hybrids : show, for every ,
and conclude Le Cam’s inequality:
- Dividends. (a) For : the bound is — the law of rare events (Exercise 23.5) upgraded to an explicit rate, uniform over all events, and valid for unequal as well. (b) letters are delivered, each going astray independently with probability : bound the error of the Poisson model of parameter , and estimate the probability that no letter goes astray. (c) Close the problem: compare the two universality classes met here — Gaussian (many small spread-out contributions; two moments matched; Taylor) and Poisson (many rare contributions; one mean matched; an exact total-variation coupling) — and the single replacement method behind both.
(Relative error and the log transform) Let be i.i.d., positive, mean , variance , and the empirical mean. Show by the delta method that
the asymptotic parameter of is the coefficient of variation — relative, scale-free error. Deduce a confidence interval for of the multiplicative form , and explain when it is preferable to the additive one.
- (The third moment steers the error) For Bernoulli() centered, compute . Using Part IV’s analysis (the swap error is driven by third moments), explain why the normal approximation of is asymmetric for — overshooting on one side, undershooting on the other — and why enjoys the faster matched-moment rate. Verify the sign of the skew numerically on against : compare with the Gaussian mass of .
Solution
Solution of Problem 23.1.
1. The family is independent: the two blocks are independent of each other by construction and each block is i.i.d. is a measurable function of the variables and only, all distinct from and : by the coalition principle (Theorem 22.5), is independent of the pair . The decompositions and are immediate from the definitions: passing from to swaps the single summand for .
2. Taylor–Lagrange at order : there is between and with , and gives the bound.
3. Subtracting the two expansions at the common base point :
Take expectations. By question 1, and are independent of , so the mixed expectations factor:
the first two moments of and match, and only the remainder survives:
The Gaussian third moment: (substitution , then ).
4. Telescoping and applying question 3 to each of the terms:
5. is exactly for every (a normalized sum of independent standard Gaussians, Exercise 23.3), so and question 4 reads for . Upgrade. Let be bounded continuous, , . Choose with : Chebyshev with gives for all , and likewise . Let be with (a smooth plateau, built by mollifying , Theorem 12.9); is continuous with compact support, hence uniformly continuous, so its mollification is with bounded derivatives of all orders and for small enough. For or , since on and everywhere:
Combining with (question 4 applies: ):
and was arbitrary: for every bounded continuous , i.e. .
6. Identical distribution entered only through one sentence: “ and have the same first two moments”. So let be independent, centered, with variances and finite third moments, , and take independent of everything. Define the hybrids with normalization : . In the -th swap, and again kill the and terms, and the remainder gives (using by scaling):
Telescoping:
Since (the power–mean inequality, i.e. Jensen for applied to ), the right side is at most : under Lyapunov’s condition , the normalized sums converge in law to — the CLT without identical distribution.
7. Let with and set : is , nonincreasing, on , on ; let . For and define and : these are with third derivative bounded by , and
Upper bound: by question 4 applied to (with ),
because is Lipschitz with constant (its density is bounded by ). The symmetric lower bound via gives the two-term estimate
The two terms balance when , i.e. : both are then , an explicit uniform rate valid for every . (The optimal Berry–Esseen rate requires the Fourier smoothing inequality; swapping trades sharpness for complete elementarity.)
8. For (fair signs): centered, variance , and so . Question 7 then bounds explicitly and uniformly for every finite — a global, non-asymptotic statement about the distribution function. The local estimate of Problem 11.1, question 7, gives instead the exact asymptotics of an individual atom, : it resolves probabilities of size , far below question 7’s resolution, but it is pointwise, asymptotic (no explicit error at fixed ) and tied to this particular lattice law. Local precision versus global uniformity: the two methods are complementary, and summing the local estimate over recovers de Moivre–Laplace on intervals — with a sharper rate, but only for this law.
9. The swapping argument used nothing about the law of the beyond , and the finiteness of : had we replaced the Gaussians by any other i.i.d. family with the same first two moments and finite third moment, the same telescoping would bound by for every smooth . Smooth statistics of large independent sums are therefore universal: up to a quantified error, they depend on the summands’ law only through two numbers. This is the invariance principle: prove a limit theorem for the most computable law (the Gaussian, where everything is exact), then transfer it to all laws by swapping. The same scheme — with sums replaced by more elaborate functionals — drives Wigner’s semicircle law for random matrices, the universality of roots of random polynomials, and much of modern probability; the central limit theorem is its first and simplest instance.
10. is a Borel function of only, while and are functions of the remaining variables of the independent family : by the coalition principle (Theorem 22.5), is independent of . As a sum of the independent , with (Exercise 23.3), with density bounded by . The law of is the product of the two marginal laws, so Tonelli (transfer) freezes the first block: with for every ,
11. The integral form of Taylor’s formula follows by integrating by parts twice in . Taking expectations in the -th swap, the orders cancel exactly as in question 3, and the two remainders (for and ) are bounded, for , by question 10 with :
Summing, with , and adding question 3’s bound for the last swap (, no Gaussian left):
Ramps: , so with and (substitution). The sandwich of question 7 then gives
At the first and third terms are and the middle one : a uniform rate , strictly better than question 7’s — the Gaussian half of the hybrid did the extra smoothing.
12. (odd integrand), and integration by parts gives (). For of class with bounded derivatives, expand each swap to fourth order: the third-order terms carry the factor (independence factors them as in question 3), so only the fourth-order remainder survives, with and or . Question 10 (with ) bounds the swaps , and summing as in question 11:
With and , the distribution-function bound becomes ; at the outer terms are and the middle : rate .
13. With matched moments the surviving remainder per swap is of order ; the hidden-Gaussian bound charges and the sum over swaps contributes the factor , giving for smooth . Ramps cost , so the distribution-function error is , balanced at : rate , which is for , for , and tends to only as — but would force and beyond, i.e. a law that already imitates the Gaussian. The saturation is structural: swapping adds swap errors in absolute value, renouncing all cancellation between swaps. The Fourier proof compares characteristic functions, where the errors appear with their oscillating phases; Esseen’s smoothing inequality converts , integrated against , into a distribution-function bound at only logarithmic cost, and delivers Berry–Esseen’s from three moments. Replacement trades optimality for robustness — and, as Part VII shows, for portability.
14. is symmetric positive semidefinite; with ( orthogonal, diagonal, Exercise 20.8), the symmetric satisfies . For any , is a linear combination of independent Gaussians, hence Gaussian (Exercise 23.3): is a Gaussian vector; its mean is and its covariance . Moments: (convexity of on ), and each coordinate is a real Gaussian with moments of all orders (Exercise 11.10): . Finally each is a normalized sum of i.i.d. , hence exactly : is a Gaussian vector with mean and covariance , and its law is (Definition 23.10: the law is determined by these data).
15. Let , : is with
Taylor–Lagrange at order for between and gives the first inequality; Cauchy–Schwarz gives , whence the constant .
16. Define and as in question 1, now in ; the coalition argument is unchanged. In the -th swap, the first-order terms give and the second-order terms give : means and covariances match. Question 15 bounds the two remainders:
and telescoping over the swaps:
Upgrade: (cross terms vanish by independence and centering), so , and likewise for : tightness. Given a bounded continuous and , multiply by a smooth plateau equal to on the ball of radius and supported in radius (mollify an indicator in , Theorem 12.9); is uniformly continuous with compact support, so its two-dimensional mollification is with bounded derivatives of all orders and for small . The three- chain of question 5 then transfers verbatim: for every bounded continuous . This is Theorem 23.12 for , now proved — swapping sidesteps the two-dimensional Lévy theorem that the chapter had left admitted.
17. For bounded continuous , the map is bounded continuous on , so question 16 gives : every projection converges in distribution, and . (This is the easy direction of Cramér–Wold: joint convergence implies convergence of all linear images.) Application: are i.i.d. centered vectors (), with covariance entries , and ; the third moment is finite when . Question 16 yields the displayed joint Gaussian limit, and Theorem 23.11(2): the two limit coordinates are independent exactly when the covariance vanishes — for symmetric laws, empirical mean and empirical variance decouple asymptotically.
18. Write where for and : differentiability at means precisely as . Step 1: in probability: for and any , eventually , so (distribution functions converge at the continuity points ), and the right side tends to as . Step 2: in probability: given , pick with on ; then . Step 3:
By Slutsky’s product rule (Exercise 23.8, with the sequence in probability and the convergent-in-law ), the second term converges in law to , hence to in probability (Exercise 23.4(b)); the first converges in law to (Slutsky again, or the affine rule for characteristic functions); Slutsky’s sum rule assembles them: the limit is .
19. (a) The CLT gives ; the delta method with , , gives — degenerate (limit ) when . In that case the fluctuation lives one scale up: , and for
, the square of a Gaussian (a “chi-squared” law) — when the first derivative dies, the second-order term of Taylor dictates a non-Gaussian limit. (b) Here and has , so : the limit is for every . On the scale the asymptotic error bar is , known in advance — whereas in Example 23.9 the width involved the unknown , to be worst-cased by or estimated: the transformation stabilizes the variance.
20. Let and , so . For any : , with equality at ; and since the positive and negative parts of have equal total mass, . Exchanging handles the sign: . Coupling: for any ,
and take the supremum over .
21. The two laws charge: : versus , with ; : versus ; : versus the Poisson remainder . Hence
and gives the bound .
22. Write and with , independent of the pair (coalitions). For , conditioning on the countably many values by independence,
and likewise for with . Subtracting, with and of zero sum:
Telescoping from to (Exercise 23.1, iterated) and using question 21:
Le Cam’s inequality. (Question 20’s coupling bound gives an alternative route: couple each pair on one uniform variable so that and bound ; swapping needs no construction at all.)
23. (a) With : . This sharpens Exercise 23.5 three times over: an explicit error at every finite , uniformity over all events at once (not one interval at a time), and no need for equal — only small, e.g. : many rare events, none dominant. (b) Here , , : the Poisson model errs by at most on every event; in particular, taking ,
so the answer is up to a guaranteed (the true discrepancy is about ). (c) The problem closes on one method with two regimes. When comparable contributions each carry variance , matching two moments against the Gaussian makes the swap errors each: sums go Gaussian — with Taylor as the local comparison tool. When contributions are indicators of probability , matching the mean against a Poisson atom makes each swap cost : counts of rare events go Poisson — with total variation as the exact local comparison. Same hybrids, same telescope, different local estimate: replacement is a strategy, not a theorem, and the Gaussian and Poisson limits are its two oldest dividends.
24. The CLT gives , and is differentiable at with : the delta method (Part VI) yields . Unwinding the interval by exponentiation:
(in practice is replaced by its empirical version, Slutsky as in Exercise 23.8). The multiplicative interval is the natural one when the data are positive with errors proportional to their size — incomes, concentrations, half-lives: quantities that live on a log scale, where symmetric additive intervals could even cross zero.
25. . In the swapping analysis (Part IV), the leading error term after matching two moments carries the signed third moment: for it is positive (the law leans right: rare large excursions above the mean), and the normal approximation systematically misplaces mass — underestimating the short left tail and overestimating the right — with an error of order ; at the third moment vanishes, the Bernoulli matches the Gaussian to third order, and the rate improves (Part IV’s matched-moment question). Numerically: , while the Gaussian gives : the normal curve, ignorant of the wall at and of the rightward skew, puts too much mass at the bottom — the predicted sign of the error, visible at .