University Mathematics — Year 3 · Bachelor Year 3
11Product Measures, Fubini, Change of Variables
One-dimensional Lebesgue theory becomes multi-dimensional calculus through two theorems. Tonelli–Fubini says that integrals over products are iterated integrals — slicing is legitimate, in either order, under hypotheses one can actually check. The change of variables formula transports integrals along diffeomorphisms, with the Jacobian determinant as the exchange rate for volume; we prove it completely, starting from the linear case where it explains what the determinant is. Applications cascade: the layer cake formula, convolution, polar coordinates, the volume of the -ball — and, in the weekend problem, Stirling’s formula with an honest error analysis.
11.1 Product -algebras and product measures
Definition 11.1
For measurable spaces , , the product -algebra on is generated by the rectangles (, ) — a -system. For and , the section is ; for a function on the product, .
Proposition 11.2
(a) If , every section (and symmetrically); if is -measurable, every is -measurable. (b) .
Proof. (a) Good sets: is a -algebra (sections commute with complements and countable unions) containing the rectangles. For : . (b) () Rectangles of Borel sets: it suffices that openopen boxes are Borel in (they are open) and that general Borel rectangles are limits — good sets again: is a -algebra containing the opens; intersect two such. () Every open set of is a countable union of rational open boxes : contained in the product -algebra. ∎
Theorem 11.3 (Product measure)
Let and be -finite. For every , the function is measurable, and
defines the unique measure on with . It is -finite, and symmetric: the same measure is obtained by integrating -sections against .
Proof. Measurability of . First let be finite. The class of for which the map is measurable contains the rectangles () and is a -system: for in , (finiteness); for , (continuity from below), and monotone limits of measurable functions are measurable. Rectangles form a -system: Dynkin (Theorem 9.4) gives . If is -finite, write , , : with finite.
Measure. -additivity of follows from Corollary 10.7 (sections of disjoint sets are disjoint). On rectangles it gives . Uniqueness: two candidates agree on the -system of rectangles; -finiteness provides rectangles of finite measure: Theorem 9.7. Symmetry: the other-order construction is also a measure agreeing on rectangles — unique, hence the same. ∎
Definition 11.4
Lebesgue measure on is ( factors; associativity of the construction is checked on boxes and propagated by uniqueness). It is the unique Borel measure giving each box its volume ; it is translation-invariant (translates agree on boxes), -finite, and complete after Carathéodory completion — we write for the completed measure and integrate accordingly.
11.2 Tonelli and Fubini
Theorem 11.5 (Tonelli)
-finite, measurable. Then is measurable and
Proof. The standard machine. For this is Theorem 11.3 (and its symmetric form). By linearity it holds for simple . For general : take simple (Theorem 10.4); then for each (MCT in ), so the left members converge by MCT in , while by MCT on the product. ∎
Theorem 11.6 (Fubini)
-finite, . Then for -a.e. the section is -integrable, the a.e.-defined function is integrable, and the two iterated integrals both equal .
Proof. Tonelli applied to shows has finite integral, hence is finite a.e.: for a.e. . Split (real case; complex by components): Tonelli computes each iterated integral of as , and the a.e.-defined difference integrates to the difference. Symmetrically for the other order. ∎
Method 11.7
To interchange two integrals (or an integral and a sum, or two sums): if the integrand is nonnegative, interchange freely (Tonelli). Otherwise, first apply Tonelli to in whichever order is easier to estimate; if the result is finite, Fubini legitimizes the interchange. Never skip the check: Exercise 11.4’s integrand has two iterated integrals with different values.
Proposition 11.8 (Layer cake)
For measurable on -finite:
Proof. Apply Tonelli to on (measurable: it is , a Borel-type combination of the measurable ): integrating in first gives ; in first, . For : substitute in , i.e. apply the first formula to and change variables in the one-dimensional integral (). ∎
Theorem 11.9 (Convolution on )
For , the integral
converges absolutely for a.e. , defines with , and is commutative and associative.
Proof. is measurable ( is continuous; compose and multiply). Tonelli:
(translation invariance of in the inner integral). So the double integral is finite; Fubini gives a.e. absolute convergence and the norm bound . Commutativity: substitute (translation and reflection invariance — reflection invariance holds on boxes, hence everywhere by uniqueness). Associativity: Tonelli–Fubini on a triple integral. ∎
11.3 Change of variables
Theorem 11.10 (Linear change of variables)
For and : ; consequently for or integrable.
Proof. The measure is a Borel measure (homeomorphisms preserve Borel sets, Problem 9.1), translation-invariant (), finite on the unit box: by the characterization of Lebesgue measure (Exercise 9.6, whose proof works verbatim in with dyadic cubes), with . The map is multiplicative (, by composing), so it suffices to compute on generators of : elementary matrices. Diagonal : maps the unit cube to a box of volume : . Transposition of coordinates: permutes the cube: . Transvection : the image of the unit cube is a sheared prism; by Tonelli its measure is , each -section being an interval of length : . Every invertible matrix is a product of these (Gaussian elimination), and both and are multiplicative: . The integral formula follows by the standard machine (indicators, simple, MCT). ∎
Theorem 11.11 (Change of variables)
Let be open and a diffeomorphism. For every measurable (or ):
Proof. Write . The heart of the proof is the inequality
Step 4 below upgrades — applied to both and — to the equality of the theorem. Note that is itself a diffeomorphism with Jacobian (chain rule on ).
Step 1: for cubes with a distortion factor. Fix a closed cube of center and side (sup-norm ball). Claim: for every , if is differentiable on with on (operator norm for the sup-norm), then
because for , the mean value inequality applied to gives , and pulls this defect into an -enlargement of the cube. By Theorem 11.10,
Step 2: for compact cubes, by subdivision. Let be a compact cube and . On , is uniformly continuous and bounded (compactness); subdivide into subcubes small enough that the oscillation of on each is . Step 1 on each subcube (center ):
the last step because uniformly on (continuity of ) — Riemann-sum comparison. Let : holds for compact cubes.
Step 3: for all Borel . The set function , on Borel subsets of , is a measure ( is a bijection onto preserving Borel sets and countable disjointness), and so is . Every open subset of is a countable union of almost-disjoint dyadic compact cubes (standard dyadic decomposition: take maximal dyadic cubes contained in the open set), and both measures are additive across them (boundaries of cubes are -null, and their -images are null by Step 2 applied to thin cube-coverings of the faces): passes from cubes to open sets. General Borel : exhaust by compacts with , and fix ; on , is bounded by some . By outer regularity of (proof as in Theorem 9.13, with boxes), choose open sets with and . Then
Let : continuity from below on the left, MCT on the right. This establishes .
Step 4: equality and the integral formula. First extend from sets to integrals: for every measurable on ,
Indeed, for this is with ; linearity extends it to simple , and MCT to all (the standard machine). Now apply twice: first to , then — for the diffeomorphism — to the function :
since (chain rule on ). All inequalities are equalities: the formula holds for , and for functions by decomposition. ∎
Example 11.12 (Polar coordinates; the Gaussian again)
is a diffeomorphism from onto minus a half-line (null set), with :
For , Tonelli and this formula give
the classical two-line proof of , now fully justified (compare the parameter proof of Problem 10.1).
Theorem 11.13 (Volume of the unit ball)
Let in . Then
Proof. Compute twice. By Tonelli it factors: . By the layer cake formula (Proposition 11.8) with , whose level sets are balls: for , of measure (dilation by scales by : Theorem 11.10), so
Equate. (The values: , , etc.) Note as — the weekend problem quantifies how fast, via Stirling. ∎
11.4 Exercises
Exercise 11.1 ★
Let be counting measure on (not -finite) and Lebesgue measure, and let be the diagonal in . Show that is measurable, and compute the two iterated integrals of against and : they differ. Which hypothesis of Theorem 11.5 fails?
Solution
Solution of Exercise 11.1.
is closed in , hence Borel, and is the product -algebra (Proposition 11.2(b)). Iterating one way:
the other way:
The failing hypothesis is -finiteness of the counting measure on the uncountable : no countable family of finite- sets covers it.
Exercise 11.2 ★
Justify the interchange and re-derive Dirichlet’s integral: for ,
compute the inner integral in closed form, and let (dominate the -integral) to get .
Solution
Solution of Exercise 11.2.
On : (Tonelli for the absolute value): Fubini applies, and since ,
(the inner integral: , computed directly). As , the correction term is bounded by ; the main term is . Hence — Dirichlet’s integral by Fubini.
Exercise 11.3 ★★
(a) Prove that for measurable and finite: : integrability is summability of the tail measures. (b) Deduce that ( finite) iff .
Solution
Solution of Exercise 11.3.
(a) Layer cake (Proposition 11.8): , and is nonincreasing. On : ; summing the integrals over the unit intervals:
(b) Apply (a) to : finiteness of the integral and of the series are equivalent (the extra is finite).
Exercise 11.4 ★★
For on , show
(note ), and verify directly that : Fubini’s integrability hypothesis is not decorative.
Solution
Solution of Exercise 11.4.
Since for :
by the antisymmetry , the other order gives . Absolute values: for ,
No contradiction with Fubini: its hypothesis fails, and the two iterated integrals are simply two different numbers.
Exercise 11.5 ★★
(a) Compute explicitly (a tent function), and ’s general shape. (b) Show . (c) Show that if and is bounded and continuous, is continuous. (DCT via continuity of translation on the bounded .)
Solution
Solution of Exercise 11.5.
(a) : for , for , for : the tent. Convolving again gives a piecewise-quadratic bump on (the quadratic B-spline): each convolution gains one degree of smoothness — the smoothing principle behind Chapter 12’s mollifiers.
(b) If , there is a ball around disjoint from the sum set; for , , so the integrand vanishes identically: near .
(c) For : ; the integrands converge pointwise (continuity of ) and are dominated by : DCT gives .
Exercise 11.6 ★★
(a) Show that the simplex has volume (induction and Fubini). (b) Recover , from Theorem 11.13, and show .
Solution
Solution of Exercise 11.6.
(a) By Fubini and induction, slicing along the last coordinate:
using the dilation rule (Theorem 11.10); with : volume .
(b) ; . The ellipsoid is with : Theorem 11.10 gives volume .
Exercise 11.7 ★★
For which are the following finite? Justify with polar coordinates:
Generalize to (the thresholds and ).
Solution
Solution of Exercise 11.7.
In , polar coordinates (Example 11.12):
finite iff (), resp. (). In , avoid spherical coordinates with the layer cake: , and iff , i.e. ; the exterior integral converges iff (same computation on the complementary region).
Exercise 11.8 ★★★
(Beta–Gamma) For , let . Starting from as a double integral, substitute (a diffeomorphism of the open quadrant onto ; compute its Jacobian ) and conclude
Deduce and the value of the Wallis integrals .
Solution
Solution of Exercise 11.8.
By Tonelli (positive integrands) and the change of variables , a diffeomorphism of onto the open quadrant with
Substituting in gives . Wallis: — e.g. using .
Exercise 11.9 ★★
(Transfer formula) Let be measurable and the pushforward measure. Show that for every measurable on :
(standard machine). Then compare with Theorem 11.10: what extra information does the change of variables formula carry that the abstract transfer formula does not? (The transfer formula never identifies ; the change of variables theorem computes explicitly as a density measure.)
Solution
Solution of Exercise 11.9.
Indicators: ; linearity extends to simple , MCT to — the transfer formula. It is purely formal: it re-expresses integrals against but says nothing about what is. The content of Theorem 11.10 and Theorem 11.11 is the identification
i.e. a computation of the pushforward of Lebesgue measure — the analytic input being the differential geometry of , not measure-theoretic formalism.
Exercise 11.10 ★★★
(Gaussian moments) Using polar coordinates and Fubini, compute for the standard Gaussian weight on :
check the consistency (), and deduce the second moment of the measure .
Solution
Solution of Exercise 11.10.
By Tonelli the Gaussian factorizes, so with and (integrate by parts):
(by symmetry, contributes equal terms — the consistency check). For the normalized measure , the second moment is .
Exercise 11.11 ★★
(Graph and hypograph) Let be measurable. (a) Show that the hypograph is measurable in with
“the integral is the area under the graph”, at last a theorem. (Sections; Tonelli.) (b) Show that the graph is a null set of . (c) Deduce a two-line proof that the sphere is Lebesgue-null in .
Solution
Solution of Exercise 11.11.
(a) for intersected with : measurable, since and are measurable on the product (compositions with the projections). The -section of is , of measure : Tonelli integrates the sections,
(b) The graph is , measurable; its -sections are singletons, of measure : Tonelli gives .
(c) is the union of the two graphs over the unit ball of (splitting the last coordinate): a union of two null sets by (b), null.
Exercise 11.12 ★★
(A famous double integral) Using the geometric series and Tonelli on , prove
(The second series identity: split even and odd indices.) With (Chapter 15), two innocent-looking integrals evaluate to and ; where exactly does Tonelli’s positivity hypothesis do its work?
Solution
Solution of Exercise 11.12.
On , with nonnegative terms: Tonelli permits term-by-term integration,
For the alternating case, is not a positive series; but the integral of the absolute series is , so Fubini (integrability now established) applies: . The series identity:
With (Problem 15.1): the integrals are and . Tonelli’s positivity was the whole ballgame in the first computation — no integrability check needed before interchanging; in the second, positivity of the absolute series is what certifies integrability so that Fubini may run on the signed one.
11.5 Problem: Stirling’s formula
Problem 11.1
Weekend problem — , by dominated convergence
Stirling’s formula governs every asymptotic count in this book — ball volumes, binomial coefficients, the central limit theorem’s local form. We prove it from the integral (Example 10.16) with the Laplace method, in its cleanest dominated-convergence form, then collect dividends.
Part I — The formula. For , .
Substitute and show
where .
- Show the pointwise limit: for every fixed , as (expand to second order).
Domination. Let , so that for . Prove the two bounds
(study and : compute the derivatives and check the sign on each range). Deduce, for :
so that : an integrable dominator independent of .
Conclude with the dominated convergence theorem and the Gaussian integral (Example 11.12):
and in particular .
Part II — Dividends.
- (Wallis) From Exercise 11.8, -type formulas: derive from Stirling, and check it against the recursion .
(Ball volumes collapse) Show
so faster than any geometric sequence; find the dimension maximizing (numerically: ).
(Concentration of the binomial — a preview of Chapter 23) Using Stirling, show the local estimate, for with fixed and even:
the discrete Gaussian profile: de Moivre–Laplace in embryo.
- Where exactly did the proof of Part I use: (i) MCT or DCT; (ii) the Gaussian integral; (iii) the invariance properties of Lebesgue measure? One sentence each.
Part III — The error term: Stirling with bars. Set , so that Part I says .
- Show .
With , verify and expand:
and deduce the two-sided bounds
Telescope (using ) and check the pleasant algebraic identity for , to obtain the classical bracketing
- Two consequences: (a) the relative error of Stirling’s formula is as soon as ; (b) estimate to four significant digits by hand from the bracketing (), and marvel briefly at the precision of an asymptotic formula at a very finite .
Part IV — The Wallis route: Stirling without the Gaussian. Historically the constant came from Wallis, not from Gauss; this part re-proves Stirling independently of Parts I–II, and thereby re-proves the Gaussian integral. Let .
Establish (integrate by parts), the closed forms
and the identity .
From the monotonicity of deduce , then
Wallis’ theorem, obtained without Stirling.
- Show, by the telescoping of Part III alone (no value of the constant needed), that converges to some limit ; equivalently with not yet identified.
- Insert this asymptotic into and identify, using question 14, the only possible value: . Assemble the logic: Parts III–IV together give a complete second proof of Stirling — and hence, running Part I’s substitution backwards, an independent evaluation of . Two pillars, either of which supports the other.
Part V — Last dividends.
(The full local profile) For integers ( fixed), show
uniformly in (take logarithms and use ). This is the two-sided version of question 7 and the exact estimate quoted in Chapter 23’s weekend problem.
- (A Poisson preview) Show with Stirling that : the mode of a Poisson law of large mean carries mass , exactly as the central limit theorem will predict.
- (Gamma ratios) For , prove using the log-convexity slope bounds of Problem 10.1 (question 14 there), and extend to every real by the functional equation. (This is what “ Stirling” means between the integers.)
(Balls, encore) From : tabulate exactly, verify unimodality via (increasing while , decreasing after), and prove the striking generating identity
all even-dimensional unit-ball volumes packed into one exponential.
(Entropy asymptotics) For fixed with , deduce from Stirling
the exponential growth rate of binomial coefficients is the entropy — check that recovers question 5, and that for (so off-center binomials are exponentially negligible in ).
- (Surface areas) The area of the unit sphere is (proved as Exercise 21.6 in the differential-forms chapter; here, take it as the definition). Tabulate , locate the maximal one (, ), and show super-geometrically as well — high-dimensional spheres are, by every Euclidean yardstick, vanishingly small.
(The first correction term) Deduce from the bracketing of question 11 that , hence
Verify at : the bare formula gives (relative error ), the corrected one against (relative error ) — one term of the series buys two and a half digits.
(The median of ) Show that
asymptotically, exactly half of the mass of the integrand sits below its mode . (Run Part I’s substitution on the truncated integral; the dominator of question 3 is already in place.)
(Entropy, non-asymptotically) For prove the bound, valid for every :
by comparing the sum with for the tilt . Check that this choice of is optimal, and reconcile with question 21: the exponential rate of the asymptotic statement is attained by a one-line inequality with no asymptotics at all.
Solution
Solution of Problem 11.1.
1. With (; ranges over as ranges over ):
since and .
2. For fixed and : : .
3. Set on : and , which is on and on : , i.e. there. Set on , : and : for . Now for : if , ; if , then (as ), so . Hence , integrable, independent of .
4. DCT: (Example 11.12 plus the scaling ). With question 1:
5. From Exercise 11.8, . Stirling:
Then , consistent with the recursion (which forces , and with — the classical Wallis relation — gives ; the two asymptotics agree).
6. , so
super-geometrically (for , each factor and shrinking). Numerically , , , , , : the maximum is at .
7. With (integer, even, fixed): take logarithms in and apply Stirling to the three factorials. Writing , with :
and the bracket is : the display tends to up to , i.e.
the Gaussian profile of coin-tossing, quantified — de Moivre–Laplace’s local form, to be globalized in Chapter 23.
8. (i) DCT converts the pointwise limit of question 2 into convergence of the integrals, using question 3’s dominator. (ii) The Gaussian integral evaluates the limit — Stirling’s constant is the Gaussian integral. (iii) The substitution is an affine change of variables: translation invariance and the scaling rule of Lebesgue measure (Theorem 11.10 in dimension ).
9. Expand both terms:
the terms and combining into .
10. For : , and ; the odd series gives
Subtract . Lower bound: the first term alone, . Upper bound: lower all denominators to and sum the geometric series: .
11. Summing the upper bound from to (with ): . For the lower bound: , and
true for . Summing this telescoping minorant: . Exponentiating gives the classical bracketing of .
12. (a) Relative error for large; as soon as , and the stated suffices ( already implies it). (b) , so ; the guaranteed window has width under in relative terms — an “asymptotic” formula that is, at , an instrument of precision.
13. Write , and integrate the second term by parts (, , ):
Hence , i.e. . From , :
converting double factorials by and . Finally by the recursion: constant, equal to .
14. (pointwise monotonicity of ) and squeeze . Combined with (question 13): , so and .
15. Questions 9–10 never used the value of the constant: with , the differences lie in , so decreases while increases: adjacent sequences, converging to a common . Hence , .
16. Substituting the unknown-constant Stirling into the central binomial:
and question 14 forces : . Parts III–IV thus reprove Stirling from scratch; feeding it into Part I’s identity evaluates without polar coordinates: Wallis and Gauss prop each other up.
17. for (and by symmetry for ). Taking logarithms, with :
and , while the error sums to : uniformly, .
18. . A Poisson variable of mean has standard deviation , and is exactly the Gaussian peak height : the local CLT, previewed at the mode.
19. For , the slope lemma of Problem 10.1 (question 14 there), applied to the convex around , gives : the ratio to is squeezed by . For (, ): , and each of the factors is : multiply the estimates.
20. The recursion (from ) gives, from , :
The ratio exceeds exactly for , so each parity increases then decreases; numerically , , : the overall maximum is . Generating function: , so — all even-dimensional ball volumes rolled into one exponential, and an instant super-geometric decay estimate for .
21. Stirling in numerator and denominator, with :
since (the powers of cancel: ), and the ’s cancel likewise. At : and the prefactor is — question 5 again. Strict concavity of (its second derivative ) puts its maximum at only: for , decays exponentially — the combinatorial engine behind every concentration statement about coin flips.
22. From and question 20:
numerically ; and : the maximum is the -sphere. The recursion shows the same -driven rise and super-geometric fall as for volumes: past dimension seven, spheres shrink away faster than any geometric sequence.
23. Question 11 says exactly , and
so and ; multiplying by gives the corrected formula. At : , low by (relative error ); multiplying by gives , low by (relative error ). The bracketing itself pins between and — the upper bound is off by eight units in seven digits.
24. Part I’s substitution , applied to the truncated integral, gives
the range becoming . The dominator of question 3 covers as well, so dominated convergence yields
while question 4 gives . The ratio tends to . Probabilistically: a Gamma random variable of large shape puts asymptotically half its mass on each side of its mode — the central limit theorem’s symmetry, read off from one substitution.
25. Let . For we have , hence
and with :
Optimality: minimizing over , the equation has the unique solution , a minimum since — the exponential-tilting (Chernoff) choice. Reconciliation: by question 21 the single term is already of order , so
the rate is exact, the whole sum costing at most a factor over its largest term. Divided by , this is the fair-coin tail bound — concentration of measure in one line.