Mathematics · Book 4 · Bachelor Year 2

University Mathematics — Year 2

University Mathematics — Year 2 · Bachelor Year 2

8Functions of a Real Variable

Before analysis moves to functions of functions (Chapter 10) it pays to know the one-variable landscape in finer detail than Year 1 required: how discontinuous a monotone function can be, how regular a convex function must be, and what special properties derivatives enjoy (Darboux). These structural results are short, sharp, and beloved of examiners.

8.1 Monotone functions

Theorem 8.1 (Regularity of monotone functions)

Let f ⁣:IRf \colon I \to \R be increasing on an interval.

  1. At every interior point aa, the one-sided limits exist:

    f(a)=supx<af(x)    f(a)    f(a+)=infx>af(x);f(a^-) = \sup_{x < a} f(x) \;\leq\; f(a) \;\leq\; f(a^+) = \inf_{x > a} f(x) ;

    every discontinuity is a jump.

  2. The set of discontinuities of ff is at most countable.

Proof. (1) The set {f(x):x<a}\{f(x) : x < a\} is nonempty, bounded above by f(a)f(a): its supremum ss satisfies f(x)sf(x) \to s as xax \to a^- (given ε\varepsilon, some f(x0)>sεf(x_0) > s - \varepsilon, and monotonicity traps f(x)(sε,s]f(x) \in \intoc{s - \varepsilon}{s} for x(x0,a)x \in \intoo{x_0}{a}). Symmetrically on the right.

(2) To each discontinuity aa attach the nonempty open interval Ja=(f(a),f(a+))J_a = \intoo{f(a^-)}{f(a^+)} (a genuine jump). For a<ba < b discontinuities, JaJ_a and JbJ_b are disjoint: f(a+)f(c)f(b)f(a^+) \leq f(c) \leq f(b^-) for any cc between. Each JaJ_a contains a rational; distinct discontinuities get distinct rationals: an injection of the discontinuity set into Q\Q, which is countable (Proposition 1.6).

Example 8.2

The bound is sharp: fix an enumeration (rn)(r_n) of Q(0,1)\Q \cap \intoo{0}{1} and set f(x)=n:rnx2nf(x) = \sum_{n : r_n \leq x} 2^{-n} (a summable-family definition, Definition 7.8). Then ff is increasing on [0,1]\intcc{0}{1} and discontinuous exactly at every rational of (0,1)\intoo{0}{1} (jump 2n2^{-n} at rnr_n): a monotone function can be discontinuous on a dense countable set.

Example 8.3 (The jumps cannot outweigh the rise)

For ff increasing on [a,b]\intcc{a}{b}, the jumps have a budget: if a<c1<<cm<ba < c_1 < \dots < c_m < b are discontinuities with jumps si=f(ci+)f(ci)>0s_i = f(c_i^+) - f(c_i^-) > 0, then choosing interlacing points a<c1<t1<c2<a < c_1 < t_1 < c_2 < \dots and using monotonicity on each piece,

i=1msi    f(b)f(a):\sum_{i=1}^{m} s_i \;\leq\; f(b) - f(a) :

the total ascent bounds the total jumping. Consequence: for each kk, at most k(f(b)f(a))k\,\bigl(f(b) - f(a)\bigr) discontinuities have jump 1k\geq \frac1k — a quantitative refinement of Theorem 8.1 (2), since the discontinuity set is the countable union over kk of these finite sets. On the rational-jump function above, the budget is spent exactly: the jumps 2n2^{-n} sum to 1=f(1+)f(0)1 = f(1^+) - f(0^-) in the obvious extended sense. Monotone functions may jump densely, but only on a strict allowance.

8.2 Convex functions

Lemma 8.4 (Slope inequality)

Let ff be convex on II and x<y<zx < y < z in II. Then

f(y)f(x)yx    f(z)f(x)zx    f(z)f(y)zy:\frac{f(y) - f(x)}{y - x} \;\leq\; \frac{f(z) - f(x)}{z - x} \;\leq\; \frac{f(z) - f(y)}{z - y} :

slopes of chords increase in both endpoints.

Proof. Write y=zyzxx+yxzxzy = \frac{z - y}{z - x}\,x + \frac{y - x}{z - x}\,z: a convex combination, since the two coefficients are positive and sum to 11. Convexity gives

f(y)    zyzxf(x)+yxzxf(z).f(y) \;\leq\; \frac{z-y}{z-x}\,f(x) + \frac{y-x}{z-x}\,f(z).

For the left inequality, subtract f(x)f(x) from both sides, using zyzx1=yxzx\frac{z-y}{z-x} - 1 = -\frac{y-x}{z-x}:

f(y)f(x)yxzx(f(z)f(x)),f(y) - f(x) \leq \frac{y - x}{z - x}\bigl(f(z) - f(x)\bigr),

and divide by yx>0y - x > 0. For the right inequality, subtract instead from f(z)f(z):

f(z)f(y)f(z)zyzxf(x)yxzxf(z)=zyzx(f(z)f(x)),f(z) - f(y) \geq f(z) - \frac{z-y}{z-x}f(x) - \frac{y-x}{z-x}f(z) = \frac{z - y}{z - x}\bigl(f(z) - f(x)\bigr),

and divide by zy>0z - y > 0. Both displayed steps are the same barycentric identity read against a different endpoint.

Theorem 8.5 (Regularity of convex functions)

Let ff be convex on an interval II.

  1. At every interior point, ff has finite one-sided derivatives fgfdf'_g \leq f'_d; both are increasing functions of the point; in particular ff is continuous on the interior of II (but possibly not at endpoints).
  2. ff lies above each support line: for aa interior and any m[fg(a),fd(a)]m \in \intcc{f'_g(a)}{f'_d(a)},

    f(x)f(a)+m(xa)(xI).f(x) \geq f(a) + m(x - a) \qquad (x \in I).
  3. (Jensen, weighted) For xiIx_i \in I and weights λi0\lambda_i \geq 0, λi=1\sum\lambda_i = 1:

    f(iλixi)iλif(xi).f\Bigl(\sum_i \lambda_i x_i\Bigr) \leq \sum_i \lambda_i f(x_i) .

Proof. (1) Fix aa interior. By Lemma 8.4, the slope τ(h)=f(a+h)f(a)h\tau(h) = \frac{f(a + h) - f(a)}{h} is an increasing function of hh (on both sides, and τ(h)τ(h+)\tau(h_-) \leq \tau(h_+) for h<0<h+h_- < 0 < h_+). Hence τ\tau has a finite limit as h0h \to 0^- (increasing, bounded above by any right slope) — this is fg(a)f'_g(a) — and as h0+h \to 0^+ (fd(a)f'_d(a)), with fg(a)fd(a)f'_g(a) \leq f'_d(a). Finite one-sided derivatives force continuity at aa. Monotonicity in the point: for a<ba < b interior, fd(a)f(b)f(a)bafg(b)f'_d(a) \leq \frac{f(b) - f(a)}{b - a} \leq f'_g(b), again by the slope inequality.

(2) For x>ax > a: f(x)f(a)xafd(a)m\frac{f(x) - f(a)}{x - a} \geq f'_d(a) \geq m; for x<ax < a: f(a)f(x)axfg(a)m\frac{f(a) - f(x)}{a - x} \leq f'_g(a) \leq m. Both rearrange to the claim.

(3) Induction on the number of points exactly as in the Year 1 volume (the two-point case is the definition) — or in one stroke: apply (2) at a=λixia = \sum\lambda_i x_i and average the support-line inequalities at the points xix_i with weights λi\lambda_i: iλif(xi)f(a)+miλi(xia)=f(a)\sum_i \lambda_i f(x_i) \geq f(a) + m\sum_i\lambda_i(x_i - a) = f(a).

Convexity in one picture: between -1.5 and 2 the graph of f(x) = x2 stays below its chord (the definition) and above the support line at x = 0.5 ( (2)) — every inequality of this chapter’s weekend problem is a rearrangement of these two positions.
Convexity in one picture: between 1.5-1.5 and 22 the graph of f(x)=x2f(x) = x^2 stays below its chord (the definition) and above the support line at x=0.5x = 0.5 (Theorem 8.5 (2)) — every inequality of this chapter’s weekend problem is a rearrangement of these two positions.

Example 8.6 (Endpoint discontinuity)

On [0,1]\intcc{0}{1}, the function f(0)=1f(0) = 1, f(x)=0f(x) = 0 for x>0x > 0 is convex but discontinuous at the endpoint 00: statement (1) is sharp.

Example 8.7 (Corners and the sheaf of support lines)

For f(x)=xf(x) = \abs x at a=0a = 0: the one-sided derivatives are fg(0)=1f'_g(0) = -1 and fd(0)=+1f'_d(0) = +1, and Theorem 8.5 (2) hands out a support line for every slope m[1,1]m \in \intcc{-1}{1}:

xmx(xR, 1m1),\abs x \geq m\,x \qquad (x \in \R,\ -1 \leq m \leq 1),

each an equality exactly on a half-line or at 00. A convex function is differentiable at aa precisely when the sheaf collapses to a single line (fg(a)=fd(a)f'_g(a) = f'_d(a)); corners carry an interval of tangents. This sheaf is the finite-dimensional germ of the subdifferential of convex optimization — and the reason convex functions are so robust: even where the derivative fails, the supporting geometry survives, which is all that Jensen’s proof used.

Example 8.8 (Power mean inequality)

For 0<p<q0 < p < q and positive xix_i with weights λi\lambda_i summing to 11, applying Jensen to the convex ttq/pt \mapsto t^{q/p} at the points xipx_i^p:

(λixip)1/p(λixiq)1/q:\Bigl(\sum \lambda_i x_i^{p}\Bigr)^{1/p} \leq \Bigl(\sum \lambda_i x_i^{q}\Bigr)^{1/q} :

power means increase with the exponent — containing AM–QM, and, in the limit p0p \to 0 (Exercise 8.6), the AM–GM inequality once more.

The power mean M_p of the values 1, 2, 4 (equal weights), as a function of the exponent p: increasing from = 1 (as p -∈fty) to = 4 (as p +∈fty), through the harmonic (p = -1), geometric (the gap at p = 0, value 2), arithmetic (p = 1) and quadratic (p = 2) means. The whole chain of classical mean inequalities is one increasing curve — proved in this chapter’s weekend problem, Part III.
The power mean MpM_p of the values 1,2,41, 2, 4 (equal weights), as a function of the exponent pp: increasing from min=1\min = 1 (as pp \to -\infty) to max=4\max = 4 (as p+p \to +\infty), through the harmonic (p=1p = -1), geometric (the gap at p=0p = 0, value 22), arithmetic (p=1p = 1) and quadratic (p=2p = 2) means. The whole chain of classical mean inequalities is one increasing curve — proved in this chapter’s weekend problem, Part III.

Example 8.9 (Maximal entropy)

For a probability vector (p1,,pn)(p_1, \dots, p_n) (positive, summing to 11), the entropy H(p)=ipilnpiH(p) = -\sum_i p_i\ln p_i satisfies

H(p)lnn,with equality iff pi=1n for all i.H(p) \leq \ln n , \qquad\text{with equality iff } p_i = \frac1n \text{ for all } i .

Proof by Jensen (Theorem 8.5 (3)) applied to the concave ln\ln with weights pip_i at the points 1pi\frac{1}{p_i}:

H(p)=ipiln1piln(ipi1pi)=lnn,H(p) = \sum_i p_i\ln\frac{1}{p_i} \leq \ln\Bigl(\sum_i p_i\,\frac1{p_i}\Bigr) = \ln n ,

equality forcing all points 1pi\frac1{p_i} equal (strict concavity), i.e. pp uniform. Equivalently, this is Exercise 8.7 with qq uniform. Uncertainty is maximized by ignorance uniformly spread — the variational principle behind coding, statistical mechanics, and the entropy appearances of Chapter 22.

Method 8.10 (Finding the convex function behind an inequality)

Most classical inequalities are Jensen in costume; to undress one: (1) normalize so that a weighted average appears (weights positive, summing to 11 — divide by a total mass if necessary); (2) look at what function is applied inside versus outside the average: the claim “f(average)f(\text{average}) \leq average of ff” names the convex ff; (3) certify convexity by the second derivative, and handle equality via strictness; (4) if no average is visible, take logarithms first — products and powers become averages, and the concavity of ln\ln carries AM–GM, Young and their relatives (this chapter’s weekend problem runs steps 1–4 on each of them). If even logarithms do not reveal an average, try reading the inequality as monotonicity of slopes (Lemma 8.4) — superadditivity statements like Exercise 8.9 live there.

Remark 8.11 (Common pitfalls)

(i) Convexity is not preserved by products: xx and (x1)2(x - 1)^2 are convex on [0,2]\intcc{0}{2}, but their product x(x1)2x(x-1)^2 has second derivative 6x46x - 4, negative on [0,23)\intco{0}{\frac23} — not convex; nor is convexity preserved by composition without monotonicity (Exercise 8.10). (ii) Jensen flips for concave functions: half the classical inequalities are the concave ln\ln-version; applying the convex form to ln\ln is the quickest way to prove AM–GM backwards. (iii) Midpoint convexity alone does not imply convexity — continuity (or mere boundedness) is needed (Exercise 8.8); the pathological counterexamples live beyond this book’s axioms. (iv) A convex function on an open interval is continuous, even locally Lipschitz (Exercise 8.12); at endpoints, nothing is free. (v) Derivatives obey Darboux but need not be continuous (Example 8.15): “ff' has no jumps” never means “ff' is continuous”.

8.3 The Darboux property

Theorem 8.12 (Darboux)

Let ff be differentiable on an interval II. Then ff' takes every value between any two of its values — even though ff' need not be continuous.

Proof. Let a<ba < b in II and vv strictly between f(a)f'(a) and f(b)f'(b), say f(a)<v<f(b)f'(a) < v < f'(b). The function g(x)=f(x)vxg(x) = f(x) - vx is differentiable with g(a)<0<g(b)g'(a) < 0 < g'(b): its minimum on [a,b]\intcc{a}{b} (attained: continuity on a compact) is not at aa (just after aa, gg decreases below g(a)g(a)) nor at bb (just before bb, gg is below g(b)g(b)): it is interior, and there g(c)=0g'(c) = 0, i.e. f(c)=vf'(c) = v. (This was a Year 1 starred exercise; its place in the theory is here.)

Example 8.13 (Which functions are derivatives?)

Darboux’s theorem is a non-existence machine. The floor function x\lfloor x\rfloor is not the derivative of any function on R\R: it takes the values 00 and 11 but skips 12\frac12 on [0,1]\intcc{0}{1}, which Theorem 8.12 forbids for derivatives. The same verdict hits every function with a jump — sign, Heaviside, all step functions — however innocent they look; their “antiderivatives” (x\abs x for sign, etc.) exist only away from the jump and knot there with a corner. Contrast: the wildly discontinuous ff' of Example 8.15 is a derivative — its discontinuity is an oscillation, which Darboux tolerates. The boundary between the two behaviours is exactly the no-jumps corollary below.

Corollary 8.14

A derivative has no jump discontinuities: if f(a)f'(a^-) and f(a+)f'(a^+) exist, they equal f(a)f'(a). The discontinuities of a derivative are always of oscillation type (x2sin1xx^2\sin\frac1x’s derivative at 00, Year 1 volume).

Proof. If f(a+)=limxa+f(x)f'(a^+) = \lim_{x\to a^+} f'(x) exists and differs from f(a)f'(a), values strictly between them would be skipped by ff' on a right neighborhood — contradicting Darboux on intervals [a,a+h]\intcc{a}{a + h}. (Alternatively: the mean value theorem forces f(a)=limh0+f(a+h)f(a)h=f(a+)f'(a) = \lim_{h\to0^+} \frac{f(a+h)-f(a)}{h} = f'(a^+), the difference quotient being an ff'-value at an intermediate point.) Same on the left.

Example 8.15 (The canonical oscillating derivative)

Let f(x)=x2sin1xf(x) = x^2\sin\frac1x for x0x \neq 0 and f(0)=0f(0) = 0. At 00: f(h)f(0)h=hsin1hh0\bigl|\frac{f(h) - f(0)}{h}\bigr| = \abs{h\sin\frac1h} \leq \abs h \to 0, so f(0)=0f'(0) = 0 exists. Away from 00,

f(x)=2xsin1xcos1x,f'(x) = 2x\sin\frac1x - \cos\frac1x ,

whose first term tends to 00 while cos1x\cos\frac1x oscillates through [1,1]\intcc{-1}{1} on every interval (0,δ)\intoo{0}{\delta}: the limit f(0+)f'(0^+) does not exist. So ff' is defined everywhere but discontinuous at 00 — and, exactly as Corollary 8.14 predicts, the discontinuity is an oscillation, not a jump: on each (0,δ)\intoo{0}{\delta}, ff' still sweeps a full interval around 00. Derivatives can be wild, but only in the Darboux-compatible way.

Remark 8.16 (Where this chapter is used)

Convexity is the engine of the inequality industry: this chapter’s weekend problem manufactures Young, Hölder, Minkowski and the power-mean chain from it, which Chapter 5’s norm theory and Chapter 9’s integral estimates consume; Jensen reappears in probability as the moment inequalities of Chapter 22. Monotone regularity returns in Chapter 9 (monotone functions are integrable) and, in the Year 3 volume, as the almost-everywhere differentiability of monotone functions — where “countably many jumps” becomes the first step of Lebesgue’s theory.

8.4 Exercises

Exercise 8.1

Determine the discontinuity sets and the jump sizes: x\lfloor x \rfloor;   xx\;x - \lfloor x\rfloor;   x+xx\;\lfloor x \rfloor + \sqrt{x - \lfloor x\rfloor}; the function of the example following Theorem 8.1 restricted to dyadic rationals rnr_n.

Solution

Solution of Exercise 8.1.

x\lfloor x\rfloor: jumps of size 11 at every integer. xxx - \lfloor x\rfloor: jumps of size 1-1 at integers (left limit 11, value 00). x+xx\lfloor x\rfloor + \sqrt{x - \lfloor x\rfloor}: at an integer nn, left limit (n1)+1=n(n - 1) + 1 = n and value nn: continuous everywhere (the square root repairs the jump), though not differentiable at integers. The rational-jump function: restricting the construction to an enumeration of the dyadics, it jumps by 2n2^{-n} exactly at the nn-th dyadic rational and is continuous elsewhere.

Exercise 8.2

Prove that an increasing function f ⁣:IRf \colon I \to \R with the intermediate value property (its image of any subinterval is an interval) is continuous.

Solution

Solution of Exercise 8.2.

Suppose ff increasing has a discontinuity at an interior aa: then f(a)<f(a+)f(a^-) < f(a^+) (Theorem 8.1) and the image of II misses the nonempty open interval (f(a),f(a+))\intoo{f(a^-)}{f(a^+)} except possibly the single value f(a)f(a): the image of any subinterval containing aa in its interior is not an interval (it has a gap on at least one side of f(a)f(a)). This contradicts the intermediate value property. Endpoint discontinuities are excluded the same way with one-sided gaps.

Exercise 8.3

Which of the following are convex on their domain? xxlnxx \mapsto x\ln x (x>0x > 0);   xln(1+ex)\;x \mapsto \ln(1 + \eu^x);   x1+x2\;x \mapsto \sqrt{1 + x^2};   xx3\;x \mapsto x^3.

Solution

Solution of Exercise 8.3.

xlnxx\ln x: second derivative 1x>0\frac1x > 0: convex. ln(1+ex)\ln(1 + \eu^x): derivative ex1+ex=111+ex\frac{\eu^x}{1 + \eu^x} = 1 - \frac{1}{1 + \eu^x}, increasing: convex. 1+x2\sqrt{1 + x^2}: second derivative (1+x2)3/2>0(1 + x^2)^{-3/2} > 0: convex. x3x^3: not convex on R\R (f=6xf'' = 6x changes sign); convex only on R+\R_+.

Exercise 8.4 ★★

Let ff be convex on R\R and bounded above. Prove that ff is constant. (If f(a)f(b)f(a) \neq f(b), the slope inequality propagates the nonzero chord slope: beyond the point with the larger value, ff grows at least linearly — contradicting boundedness. Treat both signs of the slope.) Deduce that a convex function on R\R with an asymptote at both ends is affine.

Solution

Solution of Exercise 8.4.

Suppose f(a)f(b)f(a) \neq f(b), say f(b)>f(a)f(b) > f(a) with a<ba < b (the case f(b)<f(a)f(b) < f(a) is symmetric, looking left). For x>bx > b, the slope inequality (Lemma 8.4) on a<b<xa < b < x gives

f(x)f(a)xaf(b)f(a)ba=m>0f(x)f(a)+m(xa)x++,\frac{f(x) - f(a)}{x - a} \geq \frac{f(b) - f(a)}{b - a} = m > 0 \quad\Longrightarrow\quad f(x) \geq f(a) + m(x - a) \xrightarrow[x\to+\infty]{} +\infty,

contradicting boundedness above. Hence ff is constant.

Asymptotes: if f(x)(αx+β)0f(x) - (\alpha x + \beta) \to 0 at ++\infty and f(x)(αx+β)0f(x) - (\alpha' x + \beta') \to 0 at -\infty, the convex function g(x)=f(x)(αx+β)g(x) = f(x) - (\alpha x + \beta) is bounded above near ++\infty; convexity plus an asymptote at -\infty (which forces αα\alpha' \leq \alpha then α=α\alpha' = \alpha by comparing slopes at \mp\infty: slopes of a convex function increase) makes gg bounded above on all of R\R, hence constant =0= 0 in the limit: ff is affine.

Exercise 8.5 ★★

Let ff be differentiable on II with ff' monotone. Prove that ff' is continuous (combine Theorem 8.1 and Corollary 8.14).

Solution

Solution of Exercise 8.5.

ff' is monotone, so by Theorem 8.1 its only possible discontinuities are jumps, with one-sided limits existing everywhere. By Corollary 8.14, a derivative has no jump discontinuities. Hence ff' has no discontinuities at all: continuous.

Exercise 8.6 ★★

(Geometric mean as a limit) For positive xix_i and weights λi\lambda_i summing to 11, prove

limp0+(iλixip)1/p=ixiλi,\lim_{p \to 0^+} \Bigl(\sum_i \lambda_i x_i^p\Bigr)^{1/p} = \prod_i x_i^{\lambda_i} ,

via xip=eplnxi=1+plnxi+O(p2)x_i^p = \eu^{p\ln x_i} = 1 + p\ln x_i + O(p^2), and deduce the weighted AM–GM inequality from Example 8.8.

Solution

Solution of Exercise 8.6.

Take logarithms:

1pln(iλixip)=1pln(1+piλilnxi+O(p2))=iλilnxi+O(p)p0+iλilnxi,\frac1p \ln\Bigl(\sum_i \lambda_i x_i^p\Bigr) = \frac1p \ln\Bigl(1 + p\sum_i \lambda_i \ln x_i + O(p^2)\Bigr) = \sum_i \lambda_i \ln x_i + O(p) \xrightarrow[p \to 0^+]{} \sum_i \lambda_i \ln x_i ,

using λi=1\sum\lambda_i = 1 and ln(1+u)=u+O(u2)\ln(1 + u) = u + O(u^2). Exponentiating gives the geometric mean. Now for every p(0,1)p \in \intoo{0}{1}, the power-mean inequality (Example 8.8, exponents p<1p < 1) gives

(iλixip)1/piλixi;\Bigl(\sum_i \lambda_i x_i^p\Bigr)^{1/p} \leq \sum_i\lambda_i x_i ;

letting p0+p \to 0^+ on the left yields ixiλiiλixi\prod_i x_i^{\lambda_i} \leq \sum_i \lambda_i x_i: the weighted AM–GM inequality.

Exercise 8.7 ★★

(Entropy inequality) Using strict convexity of ttlntt \mapsto t\ln t, prove that for positive pi,qip_i, q_i with pi=qi=1\sum p_i = \sum q_i = 1:

ipilnpiqi0,\sum_i p_i \ln\frac{p_i}{q_i} \geq 0 ,

with equality iff p=qp = q. (Write the left side as qiφ(piqi)\sum q_i\, \varphi\bigl(\frac{p_i}{q_i}\bigr) with φ(t)=tlnt\varphi(t) = t\ln t and apply Jensen with weights qiq_i.)

Solution

Solution of Exercise 8.7.

With φ(t)=tlnt\varphi(t) = t\ln t (convex: φ=1t>0\varphi'' = \frac1t > 0) and weights qiq_i at the points ti=piqit_i = \frac{p_i}{q_i}:

ipilnpiqi=iqiφ(piqi)    φ(iqipiqi)=φ(1)=0,\sum_i p_i \ln\frac{p_i}{q_i} = \sum_i q_i\, \varphi\Bigl(\frac{p_i}{q_i}\Bigr) \;\geq\; \varphi\Bigl(\sum_i q_i \frac{p_i}{q_i}\Bigr) = \varphi(1) = 0 ,

by Jensen (Theorem 8.5 (3)). Equality in Jensen for a strictly convex function forces all the points tit_i to coincide: piqi\frac{p_i}{q_i} constant, and summing, the constant is 11: p=qp = q. (This quantity — the Kullback–Leibler divergence — returns in Chapter 22’s world.)

Exercise 8.8 ★★★

(Midpoint convexity) f ⁣:IRf \colon I \to \R is midpoint convex when f(x+y2)f(x)+f(y)2f\bigl(\frac{x+y}{2}\bigr) \leq \frac{f(x) + f(y)}{2} always. Prove that a continuous midpoint convex function is convex. (Establish the convexity inequality for dyadic weights k2m\frac{k}{2^m} by induction on mm, then pass to the limit using density and continuity.)

Solution

Solution of Exercise 8.8.

Dyadic weights. By induction on mm: the case m=1m = 1 is the hypothesis. For weight λ=k2m+1\lambda = \frac{k}{2^{m+1}} (odd kk), write λ=12(λ1+λ2)\lambda = \frac12(\lambda_1 + \lambda_2) with λj=k12m+1\lambda_j = \frac{k \mp 1}{2^{m+1}}, both of denominator 2m2^m after simplification; then

f(λx+(1λ)y)=f(u+v2)f(u)+f(v)2λf(x)+(1λ)f(y),f\bigl(\lambda x + (1{-}\lambda)y\bigr) = f\Bigl(\tfrac{u + v}{2}\Bigr) \leq \frac{f(u) + f(v)}{2} \leq \lambda f(x) + (1 - \lambda) f(y),

where u=λ1x+(1λ1)yu = \lambda_1 x + (1 - \lambda_1)y and v=λ2x+(1λ2)yv = \lambda_2 x + (1-\lambda_2)y, using midpoint convexity then the induction hypothesis on u,vu, v.

Passage to the limit. For arbitrary λ[0,1]\lambda \in \intcc{0}{1}, take dyadics λnλ\lambda_n \to \lambda: continuity of ff and of the affine maps passes the inequality f(λnx+(1λn)y)λnf(x)+(1λn)f(y)f(\lambda_n x + (1-\lambda_n)y) \leq \lambda_n f(x) + (1-\lambda_n)f(y) to the limit: ff is convex.

Exercise 8.9 ★★★

Let ff be convex on [0,+)\intco{0}{+\infty} with f(0)0f(0) \leq 0. Prove that xf(x)xx \mapsto \frac{f(x)}{x} is increasing on (0,+)\intoo{0}{+\infty}, and deduce that for convex ff with f(0)=0f(0) = 0: f(x+y)f(x)+f(y)f(x + y) \geq f(x) + f(y) for x,y0x, y \geq 0 (superadditivity).

Solution

Solution of Exercise 8.9.

For 0<x<y0 < x < y: the slope inequality (Lemma 8.4) at the points 0<x<y0 < x < y gives

f(x)f(0)xf(y)f(0)y,i.e.f(x)xf(y)y+f(0)(1x1y).\frac{f(x) - f(0)}{x} \leq \frac{f(y) - f(0)}{y}, \qquad\text{i.e.}\qquad \frac{f(x)}{x} \leq \frac{f(y)}{y} + f(0)\Bigl(\frac1x - \frac1y\Bigr).

Since f(0)0f(0) \leq 0 and 1x1y>0\frac1x - \frac1y > 0, the last term is 0\leq 0: f(x)xf(y)y\frac{f(x)}{x} \leq \frac{f(y)}{y}. So xf(x)xx \mapsto \frac{f(x)}x increases.

Superadditivity for f(0)=0f(0) = 0: for x,y>0x, y > 0 (the cases with a zero variable are trivial),

f(x)=xf(x)xxf(x+y)x+y,f(y)yf(x+y)x+y,f(x) = x\,\frac{f(x)}{x} \leq x\,\frac{f(x+y)}{x+y}, \qquad f(y) \leq y\,\frac{f(x+y)}{x+y},

by the monotonicity just proved; adding gives f(x)+f(y)f(x+y)f(x) + f(y) \leq f(x+y).

Exercise 8.10

Let ff be convex on II and gg convex increasing on an interval containing f(I)f(I). Prove that gfg \circ f is convex, and show by a counterexample that monotonicity of gg cannot be dropped.

Solution

Solution of Exercise 8.10.

For x,yIx, y \in I and λ[0,1]\lambda \in \intcc01: convexity of ff, then monotonicity of gg, then convexity of gg:

g(f(λx+(1λ)y))g(λf(x)+(1λ)f(y))λg(f(x))+(1λ)g(f(y)).g\bigl(f(\lambda x + (1{-}\lambda)y)\bigr) \leq g\bigl(\lambda f(x) + (1{-}\lambda)f(y)\bigr) \leq \lambda\,g(f(x)) + (1{-}\lambda)\,g(f(y)).

Counterexample without monotonicity: g(t)=tg(t) = -t is convex (affine) but decreasing, f(x)=x2f(x) = x^2 is convex, and gf=x2g \circ f = -x^2 is strictly concave.

Exercise 8.11 ★★

(Hermite–Hadamard) Let ff be convex and continuous on [a,b]\intcc{a}{b}. Prove

f(a+b2)    1baabf(t) ⁣dt    f(a)+f(b)2.f\Bigl(\frac{a+b}{2}\Bigr) \;\leq\; \frac{1}{b - a}\int_a^b f(t)\,\dd t \;\leq\; \frac{f(a) + f(b)}{2} .

(Left: integrate a support line at the midpoint. Right: bound ff by the chord.)

Solution

Solution of Exercise 8.11.

Left inequality: let m=a+b2m = \frac{a+b}2 and take a support line at mm (Theorem 8.5 (2)): f(t)f(m)+μ(tm)f(t) \geq f(m) + \mu(t - m) for all t[a,b]t \in \intcc ab. Integrating over [a,b]\intcc{a}{b}: the linear term integrates to μab(tm) ⁣dt=0\mu\int_a^b(t - m)\dd t = 0 (symmetry around mm), so abf(ba)f(m)\int_a^b f \geq (b - a)f(m).

Right inequality: on [a,b]\intcc ab, convexity bounds ff by its chord: f(t)f(a)+f(b)f(a)ba(ta)f(t) \leq f(a) + \frac{f(b) - f(a)}{b - a}(t - a). Integrating: abf(ba)f(a)+f(b)f(a)ba(ba)22=(ba)f(a)+f(b)2\int_a^b f \leq (b-a)f(a) + \frac{f(b) - f(a)}{b - a}\cdot\frac{(b-a)^2}2 = (b - a)\,\frac{f(a) + f(b)}2. Divide by bab - a.

Exercise 8.12 ★★★

Prove that a convex function on an open interval II is locally Lipschitz: for every segment [a,b]I\intcc{a}{b} \subseteq I and margin δ>0\delta > 0 with [aδ,b+δ]I\intcc{a - \delta}{b + \delta} \subseteq I, the restriction of ff to [a,b]\intcc{a}{b} is Lipschitz, with constant max(f(a)f(aδ)δ,f(b+δ)f(b)δ)\max\Bigl(\bigl|\frac{f(a) - f(a - \delta)}{\delta}\bigr|, \bigl|\frac{f(b + \delta) - f(b)}{\delta}\bigr|\Bigr) (trap every chord slope between these two by the slope inequality).

Solution

Solution of Exercise 8.12.

Let aδ<ax<yb<b+δa - \delta < a \leq x < y \leq b < b + \delta, all in II. Two applications of the slope inequality (Lemma 8.4), first to aδ<ax<ya - \delta < a \leq x < y, then to x<yb<b+δx < y \leq b < b + \delta:

f(a)f(aδ)δf(y)f(x)yxf(b+δ)f(b)δ\frac{f(a) - f(a - \delta)}{\delta} \leq \frac{f(y) - f(x)}{y - x} \leq \frac{f(b + \delta) - f(b)}{\delta}

(chord slopes increase when both endpoints move right). Hence every chord slope inside [a,b]\intcc ab is trapped between two fixed numbers, and

f(y)f(x)Kyx,K=max(f(a)f(aδ)δ,f(b+δ)f(b)δ):\abs{f(y) - f(x)} \leq K\,\abs{y - x}, \qquad K = \max\Bigl(\Bigl|\frac{f(a) - f(a-\delta)}{\delta}\Bigr|, \Bigl|\frac{f(b+\delta) - f(b)}{\delta}\Bigr|\Bigr):

ff is Lipschitz on [a,b]\intcc ab. Every point of the open II has such a segment-with-margin around it: locally Lipschitz, hence (again) continuous on II.

8.5 Problem: The Convexity Toolbox

One definition — the chord above the graph — generates the entire toolbox of classical inequalities. This weekend problem builds it in logical order: convexity criteria and strict Jensen, then Young, Hölder and Minkowski (the birth certificates of the pp-norms), the complete power-mean chain from the minimum to the maximum, and two crown dividends — Carleman’s inequality, and Hölder read as a duality. Everything is proved; nothing is imported.

Problem 8.1

Weekend problem — Young, Hölder, Minkowski, and the power-mean chain

Throughout, p,q>1p, q > 1 are conjugate exponents: 1p+1q=1\frac1p + \frac1q = 1; vectors are a=(a1,,an)Rna = (a_1, \dots, a_n) \in \R^n; weights λi>0\lambda_i > 0 satisfy iλi=1\sum_i\lambda_i = 1.

Part I — Criteria and strict Jensen.

  1. Let ff be differentiable on an interval II. Prove that ff is convex if and only if ff' is increasing (one direction by passing to the limit in the slope inequality Lemma 8.4; the other by the mean value theorem). Deduce the C2C^2 criterion f0f'' \geq 0.
  2. Suppose f>0f''> 0 on II. Prove that ff is strictly convex (strict inequality for xyx \neq y and λ(0,1)\lambda \in \intoo01), and that a strictly convex function satisfies Jensen’s inequality (Theorem 8.5 (3)) with equality only when all the xix_i coincide.
  3. Certify the toolbox’s raw materials: ln-\ln is strictly convex on (0,+)\intoo{0}{+\infty}; ttrt \mapsto t^r is strictly convex there for r>1r > 1 and strictly concave for 0<r<10 < r < 1; exp\exp is strictly convex on R\R.
  4. (Young’s inequality) For a,b0a, b \geq 0, prove

    ab    app+bqq,ab \;\leq\; \frac{a^p}{p} + \frac{b^q}{q},

    with equality if and only if ap=bqa^p = b^q (apply the concavity of ln\ln to the two points ap,bqa^p, b^q with weights 1p,1q\frac1p, \frac1q).

  5. Re-derive weighted AM–GM in one line from the concavity of ln\ln:

    ixiλiiλixi(xi>0),\prod_i x_i^{\lambda_i} \leq \sum_i\lambda_ix_i \qquad (x_i > 0),

    with the equality case; compare with the limit route of Exercise 8.6.

Part II — Hölder and Minkowski. Write ap=(iaip)1/p\norm{a}_p = \bigl(\sum_i \abs{a_i}^p\bigr)^{1/p} and a=maxiai\norm{a}_\infty = \max_i\abs{a_i}.

  1. (Hölder) Prove

    i=1naibi    apbq,\sum_{i=1}^{n}\abs{a_ib_i} \;\leq\; \norm a_p\,\norm b_q ,

    with equality iff the vectors (aip)(\abs{a_i}^p) and (biq)(\abs{b_i}^q) are proportional (normalize ap=bq=1\norm a_p = \norm b_q = 1 and apply Young termwise).

  2. Identify the special cases: p=q=2p = q = 2 (Cauchy–Schwarz), and the endpoint pair (p,q)=(1,)(p, q) = (1, \infty): state and prove aibia1b\sum\abs{a_ib_i} \leq \norm a_1\norm b_\infty.
  3. (Minkowski) For p1p \geq 1, prove

    a+bpap+bp\norm{a + b}_p \leq \norm a_p + \norm b_p

    (write ai+bipai+bip1(ai+bi)\abs{a_i + b_i}^p \leq \abs{a_i + b_i}^{p-1}(\abs{a_i} + \abs{b_i}) and apply Hölder to each product). Conclude: p\norm\cdot_p is a norm on Rn\R^n for every p[1,+)p \in \intco{1}{+\infty}, completing the picture of Chapter 5.

  4. Integral versions: for f,gf, g continuous on [a,b]\intcc{a}{b}, state and prove Hölder and Minkowski for fp=(abfp)1/p\norm f_p = \bigl(\int_a^b\abs f^p\bigr)^{1/p} (same proofs, with the strict positivity of the integral for the equality discussion).
  5. Prove the monotonicity aqap\norm a_q \leq \norm a_p for 1pq1 \leq p \leq q, the limit apa\norm a_p \to \norm a_\infty as pp \to \infty, and the reverse comparison with the sharp constant:

    apn1p1qaq\norm a_p \leq n^{\frac1p - \frac1q}\,\norm a_q

    (Hölder against the constant vector). Identify the vectors achieving each equality.

  6. (Interpolation) For 1p<r<q1 \leq p < r < q and θ(0,1)\theta \in \intoo01 with 1r=θp+1θq\frac1r = \frac\theta p + \frac{1-\theta}q, prove

    arapθaq1θ\norm a_r \leq \norm a_p^{\theta}\, \norm a_q^{1-\theta}

    (apply Hölder with exponents pθr\frac{p}{\theta r} and q(1θ)r\frac{q}{(1-\theta)r} to aiθrai(1θ)r\abs{a_i}^{\theta r}\abs{a_i}^{(1-\theta)r}).

Part III — The power-mean chain, complete. For p0p \neq 0 set Mp=(iλixip)1/pM_p = \bigl(\sum_i\lambda_i x_i^p\bigr)^{1/p} (xi>0x_i > 0), and M0=ixiλiM_0 = \prod_i x_i^{\lambda_i}.

  1. Prove that pMpp \mapsto M_p is increasing on all of R\R^*: treat p<q<0p < q < 0 by the reciprocal identity Mp(x)=Mp(1/x)1M_{-p}(x) = M_p(1/x)^{-1}, and bridge through 00 by showing MpM0MqM_p \leq M_0 \leq M_q for p<0<qp < 0 < q (apply the concavity of ln\ln to xiqx_i^q, and the reversed inequality for negative exponents).
  2. Prove the limits MpmaxixiM_p \to \max_i x_i as p+p \to +\infty and MpminixiM_p \to \min_i x_i as pp \to -\infty.
  3. Write out the chain minHMGMAMQMmax\min \leq \mathrm{HM} \leq \mathrm{GM} \leq \mathrm{AM} \leq \mathrm{QM} \leq \max for equal weights, and prove the classic consequence: for positive a1,,ana_1, \dots, a_n,

    (iai)(i1ai)n2.\Bigl(\sum_i a_i\Bigr)\Bigl(\sum_i\frac1{a_i}\Bigr) \geq n^2 .
  4. Relate means to norms: for equal weights λi=1n\lambda_i = \frac1n, Mp(x)=n1/pxpM_p(x) = n^{-1/p}\norm x_p. Reconcile the two monotonicities — means increase with pp while norms decrease (question 10) — in one sentence about the factor n1/pn^{-1/p}.
  5. Determine the equality cases along the whole chain of question 14 (positive weights): equality anywhere forces all xix_i equal — strict convexity pays off.

Part IV — Dividends.

  1. (Young with a knob) For a,b0a, b \geq 0 and ε>0\varepsilon > 0, prove

    abεapp+εq/pbqq,ab \leq \varepsilon\,\frac{a^p}{p} + \varepsilon^{-q/p}\,\frac{b^q}{q},

    and the workhorse case abεa2+b24εab \leq \varepsilon a^2 + \frac{b^2}{4\varepsilon}: the absorption trick used throughout analysis.

  2. (Toward Carleman) Let ck=(k+1)kkk1c_k = \frac{(k+1)^k}{k^{k-1}}. Prove the telescoping identity k=1nck=(n+1)n\prod_{k=1}^{n}c_k = (n+1)^n, and deduce, by AM–GM applied to the numbers ckakc_ka_k,

    (a1a2an)1/n1n(n+1)k=1nckak(ak>0).(a_1a_2\cdots a_n)^{1/n} \leq \frac{1}{n(n+1)}\sum_{k=1}^{n} c_k a_k \qquad (a_k > 0).
  3. (Carleman’s inequality) Sum over nn, exchange the order of summation (positive summable families, Theorem 7.14), and use nk1n(n+1)=1k\sum_{n \geq k}\frac{1}{n(n+1)} = \frac1k and ck/k=(1+1k)k<ec_k/k = \bigl(1 + \frac1k\bigr)^k < \eu to conclude: for every convergent ak\sum a_k with positive terms,

    n=1(a1a2an)1/n    ek=1ak.\sum_{n=1}^{\infty}(a_1a_2\cdots a_n)^{1/n} \;\leq\; \eu\sum_{k=1}^{\infty}a_k .
  4. For ff continuous and positive on [0,1]\intcc{0}{1}, prove

    (01f)(011f)1,\Bigl(\int_0^1 f\Bigr)\Bigl(\int_0^1\frac1f\Bigr) \geq 1,

    with equality iff ff is constant (Cauchy–Schwarz on f1f\sqrt f\cdot\frac1{\sqrt f}).

  5. (Geometry of the balls) Using the equality case of Minkowski, show that for 1<p<1 < p < \infty the unit sphere of p\norm\cdot_p contains no segment (the norm is strictly convex in the sense of Problem 5.1), whereas for p=1p = 1 and p=p = \infty it does: exhibit the flat pieces.

Part V — Duality and synthesis.

  1. (Hölder as duality) Prove that for every aRna \in \R^n,

    ap=maxbq1 iaibi,\norm a_p = \max_{\norm b_q \leq 1}\ \sum_i a_ib_i ,

    exhibiting a maximizing bb explicitly. (The pp-norm is the dual of the qq-norm — the finite-dimensional germ of LpL^p duality.)

  2. (Moments) Let XX be a random variable taking finitely many positive values xix_i with probabilities λi\lambda_i. Restate question 12 as: rE[Xr]1/rr \mapsto \E[X^r]^{1/r} is increasing — the moment (Lyapunov) inequality, to be reused in Chapter 22.
  3. Solve with named tools, in two lines each: (i) for positive a,b,ca, b, c: a3+b3+c3(a+b+c)39a^3 + b^3 + c^3 \geq \frac{(a+b+c)^3}{9}; (ii) for positive x1,,xnx_1, \dots, x_n: (ixi)2nixi\bigl(\sum_i\sqrt{x_i}\bigr)^2 \leq n\sum_i x_i.
  4. (Synthesis) Draw the genealogy in five sentences: chord definition to slope lemma; slopes to support lines to Jensen; ln\ln’s concavity to Young to Hölder to Minkowski to the pp-norms; Jensen to the power-mean chain to moments; AM–GM to Carleman. Name the summits (Hölder–Minkowski; Carleman), and state where the toolbox is headed: the LpL^p spaces of the Year 3 volume, whose axioms are exactly questions 6 and 8.
Solution

Solution of Problem 8.1.

1. Convex \Rightarrow ff' increasing: for a<ba < b, the slope inequality gives, for small h>0h > 0, f(a+h)f(a)hf(b)f(a)baf(b)f(bh)h\frac{f(a+h) - f(a)}h \leq \frac{f(b) - f(a)}{b-a} \leq \frac{f(b) - f(b-h)}{h}; letting h0h \to 0: f(a)f(b)f(a)baf(b)f'(a) \leq \frac{f(b) - f(a)}{b - a} \leq f'(b). Conversely, if ff' increases and x<y<zx < y < z: the mean value theorem gives c1(x,y)c_1 \in \intoo{x}{y}, c2(y,z)c_2 \in \intoo yz with

f(y)f(x)yx=f(c1)f(c2)=f(z)f(y)zy,\frac{f(y) - f(x)}{y - x} = f'(c_1) \leq f'(c_2) = \frac{f(z) - f(y)}{z - y},

and this three-point slope inequality, applied with y=λx+(1λ)zy = \lambda x + (1 - \lambda)z, rearranges into the convexity inequality. For C2C^2: f0f'' \geq 0 iff ff' increases.

2. If f>0f'' > 0, ff' is strictly increasing, and the mean value computation above gives a strict inequality between the two chord slopes: strict convexity. Strict support: at an interior aa with support slope mm, if f(x0)=f(a)+m(x0a)f(x_0) = f(a) + m(x_0 - a) for some x0ax_0 \neq a, then on the segment from aa to x0x_0 the support line and the chord coincide, and strict convexity at the midpoint gives f(a+x02)<f(a)+mx0a2f\bigl(\frac{a + x_0}2\bigr) < f(a) + m\,\frac{x_0 - a}2, contradicting the support inequality. So f(x)>f(a)+m(xa)f(x) > f(a) + m(x - a) for all xax \neq a. Strict Jensen: with a=λixia = \sum\lambda_ix_i, averaging the support inequalities gives λif(xi)f(a)\sum\lambda_if(x_i) \geq f(a), with equality iff each term is an equality, i.e. iff every xi=ax_i = a.

3. (ln)=1t2>0(-\ln)'' = \frac1{t^2} > 0; (tr)=r(r1)tr2(t^r)'' = r(r - 1)t^{r-2}, positive for r>1r > 1, negative for 0<r<10 < r < 1; exp=exp>0\exp'' = \exp > 0. All strict by question 2.

4. The cases ab=0ab = 0 are trivial. For a,b>0a, b > 0, concavity of ln\ln at the points ap,bqa^p, b^q with weights 1p,1q\frac1p, \frac1q:

ln(app+bqq)1pln(ap)+1qln(bq)=ln(ab),\ln\Bigl(\frac{a^p}p + \frac{b^q}q\Bigr) \geq \frac1p\ln(a^p) + \frac1q\ln(b^q) = \ln(ab),

and ln\ln increases: abapp+bqqab \leq \frac{a^p}p + \frac{b^q}q. Equality iff the two points coincide (strict concavity): ap=bqa^p = b^q.

5. Concavity of ln\ln with weights λi\lambda_i: ln(λixi)λilnxi=lnxiλi\ln\bigl(\sum\lambda_ix_i\bigr) \geq \sum\lambda_i\ln x_i = \ln\prod x_i^{\lambda_i}; exponentiate. Equality iff all xix_i equal (question 2). The route of Exercise 8.6 obtained the same inequality as a limit of power means; here it is one application of Jensen — the toolbox has redundancy built in.

6. If a=0a = 0 or b=0b = 0 the inequality is trivial. Normalize: replacing aa by a/apa/\norm a_p and bb by b/bqb/\norm b_q, we may assume ap=bq=1\norm a_p = \norm b_q = 1 and must show aibi1\sum\abs{a_ib_i} \leq 1. Young termwise:

iaibii(aipp+biqq)=1p+1q=1.\sum_i\abs{a_i}\abs{b_i} \leq \sum_i\Bigl(\frac{\abs{a_i}^p}{p} + \frac{\abs{b_i}^q}{q}\Bigr) = \frac1p + \frac1q = 1 .

Equality iff each Young inequality is tight: aip=biq\abs{a_i}^p = \abs{b_i}^q for all ii — after undoing the normalization, (aip)(\abs{a_i}^p) proportional to (biq)(\abs{b_i}^q).

7. p=q=2p = q = 2 is Cauchy–Schwarz with the same equality case (proportionality). Endpoint: aibi(maxibi)iai=a1b\sum\abs{a_ib_i} \leq \bigl(\max_i\abs{b_i}\bigr)\sum_i\abs{a_i} = \norm a_1\norm b_\infty, immediate termwise.

8. For p=1p = 1 it is the triangle inequality termwise. For p>1p > 1, with qq conjugate:

a+bpp=iai+bipiai+bip1ai+iai+bip1bi,\norm{a+b}_p^p = \sum_i\abs{a_i + b_i}^p \leq \sum_i\abs{a_i+b_i}^{p-1}\abs{a_i} + \sum_i\abs{a_i+b_i}^{p-1}\abs{b_i},

and Hölder on each sum, noting (p1)q=p(p - 1)q = p:

iai+bip1ai(iai+bip)1/qap=a+bpp/qap,\sum_i\abs{a_i+b_i}^{p-1}\abs{a_i} \leq \Bigl(\sum_i\abs{a_i+b_i}^{p}\Bigr)^{1/q}\norm a_p = \norm{a + b}_p^{p/q}\,\norm a_p ,

likewise with bb. Hence a+bppa+bpp/q(ap+bp)\norm{a+b}_p^p \leq \norm{a + b}_p^{p/q}\bigl(\norm a_p + \norm b_p\bigr); if a+b0a + b \neq 0, divide by a+bpp/q\norm{a+b}_p^{p/q} and use ppq=1p - \frac pq = 1. With homogeneity and separation (clear), p\norm\cdot_p is a norm on Rn\R^n.

9. For continuous f,gf, g on [a,b]\intcc ab: Hölder

abfg(abfp)1/p(abgq)1/q\int_a^b\abs{fg} \leq \Bigl(\int_a^b\abs f^p\Bigr)^{1/p}\Bigl(\int_a^b\abs g^q\Bigr)^{1/q}

by the same normalization plus pointwise Young, integrated; and Minkowski f+gpfp+gp\norm{f + g}_p \leq \norm f_p + \norm g_p by the same splitting, Hölder on each piece. Separation of the norm uses strict positivity: a continuous fp\abs f^p with zero integral vanishes identically (Year 1 volume).

10. Monotonicity: we may assume ap=1\norm a_p = 1; then each ai1\abs{a_i} \leq 1, so aiqaip\abs{a_i}^q \leq \abs{a_i}^p and aqq1\norm a_q^q \leq 1: aq1=ap\norm a_q \leq 1 = \norm a_p. Equality requires aiq=aip\abs{a_i}^q = \abs{a_i}^p for every ii, i.e. each ai{0,1}\abs{a_i} \in \{0, 1\}; with aip=1\sum\abs{a_i}^p = 1 this leaves exactly one coordinate of modulus 11: equality iff aa has at most one nonzero coordinate. Limit: aapn1/pa\norm a_\infty \leq \norm a_p \leq n^{1/p}\norm a_\infty, and n1/p1n^{1/p} \to 1. Reverse comparison: Hölder with exponents qp\frac qp and its conjugate qqp\frac{q}{q-p}, applied to aip1\abs{a_i}^p\cdot 1:

app=iaip1(iaiq)p/qn1p/q=aqp  n1p/q,\norm a_p^p = \sum_i\abs{a_i}^p\cdot 1 \leq \Bigl(\sum_i\abs{a_i}^{q}\Bigr)^{p/q}\,n^{1 - p/q} = \norm a_q^{p}\; n^{1-p/q},

whence apn1p1qaq\norm a_p \leq n^{\frac1p - \frac1q}\norm a_q, with equality iff all ai\abs{a_i} are equal (the Hölder equality case against the constant vector).

11. Write air=aiθrai(1θ)r\abs{a_i}^r = \abs{a_i}^{\theta r}\,\abs{a_i}^{(1-\theta)r} and apply Hölder with the conjugate exponents pθr\frac{p}{\theta r} and q(1θ)r\frac{q}{(1-\theta)r} (conjugate precisely because θrp+(1θ)rq=1\frac{\theta r}p + \frac{(1-\theta)r}q = 1):

arr=iaiθrai(1θ)r(iaip)θr/p(iaiq)(1θ)r/q=apθraq(1θ)r.\norm a_r^r = \sum_i \abs{a_i}^{\theta r}\abs{a_i}^{(1-\theta)r} \leq \Bigl(\sum_i\abs{a_i}^{p}\Bigr)^{\theta r/p} \Bigl(\sum_i\abs{a_i}^{q}\Bigr)^{(1-\theta)r/q} = \norm a_p^{\theta r}\,\norm a_q^{(1-\theta)r} .

Take rr-th roots: the pp-norms are log-convex in 1p\frac1p.

12. Both negative: if p<q<0p < q < 0 then 0<q<p0 < -q < -p, and Mq(y)Mp(y)M_{-q}(y) \leq M_{-p}(y) for the positive exponents (course case, Example 8.8) applied to y=(1/xi)y = (1/x_i); inverting the identity Mp(x)=Mp(1/x)1M_p(x) = M_{-p}(1/x)^{-1} reverses the inequality into Mp(x)Mq(x)M_p(x) \leq M_q(x). Bridge: for q>0q > 0, concavity of ln\ln gives lnMq=1qln(λixiq)1qλilnxiq=lnM0\ln M_q = \frac1q\ln\bigl(\sum\lambda_ix_i^q\bigr) \geq \frac1q\sum\lambda_i\ln x_i^q = \ln M_0; for p<0p < 0, the same concavity gives ln(λixip)pλilnxi\ln\bigl(\sum\lambda_ix_i^p\bigr) \geq p\sum\lambda_i\ln x_i, and dividing by p<0p < 0 flips: lnMplnM0\ln M_p \leq \ln M_0. Hence MpM0MqM_p \leq M_0 \leq M_q whenever p<0<qp < 0 < q: with the two same-sign cases, MM increases on all of R\R^* (and through 00).

13. Let xmax=maxxix_{\max} = \max x_i, attained at ii^*. For p>0p > 0:

λi1/pxmaxMpxmax,\lambda_{i^*}^{1/p}\,x_{\max} \leq M_p \leq x_{\max},

and λi1/p1\lambda_{i^*}^{1/p} \to 1: MpxmaxM_p \to x_{\max}. For pp \to -\infty: Mp(x)=Mp(1/x)1(maxi1xi)1=minixiM_p(x) = M_{-p}(1/x)^{-1} \to \bigl(\max_i\frac1{x_i}\bigr)^{-1} = \min_ix_i.

14. With λi=1n\lambda_i = \frac1n, the chain MM1M0M1M2M+M_{-\infty} \leq M_{-1} \leq M_0 \leq M_1 \leq M_2 \leq M_{+\infty} reads

minn1ai(ai)1/nainai2nmax.\min \leq \frac{n}{\sum\frac1{a_i}} \leq \Bigl(\prod a_i\Bigr)^{1/n} \leq \frac{\sum a_i}{n} \leq \sqrt{\frac{\sum a_i^2}{n}} \leq \max .

AM–HM (M1M1M_{-1} \leq M_1) rearranges directly into (ai)(1ai)n2\bigl(\sum a_i\bigr)\bigl(\sum\frac1{a_i}\bigr) \geq n^2.

15. With equal weights, Mp(x)=(1nxip)1/p=n1/pxpM_p(x) = \bigl(\frac1n\sum\abs{x_i}^p\bigr)^{1/p} = n^{-1/p}\norm x_p. As pp grows, xp\norm x_p decreases (question 10) but the normalizer n1/pn^{-1/p} increases faster, and the product increases (question 12): means average, norms accumulate, and the factor n1/pn^{-1/p} is exactly the exchange rate between the two bookkeeping conventions.

16. Each link is an instance of strict Jensen (question 2) with the strictly convex/concave functions of question 3 (tq/pt^{q/p}, ln\ln), so equality at any link forces all the xix_i equal; and min=Mp\min = M_p or Mp=maxM_p = \max likewise forces all values equal to the common extremum. The chain is strict as soon as two xix_i differ.

17. Apply Young (question 4) to the pair ε1/pa\varepsilon^{1/p}a and ε1/pb\varepsilon^{-1/p}b:

ab=(ε1/pa)(ε1/pb)εapp+εq/pbqq.ab = (\varepsilon^{1/p}a)(\varepsilon^{-1/p}b) \leq \varepsilon\,\frac{a^p}p + \varepsilon^{-q/p}\,\frac{b^q}q .

For p=q=2p = q = 2, replacing ε\varepsilon by 2ε2\varepsilon: abεa2+b24εab \leq \varepsilon a^2 + \frac{b^2}{4\varepsilon} — the absorption inequality: a product is traded for a small multiple of one square plus a large multiple of the other.

18. Telescoping:

k=1nck=k=1n(k+1)kk=1nkk1=2132(n+1)n1021nn1=(n+1)n,\prod_{k=1}^{n}c_k = \frac{\prod_{k=1}^n(k+1)^k} {\prod_{k=1}^{n}k^{k-1}} = \frac{2^1\,3^2\cdots(n+1)^n}{1^0\,2^1\cdots n^{n-1}} = (n+1)^n,

every factor (k+1)k(k+1)^k of the numerator cancelling against the denominator’s next term. AM–GM on the nn numbers ckakc_ka_k:

(a1an)1/n=(kckak)1/n(n+1)1n+11nk=1nckak.(a_1\cdots a_n)^{1/n} = \frac{\bigl(\prod_k c_ka_k\bigr)^{1/n}}{(n+1)} \leq \frac{1}{n+1}\cdot\frac1n\sum_{k=1}^{n}c_ka_k .

19. Summing over nn and exchanging the two summations (all terms positive: Theorem 7.14):

n1(a1an)1/nn11n(n+1)k=1nckak=k1ckaknk1n(n+1)=k1ckakk,\sum_{n\geq1}(a_1\cdots a_n)^{1/n} \leq \sum_{n\geq1}\frac{1}{n(n+1)}\sum_{k=1}^{n}c_ka_k = \sum_{k\geq1}c_ka_k\sum_{n\geq k}\frac1{n(n+1)} = \sum_{k\geq1}\frac{c_ka_k}{k},

using the telescoping nk(1n1n+1)=1k\sum_{n\geq k}\bigl(\frac1n - \frac1{n+1}\bigr) = \frac1k. Finally ckk=(k+1)kkk=(1+1k)k<e\frac{c_k}k = \frac{(k+1)^k}{k^k} = \bigl(1 + \frac1k\bigr)^k < \eu (increasing sequence with limit e\eu, Year 1 volume):

n1(a1an)1/nek1ak:\sum_{n\geq1}(a_1\cdots a_n)^{1/n} \leq \eu\sum_{k\geq1}a_k :

Carleman’s inequality. (The constant e\eu is optimal, though we do not prove it.)

20. Cauchy–Schwarz (question 9, p=q=2p = q = 2) applied to f\sqrt f and 1f\frac1{\sqrt f}:

1=(01f1f)2(01f)(011f).1 = \Bigl(\int_0^1\sqrt f\cdot\frac{1}{\sqrt f}\Bigr)^{2} \leq \Bigl(\int_0^1 f\Bigr)\Bigl(\int_0^1\frac1f\Bigr).

Equality iff f\sqrt f and 1f\frac1{\sqrt f} are proportional, i.e. f2f^2 constant, i.e. ff constant (f>0f > 0 continuous).

21. Let 1<p<1 < p < \infty, ap=bp=1\norm a_p = \norm b_p = 1, aba \neq b, and suppose a+b2p=1\bigl\Vert\frac{a+b}2\bigr\Vert_p = 1, i.e. Minkowski is an equality for a,ba, b. Tracing question 8’s proof, equality forces equality in both Hölder applications and in the termwise triangle inequalities: (aip)(\abs{a_i}^p) and (bip)(\abs{b_i}^p) both proportional to (ai+bip)(\abs{a_i + b_i}^p), and ai,bia_i, b_i of the same sign — hence b=tab = ta for some t0t \geq 0, and bp=ap\norm b_p = \norm a_p gives t=1t = 1: b=ab = a, contradiction. So the pp-sphere contains no midpoint of distinct sphere points: no segment. For p=p = \infty in R2\R^2: all (1,t)(1, t), t1\abs t \leq 1, lie on the unit sphere — a flat edge; for p=1p = 1: the segment (t,1t)(t, 1 - t), t[0,1]t \in \intcc01, does.

22. For a=0a = 0 both sides vanish. Otherwise Hölder bounds every aibi\sum a_ib_i by apbqap\norm a_p\norm b_q \leq \norm a_p. Attainment: take

bi=sign(ai)aip1app/q:bqq=iai(p1)qapp=appapp=1,iaibi=appapp/q=ap,b_i = \frac{\operatorname{sign}(a_i)\,\abs{a_i}^{p-1}} {\norm a_p^{p/q}} : \qquad \norm b_q^q = \frac{\sum_i\abs{a_i}^{(p-1)q}}{\norm a_p^{p}} = \frac{\norm a_p^p}{\norm a_p^p} = 1, \quad \sum_ia_ib_i = \frac{\norm a_p^p}{\norm a_p^{p/q}} = \norm a_p ,

using (p1)q=p(p-1)q = p and ppq=1p - \frac pq = 1. So the supremum is a maximum, equal to ap\norm a_p: each pp-norm is the dual norm of its conjugate — the germ of LpL^pLqL^q duality.

23. E[Xr]=iλixir\E[X^r] = \sum_i\lambda_ix_i^r, so E[Xr]1/r=Mr(x;λ)\E[X^r]^{1/r} = M_r(x; \lambda), increasing in rr by question 12 (and through r0,±r \to 0, \pm\infty by questions 12–13): Lyapunov’s moment inequality, purely a statement about weighted power means. It returns for genuine random variables in Chapter 22.

24. (i) Power means M1M3M_1 \leq M_3 with equal weights: a+b+c3(a3+b3+c33)1/3\frac{a+b+c}3 \leq \bigl(\frac{a^3+b^3+c^3}3\bigr)^{1/3}; cube and multiply by 33: a3+b3+c3(a+b+c)39a^3 + b^3 + c^3 \geq \frac{(a+b+c)^3}9. (ii) Cauchy–Schwarz against the constant vector: ixi1(ixi)1/2n1/2\sum_i\sqrt{x_i}\cdot1 \leq \bigl(\sum_ix_i\bigr)^{1/2}n^{1/2}; square.

25. The chord definition yields the slope lemma by one algebraic rearrangement; slopes squeezed at a point produce one-sided derivatives and support lines, whose weighted average is Jensen. Applied to ln-\ln, Jensen becomes Young, which summed against normalized vectors is Hölder, which split and reabsorbed is Minkowski — and the pp-norms of Chapter 5 are born, with their duality (question 22) and their geometry (question 21). Jensen applied along the scale of powers chains all the means from min\min to max\max (questions 12–14), which read on random variables is the moment inequality (question 23). And AM–GM, weighted by one telescoping trick, yields Carleman’s bound with its irreducible constant e\eu (questions 18–19). Summits: Hölder–Minkowski, and Carleman. Destination: the LpL^p spaces of the Year 3 volume, whose founding axioms are exactly questions 6 and 8 with integrals in place of sums.