Mathematics · Book 5 · Bachelor Year 3

University Mathematics — Year 3

University Mathematics — Year 3 · Bachelor Year 3

20Submanifolds of ℝn

Spheres, tori, rotation groups: the natural habitats of geometry and mechanics are not vector spaces but curved sets that look flat up close. This chapter gives that phrase a precise meaning — submanifolds of Rn\R^n — and the calculus to work on them. The foundation is the inverse function theorem, proved here by the Banach fixed point; everything else is change of coordinates: the four equivalent descriptions of a submanifold (local straightening, level sets, graphs, parametrizations), tangent spaces, and constrained optimization by Lagrange multipliers — which, as a parting demonstration, re-proves the spectral theorem for symmetric matrices in three lines of geometry. The weekend problem builds the rotation group SO(3)SO(3) and its quaternionic double cover: algebra (Chapter 1’s Q8Q_8 grown up) meeting geometry.

20.1 The inverse function theorem

Theorem 20.1 (Inverse function theorem)

Let URnU \subseteq \R^n be open, f ⁣:URnf \colon U \to \R^n of class C1\mathcal C^1, and aUa \in U with Df(a)Df(a) invertible. Then there are open sets VaV \ni a, Wf(a)W \ni f(a) such that f ⁣:VWf \colon V \to W is a bijection with C1\mathcal C^1 inverse, and

D(f1)(y)=(Df(f1(y)))1(yW).D(f^{-1})(y) = \bigl(Df(f^{-1}(y))\bigr)^{-1} \qquad (y \in W).

If ff is Ck\mathcal C^k, so is f1f^{-1}.

Proof. Normalize: replacing ff by xDf(a)1(f(a+x)f(a))x \mapsto Df(a)^{-1}\bigl(f(a + x) - f(a)\bigr), we may assume a=0a = 0, f(0)=0f(0) = 0, Df(0)=IDf(0) = I (the general statement follows by composing with the affine bijections). Write f(x)=x+g(x)f(x) = x + g(x): Dg(0)=0Dg(0) = 0, and by continuity of DgDg choose r>0r > 0 with Dg(x)12\vertiii{Dg(x)} \leq \frac12 on Bˉ(0,r)\bar B(0, r); the mean value inequality gives g(x)g(x)12xx\norm{g(x) - g(x')} \leq \frac12 \norm{x - x'} there.

Bijectivity onto a neighborhood. For yB(0,r2)y \in B(0, \frac r2), solving f(x)=yf(x) = y means finding a fixed point of Φy(x)=yg(x)\Phi_y(x) = y - g(x); Φy\Phi_y maps Bˉ(0,r)\bar B(0, r) into itself (Φy(x)y+12xr\norm{\Phi_y(x)} \leq \norm y + \frac12\norm x \leq r) and is 12\frac12-Lipschitz: Banach (Theorem 7.4) gives a unique solution x=φ(y)Bˉ(0,r)x = \varphi(y) \in \bar B(0,r). Moreover ff is injective on Bˉ(0,r)\bar B(0,r):

f(x)f(x)xxg(x)g(x)12xx.()\norm{f(x) - f(x')} \geq \norm{x - x'} - \norm{g(x) - g(x')} \geq \tfrac12\norm{x - x'} . \tag{$*$}

Set W=B(0,r2)W = B(0, \frac r2) and V=f1(W)B(0,r)V = f^{-1}(W)\cap B(0, r): open (continuity), with f ⁣:VWf \colon V \to W bijective.

Continuity and differentiability of the inverse. ()(*) says φ=f1\varphi = f^{-1} is 22-Lipschitz. Fix y0=f(x0)Wy_0 = f(x_0) \in W; invertibility of A=Df(x0)A = Df(x_0) (its distance to II is 12\leq \frac12: Neumann, Proposition 8.4) and differentiability of ff give, for y=f(x)y = f(x) near y0y_0:

φ(y)φ(y0)A1(yy0)=A1(f(x)f(x0)A(xx0))=A1o(xx0)=o(yy0),\varphi(y) - \varphi(y_0) - A^{-1}(y - y_0) = -A^{-1}\bigl(f(x) - f(x_0) - A(x - x_0)\bigr) = -A^{-1}\,o(\norm{x - x_0}) = o(\norm{y - y_0}),

using ()(*) to convert xx02yy0\norm{x - x_0} \leq 2\norm{y - y_0}: φ\varphi is differentiable at y0y_0 with the inverse differential. Continuity of yDφ(y)=Df(φ(y))1y \mapsto D\varphi(y) = Df(\varphi(y))^{-1}: composition of continuous maps (inversion is continuous, Proposition 8.4): φC1\varphi \in \mathcal C^1; bootstrapping the same formula gives Ck\mathcal C^k.

Theorem 20.2 (Implicit function theorem)

Let F ⁣:URp×RqRqF \colon U \subseteq \R^p\times\R^q \to \R^q be C1\mathcal C^1 near (a,b)(a, b), F(a,b)=0F(a,b) = 0, and suppose the partial differential DyF(a,b)L(Rq)D_yF(a,b) \in \mathcal L(\R^q) is invertible. Then there are neighborhoods AaA \ni a, BbB \ni b and a C1\mathcal C^1 map ψ ⁣:AB\psi \colon A \to B with

{(x,y)A×B:F(x,y)=0}={(x,ψ(x)):xA},\bigl\{(x, y)\in A\times B : F(x,y) = 0\bigr\} = \{(x, \psi(x)) : x \in A\},

and Dψ(x)=DyF(x,ψ(x))1DxF(x,ψ(x))D\psi(x) = -D_yF(x, \psi(x))^{-1}\,D_xF(x, \psi(x)).

Proof. Apply Theorem 20.1 to Θ(x,y)=(x,F(x,y))\Theta(x, y) = (x, F(x,y)): its differential at (a,b)(a,b), block-triangular with invertible diagonal blocks II and DyFD_yF, is invertible. The local inverse has the form Θ1(x,z)=(x,h(x,z))\Theta^{-1}(x, z) = (x, h(x, z)); set ψ(x)=h(x,0)\psi(x) = h(x, 0): then F(x,y)=0F(x, y) = 0 iff Θ(x,y)=(x,0)\Theta(x,y) = (x, 0) iff y=ψ(x)y = \psi(x), locally. The formula: differentiate F(x,ψ(x))=0F(x, \psi(x)) = 0 by the chain rule.

20.2 Submanifolds: four definitions

Theorem 20.3 (Equivalent characterizations)

Let MRnM \subseteq \R^n, d{0,,n}d \in \{0, \dots, n\}, and k1k \geq 1. The following are equivalent, for each point aMa \in M (and MM is a dd-dimensional submanifold of class Ck\mathcal C^k if they hold at every aMa \in M):

  1. (Straightening) There is a Ck\mathcal C^k diffeomorphism Φ\Phi from an open Ωa\Omega \ni a onto an open ΩRn\Omega' \subseteq \R^n with

    Φ(MΩ)=Ω(Rd×{0}).\Phi(M\cap\Omega) = \Omega' \cap \bigl(\R^d\times\{0\}\bigr).
  2. (Level set) There is a Ck\mathcal C^k submersion F ⁣:ΩRndF \colon \Omega \to \R^{n-d} (i.e. DF(x)DF(x) surjective) on an open Ωa\Omega \ni a with MΩ=F1(0)M\cap\Omega = F^{-1}(0).
  3. (Graph) Up to permuting coordinates, MM is locally the graph of a Ck\mathcal C^k map ψ ⁣:ARdRnd\psi \colon A \subseteq \R^d \to \R^{n-d}.
  4. (Parametrization) There is a Ck\mathcal C^k immersion φ ⁣:ARdRn\varphi \colon A \subseteq \R^d \to \R^n (Dφ(u)D\varphi(u) injective) with AA open, φ\varphi a homeomorphism from AA onto MΩM \cap \Omega for some open Ωa\Omega \ni a.

Proof. (1)\Rightarrow(2): F=(Φd+1,,Φn)F = (\Phi_{d+1}, \dots, \Phi_n) (last coordinates of Φ\Phi): a submersion (DΦD\Phi invertible). (2)\Rightarrow(3): DF(a)DF(a) surjective: some q×qq \times q minor of the Jacobian is invertible (q=ndq = n - d); after permuting coordinates, DyF(a)D_yF(a) is invertible, and the implicit function theorem (Theorem 20.2) expresses MM locally as a graph y=ψ(x)y = \psi(x). (3)\Rightarrow(4): φ(x)=(x,ψ(x))\varphi(x) = (x, \psi(x)): an immersion (differential (IDψ)\bigl(\begin{smallmatrix}I\\ D\psi\end{smallmatrix}\bigr) injective), a homeomorphism onto the graph (inverse: the projection, continuous). (4)\Rightarrow(1): let φ(u0)=a\varphi(u_0) = a; complete imDφ(u0)\operatorname{im}D\varphi(u_0) by a supplement EE (dimE=nd\dim E = n - d) and define Θ(u,v)=φ(u)+v\Theta(u, v) = \varphi(u) + v on A×EA\times E: DΘ(u0,0)D\Theta(u_0, 0) is bijective (image contains imDφ(u0)\operatorname{im}D\varphi(u_0) and EE), so Θ\Theta is a local diffeomorphism (Theorem 20.1); its inverse Φ\Phi straightens: near aa, points of MM are exactly the φ(u)=Θ(u,0)\varphi(u) = \Theta(u, 0) — for this, the homeomorphism hypothesis in (4) guarantees that MΩM\cap\Omega, for small Ω\Omega, contains no other sheets (φ(u)+v=mM\varphi(u') + v = m \in M close to aa with small v0v \neq 0 must be excluded: m=φ(u)m = \varphi(u'') for some uu'' near u0u_0 by the homeomorphism property, and local injectivity of Θ\Theta forces v=0v = 0). Then Φ(MΩ)=(A×{0})Φ(Ω)\Phi(M\cap\Omega) = (A\times\{0\}) \cap\Phi(\Omega) up to shrinking.

Example 20.4

The sphere Sn1={x2=1}S^{n-1} = \{\norm x^2 = 1\}: level set of the submersion F(x)=x221F(x) = \norm x_2^2 - 1 on Rn{0}\R^n\setminus\{0\} (DF(x)=2xT0DF(x) = 2x^{\mathsf T} \ne 0): a C\mathcal C^\infty submanifold of dimension n1n - 1. The torus in R3\R^3: level set of (x2+y2R)2+z2r2\bigl(\sqrt{x^2 + y^2} - R\bigr)^2 + z^2 - r^2 (0<r<R0 < r < R). The cone {x2+y2=z2}\{x^2 + y^2 = z^2\} is not a submanifold at 00 (Exercise 20.1). Matrix groups: SLnSL_n and OnO_n are submanifolds of Mn(R)M_n(\R) (Exercises 20.5 and 20.6) — the starting point of Lie theory.

20.3 Tangent spaces

Definition 20.5

Let MM be a dd-submanifold and aMa \in M. The tangent space TaMT_aM is the set of velocity vectors γ(0)\gamma'(0) of C1\mathcal C^1 curves γ ⁣:(ε,ε)M\gamma \colon \intoo{-\varepsilon}\varepsilon \to M with γ(0)=a\gamma(0) = a.

Proposition 20.6

TaMT_aM is a dd-dimensional vector subspace of Rn\R^n, and:

  1. if M=F1(0)M = F^{-1}(0) locally with FF a submersion: TaM=kerDF(a)T_aM = \ker DF(a);
  2. if MM is parametrized by the immersion φ\varphi (φ(u0)=a\varphi(u_0) = a): TaM=imDφ(u0)T_aM = \operatorname{im}D\varphi(u_0).

Proof. Curves in MM satisfy F(γ(t))=0F(\gamma(t)) = 0; the chain rule at 00 gives DF(a)γ(0)=0DF(a)\gamma'(0) = 0: TaMkerDF(a)T_aM \subseteq \ker DF(a). Conversely, straightening (Theorem 20.3(1)) transports lines of Rd×{0}\R^d\times\{0\} to curves of MM: every vector of a dd-dimensional subspace is realized; comparing dimensions (dimkerDF(a)=n(nd)=d\dim\ker DF(a) = n - (n - d) = d) forces equality in (1), and the same transport argument gives (2) (Dφ(u0)D\varphi(u_0) applied to straight lines in AA; dimensions again).

Theorem 20.7 (Lagrange multipliers)

Let M=F1(0)M = F^{-1}(0) with F=(F1,,Fq) ⁣:ΩRqF = (F_1, \dots, F_q) \colon \Omega \to \R^q a C1\mathcal C^1 submersion, and let f ⁣:ΩRf \colon \Omega \to \R be C1\mathcal C^1. If the restriction fMf\restriction_M has a local extremum at aMa \in M, then there are unique reals λ1,,λq\lambda_1, \dots, \lambda_q (Lagrange multipliers) with

f(a)=λ1F1(a)++λqFq(a).\nabla f(a) = \lambda_1\nabla F_1(a) + \dots + \lambda_q\nabla F_q(a) .

Proof. For every curve γ\gamma in MM through aa: tf(γ(t))t \mapsto f(\gamma(t)) has a local extremum at 00, so 0= ⁣d ⁣dtf(γ(t))0=f(a),γ(0)0 = \frac{\dd}{\dd t}f(\gamma(t))\big|_0 = \langle\nabla f(a), \gamma'(0)\rangle: f(a)TaM=kerDF(a)\nabla f(a) \perp T_aM = \ker DF(a) (Proposition 20.6). Now kerDF(a)=imDF(a)T\ker DF(a)^\perp = \operatorname{im}DF(a)^{\mathsf T} (the rank/orthogonality identity of Year 2, or Exercise 13.8 in finite dimension: (kerT)=imT(\ker T)^\perp = \operatorname{im}T^*), which is spanned by the gradients Fi(a)\nabla F_i(a) — independent, as DF(a)DF(a) is surjective: the multipliers exist and are unique.

Example 20.8 (The spectral theorem, geometrically)

Let AA be a real symmetric n×nn\times n matrix and maximize f(x)=Ax,xf(x) = \langle Ax, x\rangle on the sphere Sn1S^{n-1} (compact: the maximum is attained, at some v1v_1). Lagrange with F(x)=x21F(x) = \norm x^2 - 1: f=2Ax\nabla f = 2Ax and F=2x\nabla F = 2x give Av1=λ1v1Av_1 = \lambda_1v_1 — an eigenvector, with λ1=maxSn1Ax,x\lambda_1 = \max_{S^{n-1}}\langle Ax, x\rangle. Restrict AA to v1v_1^\perp (invariant: Av,v1=v,Av1=λ1v,v1=0\langle Av, v_1\rangle = \langle v, Av_1\rangle = \lambda_1\langle v, v_1\rangle = 0) and iterate: an orthonormal basis of eigenvectors. The spectral theorem of Year 2, re-proved by pure optimization — and the infinite-dimensional shadow of the same argument proved Lemma 15.6.

Method 20.9

To prove a set is a submanifold: exhibit it locally as F1(0)F^{-1}(0) with DFDF surjective on the set (the most common route), or as a graph. To compute its dimension and tangent space: d=nqd = n - q and Ta=kerDF(a)T_a = \ker DF(a). To optimize on it: Lagrange — always check compactness (or coercivity) first, so an extremum exists to which the theorem can apply, and remember that the multiplier equation is only necessary: collect all critical points, then compare values. For matrix groups, differentiate curves at the identity to identify tangent spaces.

20.4 Exercises

Exercise 20.1

(a) Verify that the following are C\mathcal C^\infty submanifolds and give their dimensions: Sn1S^{n-1}; the hyperboloid {x2+y2z2=1}\{x^2 + y^2 - z^2 = 1\}; the torus of Example 20.4. (b) Show that the cone C={x2+y2=z2}R3C = \{x^2 + y^2 = z^2\} \subseteq \R^3 is not a 22-submanifold at 00: determine the number of connected components of (C{0})B(0,ε)\bigl(C\setminus\{0\}\bigr)\cap B(0,\varepsilon), and compare with the number a straightening (Theorem 20.3(1)) would force for a plane minus a point.

Solution

Solution of Exercise 20.1.

(a) Each is F1(0)F^{-1}(0) for a submersion: x21\norm x^2 - 1 on Rn{0}\R^n\setminus\{0\} (gradient 2x02x \neq 0): dimension n1n-1; x2+y2z21x^2 + y^2 - z^2 - 1 (gradient (2x,2y,2z)0(2x, 2y, -2z) \neq 0 on the hyperboloid, where x2+y2=1+z2>0x^2 + y^2 = 1 + z^2 > 0): dimension 22; the torus function G=(ρR)2+z2r2G = (\rho - R)^2 + z^2 - r^2, ρ=x2+y2\rho = \sqrt{x^2+y^2}, is C\mathcal C^\infty near the torus (ρRr>0\rho \geq R - r > 0 there) with G0\nabla G \neq 0 (its zz-component is 2z2z, and where z=0z = 0 the radial component is 2(ρR)02(\rho - R)\ne0 since ρR=r\abs{\rho - R} = r): dimension 22.

(b) For small ε\varepsilon, (C{0})B(0,ε)(C\setminus\{0\})\cap B(0,\varepsilon) has exactly 22 connected components (upper and lower punctured nappes, each path-connected: connect through circles and rays). If CC were a 22-submanifold at 00, a straightening would give a homeomorphism from CΩC\cap\Omega onto an open piece of a plane sending 00 to a point pp; small punctured plane-neighborhoods of pp have one component, and homeomorphisms preserve the number of components of punctured neighborhoods: contradiction.

Exercise 20.2

Compute the tangent spaces: (a) TaSn1T_aS^{n-1} for any aa (answer: aa^\perp); (b) the tangent plane to the torus of Example 20.4 at an arbitrary point of the outer equator {z=0, x2+y2=(R+r)2}\{z = 0,\ x^2 + y^2 = (R + r)^2\}; (c) the tangent line to the helix φ(t)=(cost,sint,t)\varphi(t) = (\cos t, \sin t, t) at φ(t0)\varphi(t_0), checking Proposition 20.6(2).

Solution

Solution of Exercise 20.2.

(a) TaSn1=ker(2aT)=aT_aS^{n-1} = \ker\bigl(2a^{\mathsf T}\bigr) = a^\perp. (b) At p=((R+r)cosθ,(R+r)sinθ,0)p = ((R+r)\cos\theta, (R+r)\sin\theta, 0): G=(2rcosθ,2rsinθ,0)\nabla G = (2r\cos\theta, 2r\sin\theta, 0), so the tangent plane is Vect((sinθ,cosθ,0), (0,0,1))\operatorname{Vect}\bigl((-\sin\theta, \cos\theta, 0),\ (0, 0, 1)\bigr): the vertical plane tangent to the outer equator. (c) The helix is an embedded curve with φ(t0)=(sint0,cost0,1)0\varphi'(t_0) = (-\sin t_0, \cos t_0, 1) \neq 0: the tangent line at φ(t0)\varphi(t_0) is φ(t0)+Rφ(t0)\varphi(t_0) + \R\,\varphi'(t_0), as Proposition 20.6(2) prescribes.

Exercise 20.3 ★★

Let f(x,y)=(x2y2, 2xy)f(x, y) = (x^2 - y^2,\ 2xy) (i.e. zz2z \mapsto z^2). (a) At which points does Theorem 20.1 apply? (b) Show that ff is locally but not globally invertible on R2{0}\R^2\setminus\{0\}, and exhibit explicitly the two local inverses defined on a neighborhood of (1,0)(1, 0) (the two square-root branches). (c) Same discussion for the polar coordinates map (r,θ)(rcosθ,rsinθ)(r, \theta) \mapsto (r\cos\theta, r\sin\theta).

Solution

Solution of Exercise 20.3.

(a) Df(x,y)=(2x2y2y2x)Df(x,y) = \bigl(\begin{smallmatrix}2x & -2y\\ 2y & 2x\end{smallmatrix}\bigr), det=4(x2+y2)\det = 4(x^2 + y^2): the theorem applies at every point except the origin. (b) f(z)=f(z)f(-z) = f(z): never injective on a set symmetric about 00; on R2{0}\R^2\setminus\{0\} it is a local diffeomorphism everywhere yet 22-to-11 globally. Near (1,0)=f(±(1,0))(1, 0) = f(\pm(1, 0)), the two inverses are the two square-root branches: in complex notation w±ww \mapsto \pm\sqrt w (principal branch), i.e.

(u,v)±(u+u2+v22, v2(u+u2+v2)/2).(u, v) \longmapsto \pm\Bigl(\sqrt{\tfrac{u + \sqrt{u^2+v^2}}{2}},\ \frac{v}{2\sqrt{(u + \sqrt{u^2+v^2})/2}}\Bigr).

(c) Jacobian r>0r > 0: local diffeomorphism on (0,)×R\intoo0\infty\times\R, but θθ+2π\theta \mapsto \theta + 2\pi gives the same point: locally invertible (angle determined up to 2π2\pi on a half-plane), never globally.

Exercise 20.4 ★★

(Folium) Let F(x,y)=x3+y33xyF(x,y) = x^3 + y^3 - 3xy and C=F1(0)\mathcal C = F^{-1}(0). (a) Show that near every point of C\mathcal C other than the origin, C\mathcal C is a 11-submanifold, locally a graph in xx or in yy (which, where?). (b) Compute the tangent line at (32,32)(\frac32, \frac32). (c) What happens at the origin? (Two branches cross: exhibit two C1\mathcal C^1 curves in C\mathcal C through 00 with independent velocities, and conclude that no straightening exists.)

Solution

Solution of Exercise 20.4.

(a) F=3(x2y, y2x)\nabla F = 3(x^2 - y,\ y^2 - x) vanishes iff y=x2y = x^2 and x=y2x = y^2, i.e. x4=xx^4 = x: at (0,0)(0,0) and (1,1)(1,1); only (0,0)(0,0) lies on C\mathcal C (F(1,1)=1F(1,1) = -1). So on C{0}\mathcal C\setminus\{0\}, FF is a submersion: a 11-submanifold, locally a graph y=ψ(x)y = \psi(x) where Fy=3(y2x)0F_y = 3(y^2 - x) \neq 0 and x=χ(y)x = \chi(y) where Fx=3(x2y)0F_x = 3(x^2 - y) \neq 0 (at least one holds off the origin).

(b) At (32,32)(\frac32, \frac32): F=3(9432)(1,1)=94(1,1)\nabla F = 3(\frac94 - \frac32) (1, 1) = \frac94(1,1): tangent line x+y=3x + y = 3.

(c) The rational parametrization x=3t1+t3x = \frac{3t}{1 + t^3}, y=3t21+t3y = \frac{3t^2}{1+t^3} passes through 00 at t=0t = 0 with velocity (3,0)(3, 0); exchanging xyx \leftrightarrow y (the curve is symmetric, or reparametrize by 1/t1/t) gives a second C1\mathcal C^1 curve through 00 with velocity (0,3)(0, 3). Two independent tangent directions are impossible for a 11-submanifold (its tangent space is a line, Proposition 20.6): C\mathcal C is not a submanifold at the origin — a transverse self-crossing.

Exercise 20.5 ★★

Let F(M)=MTMF(M) = M^{\mathsf T}M from Mn(R)M_n(\R) to the space SnS_n of symmetric matrices. (a) Show DF(M)(H)=MTH+HTMDF(M)(H) = M^{\mathsf T}H + H^{\mathsf T}M and that DF(M)DF(M) is surjective onto SnS_n at every MOnM \in O_n (given SSnS \in S_n, try H=12MSH = \frac12MS). (b) Conclude that On=F1(I)O_n = F^{-1}(I) is a compact C\mathcal C^\infty submanifold of dimension n(n1)2\frac{n(n-1)}2, with TIOn={H:HT=H}T_IO_n = \{H : H^{\mathsf T} = -H\}, the antisymmetric matrices. (c) Show that etHOn\eu^{tH} \in O_n for every antisymmetric HH: the tangent directions integrate to curves in the group.

Solution

Solution of Exercise 20.5.

(a) F(M+H)=MTM+MTH+HTM+HTHF(M + H) = M^{\mathsf T}M + M^{\mathsf T}H + H^{\mathsf T}M + H^{\mathsf T}H: DF(M)(H)=MTH+HTMDF(M)(H) = M^{\mathsf T}H + H^{\mathsf T}M. For MOnM \in O_n and SS symmetric, H=12MSH = \frac12MS gives DF(M)(H)=12(S+ST)=SDF(M)(H) = \frac12(S + S^{\mathsf T}) = S: surjective onto SnS_n.

(b) On=F1(I)O_n = F^{-1}(I) with FF a submersion (onto SnS_n, of dimension n(n+1)2\frac{n(n+1)}2) at each of its points: a submanifold of dimension n2n(n+1)2=n(n1)2n^2 - \frac{n(n+1)}2 = \frac{n(n-1)}2. Compact: closed (FF continuous) and bounded (columns are unit vectors). Tangent at II: kerDF(I)={H:H+HT=0}\ker DF(I) = \{H : H + H^{\mathsf T} = 0\}.

(c) (etH)TetH=etHTetH=etHetH=I\bigl(\eu^{tH}\bigr)^{\mathsf T}\eu^{tH} = \eu^{tH^{\mathsf T}}\eu^{tH} = \eu^{-tH}\eu^{tH} = I (transpose passes through the series; the exponentials of commuting matrices multiply, Theorem 19.8).

Exercise 20.6 ★★

(a) Show that det ⁣:Mn(R)R\det \colon M_n(\R) \to \R has differential Ddet(M)(H)=tr(com(M)TH)D\det(M)(H) = \operatorname{tr}\bigl(\operatorname{com}(M) ^{\mathsf T}H\bigr), nonzero at every MSLnM \in SL_n. (b) Conclude that SLn(R)SL_n(\R) is a submanifold of dimension n21n^2 - 1 with TISLn={H:trH=0}T_ISL_n = \{H : \operatorname{tr}H = 0\}. (c) Is GLn(R)GL_n(\R) a submanifold? Of what dimension?

Solution

Solution of Exercise 20.6.

(a) det(M+H)=detMdet(I+M1H)=detM(1+tr(M1H)+O(H2))\det(M + H) = \det M\,\det(I + M^{-1}H) = \det M\bigl(1 + \operatorname{tr}(M^{-1}H) + O(\norm H^2)\bigr) for invertible MM (expansion of det\det near II: the linear term of (1+λi)\prod(1 + \lambda_i)); with detMM1=com(M)T\det M\cdot M^{-1} = \operatorname{com}(M)^{\mathsf T}: Ddet(M)(H)=tr(com(M)TH)D\det(M)(H) = \operatorname{tr}\bigl(\operatorname{com} (M)^{\mathsf T}H\bigr), and the formula extends to all MM by density and continuity. On SLnSL_n, detM=1\det M = 1: Ddet(M)0D\det(M) \ne 0 (its value on H=MH = M is tr(I)=ndetM=n\operatorname{tr}(I)\cdot\dots = n\det M = n).

(b) SLn=det1(1)SL_n = \det^{-1}(1) with det\det a submersion there (values in R\R): dimension n21n^2 - 1; TISLn=kerDdet(I)={H:trH=0}T_ISL_n = \ker D\det(I) = \{H : \operatorname{tr}H = 0\}.

(c) GLnGL_n is an open subset of Mn(R)M_n(\R) (Exercise 6.8): a submanifold of full dimension n2n^2 (straightening: the identity chart).

Exercise 20.7 ★★

By Lagrange multipliers: (a) find the extrema of f(x,y)=xyf(x,y) = xy on the circle x2+y2=1x^2 + y^2 = 1; (b) show that among all probability vectors (p1,,pn)(p_1, \dots, p_n) (positive, summing to 11), the entropy pilnpi-\sum p_i\ln p_i is maximized exactly at the uniform distribution; (c) find the point of the ellipse {x2/4+y2=1}\{x^2/4 + y^2 = 1\} closest to (1,0)(1, 0), and check the multiplier equation geometrically (normal alignment).

Solution

Solution of Exercise 20.7.

(a) (y,x)=λ(2x,2y)(y, x) = \lambda(2x, 2y) and x2+y2=1x^2 + y^2 = 1: y=2λxy = 2\lambda x, x=2λyx = 2\lambda y give x2=y2=12x^2 = y^2 = \frac12. Values of xyxy: ±12\pm\frac12: maximum 12\frac12 at ±12(1,1)\pm\frac1{\sqrt2}(1,1), minimum 12-\frac12 at ±12(1,1)\pm\frac1{\sqrt2}(1,-1) (the constraint set is compact: extrema exist).

(b) On the interior of the simplex (pi>0p_i > 0), Lagrange for H(p)=pilnpiH(p) = -\sum p_i\ln p_i with constraint pi=1\sum p_i = 1: lnpi1=λ-\ln p_i - 1 = \lambda for all ii: all pip_i equal, pi=1np_i = \frac1n, with H=lnnH = \ln n. The maximum over the compact simplex is attained; if it were attained on the boundary (some pi=0p_i = 0), the distribution lives on n1\leq n - 1 points and by induction Hln(n1)<lnnH \leq \ln(n-1) < \ln n: the interior critical point is the global maximum — uniform ignorance maximizes entropy.

(c) Minimize (x1)2+y2(x-1)^2 + y^2 on the compact ellipse: (2(x1),2y)=λ(x2,2y)(2(x{-}1), 2y) = \lambda(\frac x2, 2y). If y0y \neq 0: λ=1\lambda = 1, then 2(x1)=x22(x - 1) = \frac x2 gives x=43x = \frac43, y2=149=59y^2 = 1 - \frac49 = \frac59: distance2^2 =19+59=23= \frac19 + \frac59 = \frac23. If y=0y = 0: x=±2x = \pm2, distances 11 and 33. Closest points: (43,±53)\bigl(\frac43, \pm\frac{\sqrt5}3\bigr), at distance 2/3<1\sqrt{2/3} < 1. The multiplier equation says the segment from (1,0)(1,0) to the closest point is parallel to \nabla(ellipse): it meets the ellipse orthogonally, as geometry demands.

Exercise 20.8 ★★★

Write out Example 20.8 in full: prove by induction that a real symmetric matrix admits an orthonormal basis of eigenvectors, with λ1λn\lambda_1 \geq \dots \geq \lambda_n the successive constrained maxima of the Rayleigh quotient. Then deduce the Courant–Fischer formulas of Exercise 15.8 in finite dimension directly from this construction.

Solution

Solution of Exercise 20.8.

Induction on nn; n=1n = 1 trivial. The Rayleigh function f(x)=Ax,xf(x) = \langle Ax, x\rangle attains its maximum λ1\lambda_1 on the compact Sn1S^{n-1} at some v1v_1; Lagrange (Theorem 20.7, sphere as level set) gives 2Av1=2λv12Av_1 = 2\lambda v_1, and λ=Av1,v1=λ1\lambda = \langle Av_1, v_1\rangle = \lambda_1. The hyperplane v1v_1^\perp is AA-invariant (symmetry: Av,v1=v,Av1=0\langle Av, v_1\rangle = \langle v, Av_1\rangle = 0); the restriction is symmetric, and induction yields an orthonormal eigenbasis v2,,vnv_2, \dots, v_n of v1v_1^\perp with eigenvalues λ2λn\lambda_2 \geq \dots \geq \lambda_n, each the maximum of ff on the sphere of the remaining orthocomplement. Courant–Fischer follows exactly as in Exercise 15.8: expand x=civix = \sum c_iv_i; on a kk-dimensional test space intersect with Vect(vk,,vn)\operatorname{Vect}(v_k, \dots, v_n) (dimension count in Rn\R^n) to get minλk\min \leq \lambda_k, and Vect(v1,,vk)\operatorname{Vect}(v_1, \dots, v_k) achieves min=λk\min = \lambda_k.

Exercise 20.9 ★★★

(Hadamard’s inequality) For MGLn(R)M \in GL_n(\R) with columns c1,,cnc_1, \dots, c_n:

detM    i=1nci2,\abs{\det M} \;\leq\; \prod_{i=1}^n\norm{c_i}_2 ,

with equality iff the columns are orthogonal. (Reduce to columns of norm 11 by scaling; maximize det\det on the compact product of spheres (Sn1)n(S^{n-1})^n; at a maximizer, Lagrange in each column separately gives cidet=λici\nabla_{c_i}\det = \lambda_ic_i, and cidet\nabla_{c_i}\det is the ii-th column of com(M)\operatorname{com}(M): deduce MTMM^{\mathsf T}M diagonal, hence =I= I, hence det=±1\det = \pm1.) Geometric reading: the volume of a parallelepiped is at most the product of its edge lengths.

Solution

Solution of Exercise 20.9.

Scaling each column to unit norm divides det\abs{\det} by ci\prod\norm{c_i}: it suffices to prove detM1\abs{\det M} \leq 1 when all columns are unit, with equality iff MOnM \in O_n. The function det\det is continuous on the compact (Sn1)n(S^{n-1})^n: it attains a maximum mdetI=1>0m \geq \det I = 1 > 0 at some MM. Fixing all columns but the ii-th, det\det is linear in cic_i with gradient the ii-th column of com(M)\operatorname{com}(M); Lagrange on the ii-th sphere: com(M)i=λici\operatorname{com}(M)_{\cdot i} = \lambda_i c_i. The identity MTcom(M)=det(M)IM^{\mathsf T}\operatorname{com}(M) = \det(M)\,I reads cj,com(M)i=det(M)δij\langle c_j, \operatorname{com}(M)_{\cdot i}\rangle = \det(M)\,\delta_{ij}, i.e. λicj,ci=det(M)δij\lambda_i\langle c_j, c_i\rangle = \det(M)\delta_{ij}; taking j=ij = i: λi=detM=m0\lambda_i = \det M = m \neq 0, and then jij \neq i gives ci,cj=0\langle c_i, c_j\rangle = 0: the columns are orthonormal, MOnM \in O_n, m=detM=1m = \abs{\det M} = 1. Hence detci\abs{\det} \leq \prod\norm{c_i} always, with equality exactly for orthogonal columns (rescale back): a parallelepiped’s volume is largest, for given edge lengths, when the edges are perpendicular.

Exercise 20.10 ★★

Near which of its points is the circle S1S^1 a graph y=ψ(x)y = \psi(x)? A graph x=χ(y)x = \chi(y)? Verify the graph characterization (Theorem 20.3(3)) explicitly at (1,0)(1, 0), and explain in one sentence why some coordinate permutation is always sufficient but no single one always works.

Solution

Solution of Exercise 20.10.

y=±1x2y = \pm\sqrt{1 - x^2} works near every point with y0y \neq 0; x=±1y2x = \pm\sqrt{1 - y^2} near every point with x0x \neq 0; at (1,0)(1, 0): the graph x=1y2x = \sqrt{1 - y^2} over y(1,1)y \in \intoo{-1}1, which is Theorem 20.3(3) with the coordinates swapped. Some permutation always works because the tangent line, being one-dimensional, cannot be simultaneously vertical and horizontal — but it can be either, so no fixed choice of “dependent” coordinate serves at every point.

Exercise 20.11 ★★

(The orthogonal group as a submanifold, quantitatively) (a) Show that On={M:MTM=I}O_n = \{M : M^{\mathsf T}M = I\} is compact: bounded (each column is a unit vector, so Mn\norm M \leq \sqrt n for the Euclidean matrix norm) and closed. (b) Show that its tangent space at II is the space of antisymmetric matrices, of dimension n(n1)2\frac{n(n-1)}2, and at a general AOnA \in O_n: TAOn={AK:KT=K}T_AO_n = \{AK : K^{\mathsf T} = -K\}. (c) Deduce that the map tAexp(tK)t \mapsto A\exp(tK) is, for each antisymmetric KK, a curve in OnO_n through AA with velocity AKAK (verify exp(tK)On\exp(tK) \in O_n using exp(X)T=exp(XT)\exp(X)^{\mathsf T} = \exp(X^{\mathsf T}) and exp(X)exp(X)=I\exp(-X)\exp(X) = I): every tangent vector is realized by an explicit curve, with no implicit function theorem needed.

Solution

Solution of Exercise 20.11.

(a) The defining map F(M)=MTMIF(M) = M^{\mathsf T}M - I is continuous: On=F1(0)O_n = F^{-1}(0) is closed; columns of an orthogonal matrix are unit vectors, so the Euclidean (Frobenius) norm is exactly n\sqrt n: bounded. Compact by Heine–Borel in Mn(R)Rn2M_n(\R) \cong \R^{n^2}.

(b) OnO_n is the level set F=0F = 0 studied in the chapter: DF(A)H=ATH+HTADF(A)H = A^{\mathsf T}H + H^{\mathsf T}A, surjective onto symmetric matrices at each AOnA \in O_n (given symmetric SS, take H=12ASH = \frac12AS), so OnO_n is a submanifold of dimension n2n(n+1)2=n(n1)2n^2 - \frac{n(n+1)}2 = \frac{n(n-1)}2 with

TAOn=kerDF(A)={H:ATH antisymmetric}={AK:KT=K};T_AO_n = \ker DF(A) = \{H : A^{\mathsf T}H \text{ antisymmetric}\} = \{AK : K^{\mathsf T} = -K\} ;

at A=IA = I these are the antisymmetric matrices.

(c) exp(tK)Texp(tK)=exp(tKT)exp(tK)=exp(tK)exp(tK)=I\exp(tK)^{\mathsf T}\exp(tK) = \exp(tK^{\mathsf T}) \exp(tK) = \exp(-tK)\exp(tK) = I (the two matrices ±tK\pm tK commute, so the product of exponentials is the exponential of the sum): exp(tK)On\exp(tK) \in O_n, and γ(t)=Aexp(tK)\gamma(t) = A\exp(tK) is a curve in OnO_n with γ(0)=A\gamma(0) = A, γ(0)=AK\gamma'(0) = AK. As KK runs over antisymmetric matrices, AKAK sweeps TAOnT_AO_n: the exponential realizes the whole tangent space by explicit curves — the Lie-group shortcut that Problem 20.1 exploits for SO(3)SO(3).

Exercise 20.12 ★★

(Critical points of the distance) Let MRnM \subseteq \R^n be a submanifold and pMp \notin M. Show that if x0Mx_0 \in M minimizes the distance to pp (such a point exists when MM is closed and nonempty — why?), then

px0    Tx0Mp - x_0 \;\perp\; T_{x_0}M

(differentiate tγ(t)p2t \mapsto \norm{\gamma(t) - p}^2 along curves in MM). Deduce: the closest point on a sphere lies on the ray through the center; and use the condition to compute the distance from p=(2,0)p = (2, 0) to the parabola y=x2y = x^2 (reduce to a cubic and solve it numerically to three digits).

Solution

Solution of Exercise 20.12.

Existence: intersect MM with a large closed ball around pp to get a nonempty compact; the continuous distance attains its minimum there, and points outside the ball are farther. First-order condition: for a curve γ\gamma in MM with γ(0)=x0\gamma(0) = x_0, the function h(t)=γ(t)p2h(t) = \norm{\gamma(t) - p}^2 is differentiable with a minimum at 00:

0=h(0)=2γ(0), x0p,0 = h'(0) = 2\,\langle\gamma'(0),\ x_0 - p\rangle,

and γ(0)\gamma'(0) sweeps Tx0MT_{x_0}M: px0Tx0Mp - x_0 \perp T_{x_0}M. Sphere S(c,r)S(c, r): the tangent space at x0x_0 is (x0c)(x_0 - c)^\perp, so px0x0cp - x_0 \parallel x_0 - c: x0x_0 lies on the line through cc and pp, at distance rr from cc — the ray point, as geometry insists. Parabola: at x0=(x,x2)x_0 = (x, x^2) the tangent is spanned by (1,2x)(1, 2x); orthogonality to px0=(2x,x2)p - x_0 = (2 - x, -x^2) reads

(2x)2x3=0,i.e.2x3+x2=0,(2 - x) - 2x^3 = 0, \qquad\text{i.e.}\qquad 2x^3 + x - 2 = 0,

with unique real root (x2x3+xx \mapsto 2x^3 + x is strictly increasing) x0.835x \approx 0.835; then x0(0.835,0.698)x_0 \approx (0.835, 0.698) and d(p,M)=(20.835)2+0.69821.358d(p, M) = \sqrt{(2 - 0.835)^2 + 0.698^2} \approx 1.358.

20.5 Problem: SO(3)SO(3) and the quaternions

Problem 20.1

Weekend problem — rotations, the group S3S^3, and the double cover

The quaternions H={t+xi+yj+zk}\mathbb H = \{t + x\mathrm i + y\mathrm j + z\mathrm k\} — the algebra whose unit group contains Problem 1.1’s Q8Q_8 — parametrize three-dimensional rotations twice over: the map “conjugate by a unit quaternion” is a surjective morphism S3SO(3)S^3 \to SO(3) with kernel {±1}\{\pm1\}. We build everything. Recall/define: multiplication is R\R-bilinear with i2=j2=k2=ijk=1\mathrm i^2 = \mathrm j^2 = \mathrm k^2 = \mathrm{ijk} = -1; the conjugate of q=t+xi+yj+zkq = t + x\mathrm i + y\mathrm j + z\mathrm k is qˉ=txiyjzk\bar q = t - x\mathrm i - y\mathrm j - z\mathrm k; N(q)=qqˉ=t2+x2+y2+z2N(q) = q\bar q = t^2 + x^2 + y^2 + z^2.

Part I — The algebra H\mathbb H and the group S3S^3.

  1. Verify that H\mathbb H is an associative R\R-algebra with center R\R, that pq=qˉpˉ\overline{pq} = \bar q\,\bar p, and that N(pq)=N(p)N(q)N(pq) = N(p)N(q) (one clean route: represent qq as the 2×22\times2 complex matrix (αββˉαˉ)\bigl(\begin{smallmatrix}\alpha & \beta\\ -\bar\beta & \bar\alpha\end{smallmatrix}\bigr), q=α+βjq = \alpha + \beta\mathrm j, and use det\det).
  2. Deduce that every q0q \neq 0 is invertible (q1=qˉ/N(q)q^{-1} = \bar q/N(q)): H\mathbb H is a (noncommutative) field, and S3={N(q)=1}S^3 = \{N(q) = 1\} is a group — and a compact 33-submanifold of R4\R^4 (Example 20.4).

Part II — The rotation morphism. Identify R3\R^3 with the pure quaternions P={xi+yj+zk}P = \{x\mathrm i + y\mathrm j + z\mathrm k\}, and for qS3q \in S^3 define ρq(v)=qvqˉ\rho_q(v) = q\,v\,\bar q.

  1. Show that ρq\rho_q maps PP to PP (pure quaternions are those with vˉ=v\bar v = -v), is R\R-linear, preserves the norm, and that ρ ⁣:qρq\rho \colon q \mapsto \rho_q is a group morphism S3O(3)S^3 \to O(3).
  2. Compute the kernel: ρq=id\rho_q = \mathrm{id} iff qq commutes with i,j,k\mathrm i, \mathrm j, \mathrm k iff qRS3={±1}q \in \R\cap S^3 = \{\pm1\}.
  3. Write q=cosθ2+sinθ2uq = \cos\frac\theta2 + \sin\frac\theta2\,u with uPu \in P, N(u)=1N(u) = 1 (why is this always possible for qS3q \in S^3?). Show that ρq\rho_q fixes uu and, on the plane uPu^\perp\cap P, acts as the rotation of angle θ\theta (compute ρq(w)\rho_q(w) for wuw \perp u using uw=wuuw = -wu for orthogonal pure units — prove this identity from the multiplication table, or from uw+wu=2u,wuw + wu = -2\langle u, w\rangle).
  4. Conclude: imρSO(3)\operatorname{im}\rho \subseteq SO(3) (each ρq\rho_q is a rotation with axis and angle as computed — determinant +1+1 by continuity of qdetρqq \mapsto \det\rho_q on the connected S3S^3, or directly), and ρ\rho is onto SO(3)SO(3): every rotation of R3\R^3 has an axis (prove: a real 3×33\times3 orthogonal matrix with det=1\det = 1 has eigenvalue 11 — consider the characteristic polynomial) and is therefore some ρq\rho_q. Summary:

    SO(3)    S3/{±1}.SO(3) \;\cong\; S^3/\{\pm 1\} .

Part III — SO(3)SO(3) as a submanifold; Rodrigues.

  1. Show that SO(3)SO(3) is a compact 33-dimensional submanifold of M3(R)M_3(\R) with TISO(3)=T_ISO(3) = antisymmetric matrices (Exercise 20.5; the determinant condition selects a union of components).
  2. For the antisymmetric matrix AuA_u associated with uR3u \in \R^3 (Auv=uvA_uv = u\wedge v, the cross product), prove Rodrigues’ formula:

    eθAu=I+sinθAu+(1cosθ)Au2(u=1)\eu^{\theta A_u} = I + \sin\theta\,A_u + (1 - \cos\theta)\,A_u^2 \qquad (\norm u = 1)

    (from Au3=AuA_u^3 = -A_u: split the exponential series along AuA_u’s powers), and identify it as the rotation of axis uu and angle θ\theta. Deduce that exp\exp maps the antisymmetric matrices onto SO(3)SO(3).

  3. Relate the two parametrizations: show that tρq(t)t \mapsto \rho_{q(t)} with q(t)=cost2+sint2uq(t) = \cos\frac t2 + \sin\frac t2\,u is a one-parameter group of rotations whose derivative at t=0t = 0 is AuA_u — the quaternionic and matrix exponentials tell the same story at half and full speed respectively.

Part IV — The double cover, felt.

  1. Show that the path q(t)=cost2+sint2kq(t) = \cos\frac t2 + \sin\frac t2\,\mathrm k, t[0,2π]t \in \intcc0{2\pi}, is a loop in SO(3)SO(3) (its image ρq(t)\rho_{q(t)} returns to the identity) whose quaternionic lift is not a loop: q(2π)=q(0)q(2\pi) = -q(0). Continuing to t=4πt = 4\pi closes the lift. Explain in a short paragraph what this says: a 2π2\pi rotation is not continuously undoable while a 4π4\pi rotation is (the belt trick), because SO(3)SO(3)’s loops are detected in its double cover S3S^3.
  2. Deduce also the practical dividend: composition of rotations = multiplication of quaternions (44 multiplications’ worth of data instead of 99, no drift from orthogonality) — verify on the composition of two quarter-turns about i\mathrm i and j\mathrm j: compute the axis and angle of the product.

Part V — The explicit matrix: Euler–Rodrigues. Write q=a+bi+cj+dkS3q = a + b\mathrm i + c\mathrm j + d\mathrm k \in S^3, so that a2+b2+c2+d2=1a^2 + b^2 + c^2 + d^2 = 1.

  1. Compute ρq(i)\rho_q(\mathrm i) in full from the multiplication table; then obtain ρq(j)\rho_q(\mathrm j) and ρq(k)\rho_q(\mathrm k) by the cyclic substitution ijki\mathrm i \to \mathrm j \to \mathrm k \to \mathrm i, (b,c,d)(c,d,b)(b, c, d) \to (c, d, b) (justify it: cycling i,j,k\mathrm i, \mathrm j, \mathrm k extends to an automorphism of H\mathbb H, because the defining relations are cyclically symmetric). Conclude that the matrix of ρq\rho_q in the basis (i,j,k)(\mathrm i, \mathrm j, \mathrm k) is the Euler–Rodrigues matrix

    Rq=(a2+b2c2d22(bcad)2(bd+ac)2(bc+ad)a2b2+c2d22(cdab)2(bdac)2(cd+ab)a2b2c2+d2).R_q = \begin{pmatrix} a^2 + b^2 - c^2 - d^2 & 2(bc - ad) & 2(bd + ac)\\ 2(bc + ad) & a^2 - b^2 + c^2 - d^2 & 2(cd - ab)\\ 2(bd - ac) & 2(cd + ab) & a^2 - b^2 - c^2 + d^2 \end{pmatrix}.
  2. (Reading a rotation backwards) Show that

    trRq=4a21=1+2cosθ,12(RqRqT)=sinθAu,\operatorname{tr}R_q = 4a^2 - 1 = 1 + 2\cos\theta, \qquad \tfrac12\bigl(R_q - R_q^{\mathsf T}\bigr) = \sin\theta\,A_u,

    in the notation of questions 5 and 8. Deduce an algorithm recovering ±q\pm q from a rotation matrix RR: the angle from the trace; the axis from the antisymmetric part when 0<θ<π0 < \theta < \pi; and, when θ=π\theta = \pi, prove and use the identity R+I=2uuTR + I = 2\,uu^{\mathsf T}.

  3. Evaluate RqR_q for question 11’s product q=12(1+i+j+k)q = \frac12(1 + \mathrm i + \mathrm j + \mathrm k): a permutation matrix appears. Identify the rotation and reconcile with the axis and angle found in question 11.

Part VI — Inside S3S^3: SU(2)SU(2), conjugacy classes, exponentials.

  1. Show that the matrix representation of question 1 (call it Φ\Phi) restricts to a group isomorphism from S3S^3 onto the special unitary group

    SU(2)={UM2(C):UU=I, detU=1}SU(2) = \bigl\{U \in M_2(\C) : U^*U = I,\ \det U = 1\bigr\}

    (for surjectivity, write out the equations U1=UU^{-1} = U^* and detU=1\det U = 1 for a general 2×22\times2 complex matrix).

  2. Show that the real part is a conjugation invariant on S3S^3Re(pqpˉ)=Req\operatorname{Re}(pq\bar p) = \operatorname{Re}q for all pS3p \in S^3 — and, conversely, that two unit quaternions with the same real part are conjugate in S3S^3 (reduce to moving one unit pure axis onto another, which Part II provides). Describe the conjugacy classes of S3S^3 geometrically; translate into SU(2)SU(2) (level sets of the trace); and project by ρ\rho: two rotations are conjugate in SO(3)SO(3) if and only if they have the same angle θ[0,π]\theta \in \intcc0\pi.
  3. Define exp\exp on H\mathbb H by the exponential series; check absolute convergence, using pq=pq\abs{pq} = \abs p\,\abs q for q=N(q)\abs q = \sqrt{N(q)}. Show, for a unit pure uu and θR\theta \in \R,

    exp(θu)=cosθ+sinθu,\exp(\theta u) = \cos\theta + \sin\theta\,u ,

    deduce that exp\exp maps the hyperplane PP onto S3S^3, and check that ρexp(su)=e2sAu\rho_{\exp(su)} = \eu^{2sA_u}: the half-angle phenomenon of question 9 again.

  4. For pure quaternions v,wv, w prove the product rule vw=v,w+vwvw = -\langle v, w\rangle + v\wedge w, hence the commutator identity vwwv=2vwvw - wv = 2\,v\wedge w; prove also [Av,Aw]=Avw[A_v, A_w] = A_{v\wedge w} for the matrices of question 8. Conclude that the derivative of ρ\rho at 11 along the curves texp(tv)t \mapsto \exp(tv) is the linear isomorphism v2Avv \mapsto 2A_v from PP onto the antisymmetric matrices, and that it transports the quaternion commutator to the matrix commutator.

Part VII — Global structure.

  1. (No continuous section) Suppose s ⁣:SO(3)S3s \colon SO(3) \to S^3 is continuous with ρs=id\rho \circ s = \operatorname{id}. For the loop R(t)=ρq(t)R(t) = \rho_{q(t)} of question 10, set ε(t)=s(R(t))q(t)1\varepsilon(t) = s(R(t))\,q(t)^{-1} for t[0,2π]t \in \intcc0{2\pi}. Show that ε\varepsilon is continuous with values in {±1}\{\pm1\}, and derive a contradiction: there is no continuous global choice of a unit quaternion representing each rotation.
  2. (The ball model) Let BˉR3\bar B \subseteq \R^3 be the closed ball of radius π\pi and E(v)=eAvE(v) = \eu^{A_v}, with E(0)=IE(0) = I. Show that EE maps Bˉ\bar B onto SO(3)SO(3), is injective on the open ball, and on the boundary sphere identifies exactly antipodes: E(πu)=E(πu)=2uuTIE(\pi u) = E(-\pi u) = 2uu^{\mathsf T} - I, with no other coincidences. Thus SO(3)SO(3) is the ball with antipodal boundary points glued — the projective space RP3\mathbb{RP}^3 — and a diameter becomes question 10’s non-contractible loop.
  3. Show that ρpρqρp1=ρpqpˉ\rho_p\rho_q\rho_p^{-1} = \rho_{pq\bar p}; that the involutions of SO(3)SO(3) (the RIR \neq I with R2=IR^2 = I) are exactly the half-turns ρw\rho_w with ww a unit pure quaternion; and that the center of SO(3)SO(3) is trivial.
  4. Show that every rotation is a product of two half-turns: for q=cosθ2+sinθ2uq = \cos\frac\theta2 + \sin\frac\theta2\,u, choose a unit pure wuw \perp u, check that w=qww' = qw is again a unit pure quaternion, and verify ρq=ρwρw\rho_q = \rho_{w'}\rho_w. Where do the two axes lie, and what angle do they make?
  5. Conclude the topological summary: SO(3)SO(3) is compact and path-connected (give two proofs: continuous image of S3S^3 under ρ\rho; image of exp\exp), while O(3)O(3) has exactly two connected components, each homeomorphic to SO(3)SO(3).
  6. (A composition, three ways) Let R1R_1 be the rotation by π2\frac\pi2 about the zz-axis and R2R_2 the rotation by π2\frac\pi2 about the xx-axis. Compute the axis and angle of R2R1R_2R_1: (i) by multiplying the two 3×33\times3 matrices and using trace/antisymmetric part (Part V); (ii) by multiplying the corresponding unit quaternions q2q1q_2q_1. Check the two answers agree: angle 2π3\frac{2\pi}3, axis 13(1,1,1)\frac1{\sqrt3}(1, -1, 1).
  7. (The Cayley transform) For KK antisymmetric, show that I+KI + K is invertible and

    C(K)=(IK)(I+K)1SO(n),C(K) = (I - K)(I + K)^{-1} \in SO(n),

    with 1-1 never an eigenvalue of C(K)C(K); show that KC(K)K \mapsto C(K) is a bijection from antisymmetric matrices onto {RSO(n):1SpR}\{R \in SO(n) : -1 \notin \operatorname{Sp}R\}, with inverse R(IR)(I+R)1R \mapsto (I - R)(I + R)^{-1}. (A rational chart of SO(n)SO(n), companion to the transcendental exp\exp of Exercise 20.11.)

Solution

Solution of Problem 20.1.

1. Map q=t+xi+yj+zk(αββˉαˉ)q = t + x\mathrm i + y\mathrm j + z\mathrm k \mapsto \bigl(\begin{smallmatrix}\alpha & \beta\\ -\bar\beta & \bar\alpha\end{smallmatrix}\bigr) with α=t+ix\alpha = t + \iu x, β=y+iz\beta = y + \iu z: one checks that 1,i,j,k1, \mathrm i, \mathrm j, \mathrm k go to II, (i00i)\bigl(\begin{smallmatrix} \iu & 0\\ 0 & -\iu\end{smallmatrix}\bigr), (0110)\bigl(\begin{smallmatrix}0 & 1\\ -1 & 0\end{smallmatrix}\bigr), (0ii0)\bigl(\begin{smallmatrix}0 & \iu\\ \iu & 0\end{smallmatrix}\bigr), whose products reproduce the quaternion table: the map is an injective algebra morphism, so H\mathbb H inherits associativity; N(q)=α2+β2=detN(q) = \abs\alpha^2 + \abs\beta^2 = \det is multiplicative, and conjugation corresponds to the adjugate-transpose, giving pq=qˉpˉ\overline{pq} = \bar q\bar p. Center: commuting with i\mathrm i forces y=z=0y = z = 0, with j\mathrm j forces x=0x = 0: R\R.

2. qqˉ=N(q)q\bar q = N(q): for q0q \neq 0, q1=qˉ/N(q)q^{-1} = \bar q/N(q): a division algebra. On S3S^3: N(pq)=1N(pq) = 1 and N(q1)=1N(q^{-1}) = 1: a group; and S3R4S^3 \subseteq \R^4 is the unit sphere: a compact 33-submanifold.

3. vv is pure iff vˉ=v\bar v = -v; then qvqˉ=qvˉqˉ=qvqˉ\overline{qv\bar q} = q\bar v\bar q = -qv\bar q: ρq\rho_q preserves PP. Linearity is clear; N(qvqˉ)=N(q)N(v)N(q)=N(v)N(qv\bar q) = N(q)N(v)N(q) = N(v): an isometry of (P,N)(R3,2)(P, N) \cong (\R^3, \norm\cdot^2): ρqO(3)\rho_q \in O(3). And ρpq(v)=pqvpq=p(qvqˉ)pˉ=ρp(ρq(v))\rho_{pq}(v) = pqv\overline{pq} = p(qv\bar q)\bar p = \rho_p(\rho_q(v)): a morphism.

4. ρq=id\rho_q = \mathrm{id} iff qv=vqqv = vq for all pure vv, iff qq commutes with i,j,k\mathrm i, \mathrm j, \mathrm k, iff qq is central (question 1): qRS3={±1}q \in \R\cap S^3 = \{\pm1\}.

5. Write q=t+pq = t + p (tRt \in \R, pp pure): 1=N(q)=t2+N(p)1 = N(q) = t^2 + N(p), so t=cosθ2t = \cos\frac\theta2 and p=sinθ2up = \sin\frac\theta2\,u with N(u)=1N(u) = 1 for some θ\theta (if p=0p = 0, q=±1q = \pm1 acts trivially). Since u2=N(u)=1u^2 = -N(u) = -1, qq and uu commute, and ρq(u)=quqˉ=uqqˉ=u\rho_q(u) = qu\bar q = uq\bar q = u: the axis. For pure units wuw \perp u: the product rule vw=v,w+vwvw = -\langle v, w\rangle + v\wedge w (expand in coordinates from the table) gives uw=uw=wuuw = u\wedge w = -wu. Then

ρq(w)=(cosθ2+sinθ2u)w(cosθ2sinθ2u)=cosθw+sinθ(uw),\rho_q(w) = \bigl(\cos\tfrac\theta2 + \sin\tfrac\theta2u\bigr)\,w\,\bigl(\cos\tfrac\theta2 - \sin\tfrac\theta2u\bigr) = \cos\theta\,w + \sin\theta\,(u\wedge w),

using uwu=u2w=wuwu = -u^2w = w and the double-angle formulas: the rotation of angle θ\theta in the oriented plane (w,uw)(w, u\wedge w).

6. Each ρq\rho_q is a rotation about uu by θ\theta: in the orthonormal basis (u,w,uw)(u, w, u\wedge w) its matrix has determinant +1+1: imρSO(3)\operatorname{im}\rho \subseteq SO(3). Surjectivity: a matrix RSO(3)R \in SO(3) has 11 as an eigenvalue, since

det(RI)=detRdet(IRT)=det(IR)=(1)3det(RI),\det(R - I) = \det R\,\det(I - R^{\mathsf T}) = \det(I - R) = (-1)^3\det(R - I),

so det(RI)=0\det(R - I) = 0. Take a unit eigenvector uu; RR preserves uu^\perp and restricts there to a rotation of some angle θ\theta (planar orthogonal, determinant 11): R=ρqR = \rho_q for q=cosθ2+sinθ2uq = \cos\frac\theta2 + \sin\frac\theta2\,u. With question 4 and the first isomorphism theorem (Theorem 1.3): SO(3)S3/{±1}SO(3) \cong S^3/\{\pm1\}.

7. O3O_3 is a compact 33-dimensional submanifold (Exercise 20.5); det\det is continuous on it with values in {±1}\{\pm1\}, so SO(3)=O3{det=1}SO(3) = O_3\cap\{\det = 1\} is open and closed in O3O_3: a union of connected components, hence itself a compact 33-submanifold, with the same tangent space at II: the antisymmetric matrices.

8. Au2v=u(uv)=u,vuvA_u^2v = u\wedge(u\wedge v) = \langle u, v\rangle u - v (unit uu), so Au3v=u(u,vuv)=uvA_u^3v = u\wedge(\langle u,v\rangle u - v) = -u\wedge v: Au3=AuA_u^3 = -A_u. Splitting the exponential series by residues of powers mod the relation A3=AA^3 = -A:

eθAu=I+(θθ33!+)Au+(θ22!θ44!+)Au2=I+sinθAu+(1cosθ)Au2.\eu^{\theta A_u} = I + \Bigl(\theta - \frac{\theta^3}{3!} + \cdots\Bigr)A_u + \Bigl(\frac{\theta^2}{2!} - \frac{\theta^4}{4!} + \cdots\Bigr)A_u^2 = I + \sin\theta\,A_u + (1 - \cos\theta)\,A_u^2 .

On uu: Auu=0A_uu = 0: fixed. On wuw \perp u: eθAuw=w+sinθuw+(1cosθ)(w)=cosθw+sinθuw\eu^{\theta A_u}w = w + \sin\theta\,u\wedge w + (1 - \cos\theta)(-w) = \cos\theta\,w + \sin\theta\,u\wedge w: the rotation of axis uu, angle θ\theta — Rodrigues. Every rotation has this form (question 6): exp\exp is onto SO(3)SO(3) from the antisymmetric matrices.

9. With q(t)=cost2+sint2uq(t) = \cos\frac t2 + \sin\frac t2\,u: question 5 shows ρq(t)\rho_{q(t)} is the rotation of axis uu and angle tt, i.e. ρq(t)=etAu\rho_{q(t)} = \eu^{tA_u}, whose derivative at t=0t = 0 is AuA_u. The quaternion runs at half the angle — the analytic trace of the double cover.

10. ρq(t)\rho_{q(t)} is the rotation about k\mathrm k by angle tt: at t=2πt = 2\pi it returns to the identity — a loop in SO(3)SO(3). Its lift satisfies q(2π)=cosπ=1=q(0)q(2\pi) = \cos\pi = -1 = -q(0): the lifted path is not closed; only at t=4πt = 4\pi does qq return to 11. Interpretation: the loop of full rotations is not contractible in SO(3)SO(3) — its lift ends at the other sheet of the cover — while the double loop is; a body attached to its surroundings by straps (the belt trick) returns to an untwisted state after 4π4\pi but not after 2π2\pi. Rotation groups remember the parity of full turns; S3S^3, being simply connected, is where that memory lives.

11. Quarter turns: qi=cosπ4+sinπ4iq_{\mathrm i} = \cos\frac\pi4 + \sin\frac\pi4\,\mathrm i, qj=cosπ4+sinπ4jq_{\mathrm j} = \cos\frac\pi4 + \sin\frac\pi4\,\mathrm j. Product (applying the j\mathrm j-turn first):

qiqj=12(1+i)(1+j)=12(1+i+j+k),q_{\mathrm i}q_{\mathrm j} = \tfrac12(1 + \mathrm i)(1 + \mathrm j) = \tfrac12\bigl(1 + \mathrm i + \mathrm j + \mathrm k\bigr),

of norm 11, with cosθ2=12\cos\frac\theta2 = \frac12: θ=2π3\theta = \frac{2\pi}3, and axis u=i+j+k3u = \frac{\mathrm i + \mathrm j + \mathrm k}{\sqrt3} (the pure part normalized). Two successive quarter-turns about orthogonal axes make a 120120^\circ rotation about the cube’s main diagonal — four real multiplications’ worth of bookkeeping, orthogonality preserved exactly: why flight software and graphics engines compose rotations through quaternions.

12. From the table, ji=k\mathrm{ji} = -\mathrm k and ki=j\mathrm{ki} = \mathrm j, so

qi=aib+c(ji)+d(ki)=b+ai+djck.q\,\mathrm i = a\mathrm i - b + c(\mathrm{ji}) + d(\mathrm{ki}) = -b + a\mathrm i + d\mathrm j - c\mathrm k .

Multiplying by qˉ=abicjdk\bar q = a - b\mathrm i - c\mathrm j - d\mathrm k with the scalar–vector rule (t1+p1)(t2+p2)=t1t2p1,p2+t1p2+t2p1+p1p2(t_1 + p_1)(t_2 + p_2) = t_1t_2 - \langle p_1, p_2\rangle + t_1p_2 + t_2p_1 + p_1\wedge p_2, where p1=(a,d,c)p_1 = (a, d, -c) and p2=(b,c,d)p_2 = (-b, -c, -d): the scalar part is ab(ab)=0-ab - (-ab) = 0 (pure, as it must be), and the vector part is

b(b,c,d)+a(a,d,c)+(c2d2, bc+ad, bdac)=(a2+b2c2d2, 2(bc+ad), 2(bdac)):b(b, c, d) + a(a, d, -c) + (-c^2 - d^2,\ bc + ad,\ bd - ac) = \bigl(a^2 + b^2 - c^2 - d^2,\ 2(bc + ad),\ 2(bd - ac)\bigr):

the first column of RqR_q. The cyclic map σ(i)=j\sigma(\mathrm i) = \mathrm j, σ(j)=k\sigma(\mathrm j) = \mathrm k, σ(k)=i\sigma(\mathrm k) = \mathrm i preserves the relations i2=j2=k2=ijk=1\mathrm i^2 = \mathrm j^2 = \mathrm k^2 = \mathrm{ijk} = -1 (the word ijk\mathrm{ijk} is cyclically invariant up to the relation ijk=jki\mathrm{ijk} = \mathrm{jki}, which holds in any ring: conjugating ijk=1\mathrm{ijk} = -1 by the invertible i\mathrm i), so σ\sigma extends to an R\R-algebra automorphism, and σ(ρq(v))=ρσ(q)(σ(v))\sigma(\rho_q(v)) = \rho_{\sigma(q)}(\sigma(v)). Unwinding, the image of j\mathrm j is the first-column formula after the substitution (b,c,d)(c,d,b)(b, c, d) \to (c, d, b) with the basis relabeled ijki\mathrm i \to \mathrm j \to \mathrm k \to \mathrm i, which is exactly the second column displayed; one more turn gives the third.

13. Summing the diagonal, trRq=3a2(b2+c2+d2)=4a21\operatorname{tr}R_q = 3a^2 - (b^2 + c^2 + d^2) = 4a^2 - 1 (unit norm), and with a=cosθ2a = \cos\frac\theta2: 4cos2θ21=1+2cosθ4\cos^2\frac\theta2 - 1 = 1 + 2\cos\theta. Antisymmetric part: the three independent entries of RqRqTR_q - R_q^{\mathsf T} are 4ab,4ac,4ad4ab, 4ac, 4ad (in positions (3,2),(1,3),(2,1)(3,2), (1,3), (2,1)), so 12(RqRqT)=Am\frac12(R_q - R_q^{\mathsf T}) = A_m with m=2a(b,c,d)=2cosθ2sinθ2u=sinθum = 2a\,(b, c, d) = 2\cos\frac\theta2\sin\frac\theta2\,u = \sin\theta\,u. Algorithm: θ=arccostrR12[0,π]\theta = \arccos\frac{\operatorname{tr}R - 1}{2} \in \intcc0\pi; if 0<θ<π0 < \theta < \pi, read uu off RRT2sinθ\frac{R - R^{\mathsf T}}{2\sin\theta} and set q=±(cosθ2+sinθ2u)q = \pm(\cos\frac\theta2 + \sin\frac\theta2 u); if θ=0\theta = 0, q=±1q = \pm1. For θ=π\theta = \pi: a=0a = 0, and Rodrigues (question 8) gives R=I+2Au2=I+2(uuTI)=2uuTIR = I + 2A_u^2 = I + 2(uu^{\mathsf T} - I) = 2uu^{\mathsf T} - I, i.e. R+I=2uuTR + I = 2uu^{\mathsf T}; any nonzero column of R+IR + I, normalized, is ±u\pm u, and q=±uq = \pm u.

14. With a=b=c=d=12a = b = c = d = \frac12: all diagonal entries vanish, 2(bcad)=02(bc - ad) = 0, 2(bd+ac)=12(bd + ac) = 1, 2(bc+ad)=12(bc + ad) = 1, 2(cdab)=02(cd - ab) = 0, 2(bdac)=02(bd - ac) = 0, 2(cd+ab)=12(cd + ab) = 1:

Rq=(001100010),R_q = \begin{pmatrix} 0 & 0 & 1\\ 1 & 0 & 0\\ 0 & 1 & 0 \end{pmatrix},

the cyclic permutation e1e2e3e1e_1 \to e_2 \to e_3 \to e_1. Its trace is 0=1+2cosθ0 = 1 + 2\cos\theta, so θ=2π3\theta = \frac{2\pi}3, and it fixes (1,1,1)(1,1,1): the rotation by 120120^\circ about the main diagonal — precisely question 11’s answer, now visible as the matrix that cycles the coordinate axes.

15. On the basis one checks Φ(qˉ)=Φ(q)\Phi(\bar q) = \Phi(q)^* (the matrix of qˉ\bar q has α=αˉ\alpha' = \bar\alpha, β=β\beta' = -\beta, which is the conjugate transpose of (αββˉαˉ)\bigl(\begin{smallmatrix}\alpha & \beta\\ -\bar\beta & \bar\alpha\end{smallmatrix}\bigr)). Hence Φ(q)Φ(q)=Φ(qˉq)=N(q)I\Phi(q)^*\Phi(q) = \Phi(\bar qq) = N(q)I and detΦ(q)=α2+β2=N(q)\det\Phi(q) = \abs\alpha^2 + \abs\beta^2 = N(q): for qS3q \in S^3, Φ(q)SU(2)\Phi(q) \in SU(2), and Φ\Phi is an injective morphism (question 1). Surjectivity: let U=(αβγδ)U = \bigl(\begin{smallmatrix}\alpha & \beta\\ \gamma & \delta\end{smallmatrix}\bigr) with detU=1\det U = 1; then U1=(δβγα)U^{-1} = \bigl(\begin{smallmatrix}\delta & -\beta\\ -\gamma & \alpha\end{smallmatrix}\bigr), and U1=U=(αˉγˉβˉδˉ)U^{-1} = U^* = \bigl(\begin{smallmatrix}\bar\alpha & \bar\gamma\\ \bar\beta & \bar\delta\end{smallmatrix}\bigr) forces δ=αˉ\delta = \bar\alpha, γ=βˉ\gamma = -\bar\beta, and then 1=detU=α2+β21 = \det U = \abs\alpha^2 + \abs\beta^2: U=Φ(q)U = \Phi(q) for the unit quaternion qq with coordinates α=a+ib\alpha = a + \iu b, β=c+id\beta = c + \iu d. So S3SU(2)S^3 \cong SU(2).

16. Real scalars are central and N(p)=1N(p) = 1 gives pqpˉ=pqˉpˉ\overline{pq\bar p} = p\bar q\bar p, so pqpˉ+pqpˉ=p(q+qˉ)pˉ=q+qˉpq\bar p + \overline{pq\bar p} = p(q + \bar q)\bar p = q + \bar q: the real part is invariant. Conversely let Req=Req=a\operatorname{Re}q = \operatorname{Re}q' = a; then the pure parts have the same norm 1a2=s\sqrt{1 - a^2} = s. If s=0s = 0, q=q=±1q = q' = \pm1. If s>0s > 0, write q=a+suq = a + su, q=a+suq' = a + su' with u,uu, u' unit pure; question 6 provides a rotation carrying uu' to uu, i.e. pS3p \in S^3 with ρp(u)=u\rho_p(u') = u, and then pqpˉ=a+sρp(u)=qpq'\bar p = a + s\rho_p(u') = q. The classes of S3S^3 are therefore {1}\{1\}, {1}\{-1\}, and for each a(1,1)a \in \intoo{-1}1 the 22-sphere {a+su:u unit pure}\{a + su : u \text{ unit pure}\} of radius ss. Under Φ\Phi, trΦ(q)=α+αˉ=2Req\operatorname{tr}\Phi(q) = \alpha + \bar\alpha = 2\operatorname{Re}q: the classes of SU(2)SU(2) are the level sets of the trace. Projecting: if q=pqpˉq' = pq\bar p then ρq=ρpρqρp1\rho_{q'} = \rho_p\rho_q\rho_p^{-1}; conversely ρq=ρpρqρp1=ρpqpˉ\rho_{q'} = \rho_p\rho_q\rho_p^{-1} = \rho_{pq\bar p} forces q=±pqpˉq' = \pm pq\bar p (kernel), so Req=±Req\operatorname{Re}q' = \pm\operatorname{Re}q, i.e. cosθ2=Req=Req=cosθ2\cos\frac{\theta'}2 = \abs{\operatorname{Re}q'} = \abs{\operatorname{Re}q} = \cos\frac\theta2 for the angles in [0,π]\intcc0\pi: conjugate rotations have equal angles. Conversely, equal angles allow representatives with the same nonnegative real part, conjugate by the above: in SO(3)SO(3), the conjugacy class of a rotation is exactly its angle.

17. NN is multiplicative, so \abs\cdot is a multiplicative norm on HR4\mathbb H \cong \R^4 and qk=qk\abs{q^k} = \abs q^k: the series qk/k!\sum q^k/k! converges absolutely in the finite-dimensional (hence complete) space, dominated by qk/k!=eq\sum\abs q^k/k! = \eu^{\abs q}. For a unit pure uu: u2=1u^2 = -1, so (θu)2m=(1)mθ2m(\theta u)^{2m} = (-1)^m\theta^{2m} and (θu)2m+1=(1)mθ2m+1u(\theta u)^{2m+1} = (-1)^m\theta^{2m+1}u; splitting the series,

exp(θu)=m(1)mθ2m(2m)!+um(1)mθ2m+1(2m+1)!=cosθ+sinθu.\exp(\theta u) = \sum_m\frac{(-1)^m\theta^{2m}}{(2m)!} + u\sum_m\frac{(-1)^m\theta^{2m+1}}{(2m+1)!} = \cos\theta + \sin\theta\,u .

Any qS3q \in S^3 is cosα+sinαu\cos\alpha + \sin\alpha\,u with α[0,π]\alpha \in \intcc0\pi (question 5): q=exp(αu)q = \exp(\alpha u), so exp(P)=S3\exp(P) = S^3. Finally exp(su)=coss+sinsu=q(2s)\exp(su) = \cos s + \sin s\,u = q(2s) in question 9’s notation, and ρq(t)=etAu\rho_{q(t)} = \eu^{tA_u} there: ρexp(su)=e2sAu\rho_{\exp(su)} = \eu^{2sA_u}.

18. Expanding vwvw coordinatewise with the table: the products ii=1\mathrm i\cdot\mathrm i = -1, … give the scalar (v1w1+v2w2+v3w3)-(v_1w_1 + v_2w_2 + v_3w_3), and the mixed products (ij=k\mathrm{ij} = \mathrm k, ji=k\mathrm{ji} = -\mathrm k, …) give the vector (v2w3v3w2, v3w1v1w3, v1w2v2w1)(v_2w_3 - v_3w_2,\ v_3w_1 - v_1w_3,\ v_1w_2 - v_2w_1): vw=v,w+vwvw = -\langle v, w\rangle + v\wedge w. Subtracting the reversed product: vwwv=2vwvw - wv = 2\,v\wedge w (the scalar parts cancel, the cross products add). For the matrices, with a(bc)=ba,cca,ba\wedge(b\wedge c) = b\langle a, c\rangle - c\langle a, b\rangle:

[Av,Aw]x=v(wx)w(vx)=wv,xvw,x=(vw)x=Avwx.[A_v, A_w]x = v\wedge(w\wedge x) - w\wedge(v\wedge x) = w\langle v, x\rangle - v\langle w, x\rangle = (v\wedge w)\wedge x = A_{v\wedge w}x .

Derivative: exp(tv)=exp(tv)\overline{\exp(tv)} = \exp(-tv) (conjugation is continuous and negates pure quaternions), so

 ⁣d ⁣dtt=0exp(tv)xexp(tv)=vxxv=2vx=2Avx:\frac{\dd}{\dd t}\Bigr|_{t=0}\exp(tv)\,x\,\exp(-tv) = vx - xv = 2\,v\wedge x = 2A_vx :

the differential is v2Avv \mapsto 2A_v, a linear bijection from PP onto the antisymmetric matrices, and [2Av,2Aw]=4Avw=2A2vw=2A[v,w][2A_v, 2A_w] = 4A_{v\wedge w} = 2A_{2v\wedge w} = 2A_{[v,w]} shows it carries the quaternion commutator to the matrix commutator.

19. Applying ρ\rho: ρ(ε(t))=ρ(s(R(t)))ρ(q(t))1=R(t)R(t)1=id\rho(\varepsilon(t)) = \rho(s(R(t)))\,\rho(q(t))^{-1} = R(t)R(t)^{-1} = \operatorname{id}, so ε(t)kerρ={±1}\varepsilon(t) \in \ker\rho = \{\pm1\} (question 4). As a product of the continuous maps ts(R(t))t \mapsto s(R(t)) and tq(t)1=q(t)ˉt \mapsto q(t)^{-1} = \bar{q(t)}, ε\varepsilon is continuous on the connected interval [0,2π]\intcc0{2\pi} with values in the discrete pair {±1}\{\pm1\}: it is constant, say ε(t)ε\varepsilon(t) \equiv \varepsilon. But R(0)=R(2π)=IR(0) = R(2\pi) = I, so s(R(0))=s(R(2π))s(R(0)) = s(R(2\pi)), while s(R(0))=εq(0)=εs(R(0)) = \varepsilon\,q(0) = \varepsilon and s(R(2π))=εq(2π)=εs(R(2\pi)) = \varepsilon\,q(2\pi) = -\varepsilon: contradiction. No continuous section exists: the sign ambiguity ±q\pm q is global, not a defect of a particular formula.

20. Onto: every RSO(3)R \in SO(3) is eθAu\eu^{\theta A_u} for some unit uu and θ[0,2π]\theta \in \intcc0{2\pi} (questions 6 and 8); if θ>π\theta > \pi, Rodrigues gives eθAu=e(2πθ)Au\eu^{\theta A_u} = \eu^{(2\pi - \theta)A_{-u}} (both equal I+sinθAu+(1cosθ)Au2I + \sin\theta A_u + (1 - \cos\theta)A_u^2, and Au=AuA_{-u} = -A_u with sin(2πθ)=sinθ\sin(2\pi - \theta) = -\sin\theta, cos(2πθ)=cosθ\cos(2\pi - \theta) = \cos\theta), so R=E(v)R = E(v) with vπ\norm v \leq \pi. Injective inside: if E(v)=E(v)IE(v) = E(v') \neq I with v,v<π\norm v, \norm{v'} < \pi, question 13 recovers the same angle θ=v=v(0,π)\theta = \norm v = \norm{v'} \in \intoo0\pi from the trace and, since sinθ0\sin\theta \neq 0, the same axis from the antisymmetric part: v=vv = v'; and E(v)=IE(v) = I forces θ{0}\theta \in \{0\} on the open ball. Boundary: E(πu)=I+2Au2=2uuTIE(\pi u) = I + 2A_u^2 = 2uu^{\mathsf T} - I depends on uu only through uuTuu^{\mathsf T}, whence E(πu)=E(πu)E(\pi u) = E(-\pi u); conversely 2uuTI=2uuTI2uu^{\mathsf T} - I = 2u'u'^{\mathsf T} - I applied to uu gives u=u,uuu = \langle u', u\rangle u', so u=±uu' = \pm u. Interior and boundary never collide (trace >1> -1 versus =1= -1). So EE induces a continuous bijection from the ball-with-antipodal-gluing — compact — onto SO(3)SO(3): a homeomorphism, and SO(3)RP3SO(3) \cong \mathbb{RP}^3. A diameter from πu\pi u to πu-\pi u has glued endpoints: it is a loop in SO(3)SO(3), and its EE-description matches question 10’s family of rotations about uu sweeping a full turn.

21. ρ\rho is a morphism and ρp1=ρp1=ρpˉ\rho_p^{-1} = \rho_{p^{-1}} = \rho_{\bar p}, so ρpρqρp1=ρpqpˉ\rho_p\rho_q\rho_p^{-1} = \rho_{pq\bar p}; by question 16, conjugating a rotation preserves its angle and rotates its axis by ρp\rho_p. Involutions: ρq2=ρq2=id\rho_q^2 = \rho_{q^2} = \operatorname{id} iff q2=±1q^2 = \pm1. If q2=1q^2 = 1 then (q1)(q+1)=q21=0(q - 1)(q + 1) = q^2 - 1 = 0 (central scalars, so this factorization is valid) and q=±1q = \pm1 in the division ring H\mathbb H, giving ρq=I\rho_q = I, excluded; q2=1q^2 = -1 with q=a+suq = a + su gives a2s2+2asu=1a^2 - s^2 + 2as\,u = -1, so a=0a = 0, s=1s = 1: qq is a unit pure ww, and ρw\rho_w is the half-turn about ww (angle π\pi, question 5). Center: if ρq\rho_q commutes with every ρp\rho_p, then ρpqpˉ=ρq\rho_{pq\bar p} = \rho_q, so pqpˉ=ε(p)qpq\bar p = \varepsilon(p)\,q with ε(p){±1}\varepsilon(p) \in \{\pm1\}; pε(p)=(pqpˉ)q1p \mapsto \varepsilon(p) = (pq\bar p)q^{-1} is continuous on the connected S3S^3 and equals 11 at p=1p = 1, hence 1\equiv 1: qq commutes with all of S3S^3, hence with all of H\mathbb H (rescale), so qRS3={±1}q \in \R \cap S^3 = \{\pm1\} (question 1) and ρq=I\rho_q = I: the center is trivial.

22. Since uwu \perp w are unit pures, uw=uwuw = u\wedge w is pure (question 18), so

w=qw=cosθ2w+sinθ2uww' = qw = \cos\tfrac\theta2\,w + \sin\tfrac\theta2\,u\wedge w

is pure, of norm qw=1\abs q\abs w = 1. Then ww=qww=qw2=qw'w = qw\cdot w = qw^2 = -q, and

ρwρw=ρww=ρq=ρq.\rho_{w'}\rho_w = \rho_{w'w} = \rho_{-q} = \rho_q .

Both axes ww and w=cosθ2w+sinθ2(uw)w' = \cos\frac\theta2 w + \sin\frac\theta2(u\wedge w) lie in the plane uu^\perp orthogonal to the rotation axis, and w,w=cosθ2\langle w', w\rangle = \cos\frac\theta2: they make the half-angle θ2\frac\theta2. This is the classical generation: two half-turns about axes meeting at angle θ2\frac\theta2 compose to the rotation of angle θ\theta about their common perpendicular.

23. Compactness is question 7. Path-connectedness: SO(3)=ρ(S3)SO(3) = \rho(S^3) is the continuous image of the path-connected sphere; alternatively, for R=eAR = \eu^{A} with AA antisymmetric (question 8), tetAt \mapsto \eu^{tA} is a path in SO(3)SO(3) from II to RR (orthogonal since (etA)T=etA(\eu^{tA})^{\mathsf T} = \eu^{-tA}, determinant 11 by continuity from t=0t = 0). For O(3)O(3): det\det is continuous onto {±1}\{\pm1\}, so O(3)O(3) is disconnected, O(3)=SO(3)DSO(3)O(3) = SO(3) \sqcup D\,SO(3) for any fixed DD with detD=1\det D = -1 (e.g. D=ID = -I), and left multiplication by DD is a homeomorphism: exactly two components, each a copy of SO(3)SO(3). The two-to-one ρ\rho, section-free by question 19, is thus an honest double cover of a connected compact group by the simply connected S3S^3 — the geometry behind the belt trick.

24. (i) Matrices:

R1=(010100001),R2=(100001010),R2R1=(010001100).R_1 = \begin{pmatrix} 0 & -1 & 0\\ 1 & 0 & 0\\ 0 & 0 & 1 \end{pmatrix}, \quad R_2 = \begin{pmatrix} 1 & 0 & 0\\ 0 & 0 & -1\\ 0 & 1 & 0 \end{pmatrix}, \quad R_2R_1 = \begin{pmatrix} 0 & -1 & 0\\ 0 & 0 & -1\\ 1 & 0 & 0\end{pmatrix}.

Trace 0=1+2cosθ0 = 1 + 2\cos\theta gives cosθ=12\cos\theta = -\frac12: θ=2π3\theta = \frac{2\pi}3. Antisymmetric part RRT2\frac{R - R^{\mathsf T}}2 has entries encoding sinθ(v3,v2,v1)\sin\theta\,(v_3, -v_2, v_1)-wise the axis: here RRT2=12(011101110)\frac{R - R^{\mathsf T}}2 = \frac12\bigl(\begin{smallmatrix}0 & -1 & -1\\ 1 & 0 & -1\\ 1 & 1 & 0\end{smallmatrix}\bigr), which reads (Part V’s dictionary AvA_v) vsinθ=12(1,1,1)v\sin\theta = \frac12(1, -1, 1); with sin2π3=32\sin\frac{2\pi}3 = \frac{\sqrt3}2: v=13(1,1,1)v = \frac1{\sqrt3}(1, -1, 1). (ii) Quaternions: q1=cosπ4+sinπ4k=22(1+k)q_1 = \cos\frac\pi4 + \sin\frac\pi4\,k = \frac{\sqrt2}2(1 + k), q2=22(1+i)q_2 = \frac{\sqrt2}2(1 + i), and

q2q1=12(1+i)(1+k)=12(1+k+i+ik)=12(1+ij+k)q_2q_1 = \tfrac12(1 + i)(1 + k) = \tfrac12(1 + k + i + ik) = \tfrac12\bigl(1 + i - j + k\bigr)

(ik=jik = -j). So cosθ2=12\cos\frac\theta2 = \frac12: θ=2π3\theta = \frac{2\pi}3, and the vector part 12(ij+k)\frac12(i - j + k) has direction 13(1,1,1)\frac1{\sqrt3}(1, -1, 1) — the same answer, with the quaternion route requiring one line of multiplication instead of a matrix product: the practical reason flight software composes attitudes in S3S^3.

25. I+KI + K invertible: (I+K)v=0(I + K)v = 0 gives 0=v,v+v,Kv=v20 = \langle v, v\rangle + \langle v, Kv\rangle = \norm v^2 (antisymmetry kills the second term): v=0v = 0. Orthogonality of C=C(K)C = C(K): using (I±K)T=IK(I \pm K)^{\mathsf T} = I \mp K and the fact that all four matrices I±KI \pm K, (I±K)1(I \pm K)^{-1} commute (polynomial expressions in KK, plus limits):

CTC=(I+K)T(IK)T(IK)(I+K)1=(IK)1(I+K)(IK)(I+K)1=I.C^{\mathsf T}C = (I + K)^{-\mathsf T}(I - K)^{\mathsf T} (I - K)(I + K)^{-1} = (I - K)^{-1}(I + K)(I - K)(I + K)^{-1} = I .

Determinant: det(IK)=det((IK)T)=det(I+K)\det(I - K) = \det\bigl((I - K)^{\mathsf T}\bigr) = \det(I + K), so detC=1\det C = 1: CSO(n)C \in SO(n). No eigenvalue 1-1: Cv=vCv = -v means (IK)w=(I+K)w(I - K)w = -(I + K)w for w=(I+K)1vw = (I + K)^{-1}v, i.e. 2w=02w = 0: v=0v = 0. Inversion: from C(I+K)=IKC(I + K) = I - K, solve K(I+C)=ICK(I + C) = I - C; since 1SpC-1 \notin \operatorname{Sp}C, I+CI + C is invertible and K=(IC)(I+C)1K = (I - C)(I + C)^{-1}, which is antisymmetric whenever CC is orthogonal without eigenvalue 1-1 (transpose the expression and use CT=C1C^{\mathsf T} = C^{-1}: KT=(IC1)(I+C1)1=(CI)(C+I)1=KK^{\mathsf T} = (I - C^{-1})(I + C^{-1})^{-1} = (C - I)(C + I)^{-1} = -K). The two maps are mutually inverse by construction: a global rational parametrization of the dense open piece of SO(n)SO(n) avoiding eigenvalue 1-1 — no series, no trigonometry, and in dimension 33 it is the half-angle substitution K=tanθ2AvK = \tan\frac\theta2\,A_v in disguise.