Mathematics · Book 5 · Bachelor Year 3

University Mathematics — Year 3

University Mathematics — Year 3 · Bachelor Year 3

3Modules over a Principal Ideal Domain

Linear algebra over a ring instead of a field: this small change of hypothesis produces one of algebra’s great unification theorems. A module over Z\Z is an abelian group; a module over K[X]K[X] is a vector space equipped with an endomorphism. The structure theorem for finitely generated modules over a PID therefore classifies, in one stroke, all finitely generated abelian groups and all endomorphisms up to similarity — Jordan’s reduction, which Year 2 obtained by delicate inductions, falls out as a corollary, together with its subtler sibling, the rational canonical form, valid over every field. The computational engine is the Smith normal form, an arithmetic of matrices worthy of Euclid.

Throughout, AA is a commutative ring, soon a PID; “module” means AA-module.

3.1 Modules, free modules

Definition 3.1

An AA-module is an abelian group (M,+)(M, +) with a scalar multiplication A×MMA \times M \to M satisfying the vector-space axioms: a(x+y)=ax+aya(x + y) = ax + ay, (a+b)x=ax+bx(a + b)x = ax + bx, (ab)x=a(bx)(ab)x = a(bx), 1x=x1x = x. Submodules, quotients M/NM/N, morphisms (AA-linear maps), direct sums iMi\bigoplus_i M_i, and the isomorphism theorems are defined and proved word for word as for vector spaces and abelian groups; in particular M/kerfimfM/\ker f \cong \operatorname{im} f for a morphism ff.

Example 3.2

The three motivating cases.

  1. A=KA = K a field: modules are vector spaces.
  2. A=ZA = \Z: modules are exactly abelian groups (nxnx is forced to be x++xx + \dots + x), submodules are subgroups.
  3. A=K[X]A = K[X]: a module is a KK-vector space VV together with the KK-linear map u ⁣:xXxu\colon x \mapsto X\cdot x — conversely, every pair (V,u)(V, u) with uL(V)u \in \mathcal L(V) becomes a K[X]K[X]-module by Px=P(u)(x)P \cdot x = P(u)(x). The submodules are precisely the uu-stable subspaces.

An ideal of AA is exactly a submodule of AA; a quotient ring A/IA/I is an AA-module. Unlike vector spaces, modules can have torsion: in Z/6Z\Z/6\Z, the element 3ˉ0\bar 3 \ne 0 is killed by 202 \neq 0.

Definition 3.3

MM is finitely generated if M=Ax1++AxnM = Ax_1 + \dots + Ax_n for some xix_i. MM is free of rank nn if MAnM \cong A^n, i.e. if it has a basis (a generating family that is AA-linearly independent). Every finitely generated MM is a quotient of a free module: (a1,,an)aixi(a_1, \dots, a_n) \mapsto \sum a_ix_i maps AnA^n onto MM.

Proposition 3.4 (Invariance of rank)

If A0A \neq 0 and AmAnA^m \cong A^n, then m=nm = n.

Proof. Pick a maximal ideal m\mathfrak m of AA (Theorem 2.8) and set k=A/mk = A/\mathfrak m, a field. An isomorphism f ⁣:AmAnf \colon A^m \to A^n maps mAm\mathfrak m A^m into mAn\mathfrak m A^n (linearity), hence induces an isomorphism of quotients

Am/mAm    An/mAn,i.e.kmknA^m/\mathfrak m A^m \;\cong\; A^n/\mathfrak m A^n, \qquad\text{i.e.}\qquad k^m \cong k^n

as kk-vector spaces (the quotient Am/mAmA^m/\mathfrak m A^m is killed by m\mathfrak m, so the AA-action factors through kk; the images of the standard basis form a kk-basis). Dimension theory over the field kk gives m=nm = n.

Theorem 3.5 (Submodules of free modules)

Let AA be a PID and MAnM \subseteq A^n a submodule. Then MM is free of rank n\leq n.

Proof. Induction on nn. For n=1n = 1: MM is an ideal, so M=(0)M = (0) (free of rank 00) or M=dAAM = dA \cong A (xdxx \mapsto dx is injective: domain). For n>1n > 1: let π ⁣:AnA\pi \colon A^n \to A be the last coordinate. Then π(M)\pi(M) is an ideal, (0)(0) or dAdA. If (0)(0): MAn1×{0}M \subseteq A^{n-1} \times \{0\} and induction applies. Otherwise choose x0Mx_0 \in M with π(x0)=d\pi(x_0) = d. Every xMx \in M writes uniquely

x=π(x)dAx0+(xπ(x)dx0),xπ(x)dx0Mkerπx = \underbrace{\frac{\pi(x)}{d}}_{\in A}\, x_0 + \Bigl(x - \tfrac{\pi(x)}d x_0\Bigr), \qquad x - \tfrac{\pi(x)}d x_0 \in M \cap \ker\pi

(π(x)dA\pi(x) \in dA, so the coefficient is in AA). Thus M=Ax0(Mkerπ)M = Ax_0 \oplus (M \cap \ker \pi): the sum is direct since π(ax0)=ad=0\pi(ax_0) = ad = 0 forces a=0a = 0. By induction MkerπkerπAn1M \cap \ker\pi \subseteq \ker \pi \cong A^{n-1} is free of rank n1\leq n - 1; adjoining x0x_0 (independent from kerπ\ker\pi as just seen) gives a basis of MM of cardinality n\leq n.

Remark 3.6

Consequently, over a PID every finitely generated module MM has a finite presentation: a surjection φ ⁣:AnM\varphi\colon A^n \to M has free kernel with basis c1,,ckc_1, \dots, c_k (knk \leq n), and MAn/im(C)M \cong A^n / \operatorname{im}(C) where CMn,k(A)C \in M_{n,k}(A) is the matrix whose columns are the cjc_j. Understanding MM means understanding a matrix over AA up to change of bases in source and target — the subject of the next section.

3.2 Smith normal form

Definition 3.7

Two matrices B,CMn,k(A)B, C \in M_{n,k}(A) are equivalent if C=QBPC = QBP with QGLn(A)Q \in GL_n(A), PGLk(A)P \in GL_k(A) (invertible over AA: determinant in A×A^\times). Equivalent presentation matrices define isomorphic modules An/imA^n/ \operatorname{im} (change bases in AnA^n and AkA^k).

Theorem 3.8 (Smith normal form)

Let AA be a PID and BMn,k(A)B \in M_{n,k}(A). Then BB is equivalent to a diagonal matrix

diag(d1,d2,,dr,0,,0),d1d2dr0,\operatorname{diag}(d_1, d_2, \dots, d_r, 0, \dots, 0), \qquad d_1 \mid d_2 \mid \cdots \mid d_r \neq 0 ,

and the did_i are unique up to associates: d1did_1 \cdots d_i is a gcd of the i×ii \times i minors of BB (in particular that gcd is an invariant of equivalence). The did_i are the invariant factors of BB.

Proof. Existence. If B=0B = 0, done. Otherwise, consider the set of ideals (b)(b) generated by entries of matrices equivalent to BB; since AA is Noetherian, choose a matrix BB' equivalent to BB and an entry dd of BB' with (d)(d) maximal in this set. Move dd to position (1,1)(1,1) by row and column swaps.

Claim: dd divides every entry of BB'. First, column 1: if bi1b_{i1} is not a multiple of dd, let e=gcd(d,bi1)=ud+vbi1e = \gcd(d, b_{i1}) = ud + vb_{i1} (Bézout), so (e)(d)(e) \supsetneq (d). The 2×22 \times 2 matrix trick: acting on rows 11 and ii by

(uvbi1ede)GL2(A)(det=ud+vbi1e=1)\begin{pmatrix} u & v \\ -\dfrac{b_{i1}}e & \dfrac de \end{pmatrix} \in GL_2(A) \qquad \Bigl(\det = \tfrac{ud + v b_{i1}}e = 1\Bigr)

produces an equivalent matrix with entry ee in position (1,1)(1,1): this contradicts maximality of (d)(d). So dd divides column 11, and symmetrically row 11. Subtracting multiples of row 1 and column 1 clears them: BB' is equivalent to (d00B)\begin{pmatrix} d & 0\\ 0 & B''\end{pmatrix}. Next, dd divides every entry bb of BB'': add the row of bb to row 1 (an elementary operation; the new first row contains dd and bb-entries), and repeat the column-clearing argument: a non-multiple would again improve (d)(d). Now induct on the size: BB'', all of whose entries are divisible by dd, has a Smith form diag(d2,)\operatorname{diag}(d_2, \dots) whose entries remain divisible by dd (every entry of any QBPQB''P is an AA-combination of entries of BB''); set d1=dd_1 = d.

Uniqueness. Let Di(B)D_i(B) denote a gcd of all i×ii \times i minors. Row and column operations, and more generally multiplication by any matrix, cannot shrink the gcd: the i×ii \times i minors of QBQB are AA-combinations of those of BB (Cauchy–Binet expansion; or directly: each row of QBQB is a combination of rows of BB, and minors are multilinear in rows). So Di(QBP)D_i(QBP) and Di(B)D_i(B) divide each other: DiD_i is an equivalence invariant. On the diagonal form, the nonzero i×ii \times i minors are the products of ii of the djd_j’s, and divisibility d1drd_1 \mid \dots \mid d_r makes d1did_1 \cdots d_i the gcd. Hence d1di=Di(B)d_1 \cdots d_i = D_i(B) up to units, and di=Di/Di1d_i = D_i/D_{i-1} is determined.

Method 3.9

Over a Euclidean domain (Z\Z, K[X]K[X]), Smith reduction is an algorithm — no maximality argument needed: bring the entry of smallest Euclidean size to position (1,1)(1,1); if it fails to divide some entry of its row or column, a Euclidean division leaves a strictly smaller remainder there — swap it in and restart (termination: sizes decrease); when it divides its whole row and column, clear them; if it fails to divide an inner entry, add that row to row 11 and restart; recurse on the inner block. In practice on integer matrices: compute D1=gcdD_1 = \gcd of entries, D2D_2, … via minors for small sizes, or run the algorithm.

Example 3.10 (A Smith reduction, in full)

Reduce M=(123456789)M = \begin{pmatrix} 1 & 2 & 3\\ 4 & 5 & 6\\ 7 & 8 & 9 \end{pmatrix} over Z\Z. The corner 11 divides everything: clear its row and column (L2L24L1L_2 \leftarrow L_2 - 4L_1, L3L37L1L_3 \leftarrow L_3 - 7L_1, then C2C22C1C_2 \leftarrow C_2 - 2C_1, C3C33C1C_3 \leftarrow C_3 - 3C_1):

M(1000360612).M \sim \begin{pmatrix} 1 & 0 & 0\\ 0 & -3 & -6\\ 0 & -6 & -12 \end{pmatrix} .

In the inner block, the corner 3-3 divides all entries: L3L32L2L_3 \leftarrow L_3 - 2L_2 and C3C32C2C_3 \leftarrow C_3 - 2C_2 clear it to diag(3,0)\operatorname{diag}(-3, 0). Adjusting signs (multiply a row by 1-1, a legal operation):

Mdiag(1,3,0),Z3/MZ3Z/3Z×Z.M \sim \operatorname{diag}(1, 3, 0), \qquad \Z^3/M\Z^3 \cong \Z/3\Z \times \Z .

Cross-check by determinantal divisors: D1=gcd(entries)=1D_1 = \gcd(\text{entries}) = 1; every 2×22\times2 minor of MM is a multiple of 33 (e.g. det(1245)=3\det\bigl(\begin{smallmatrix}1 & 2\\ 4 & 5\end{smallmatrix}\bigr) = -3) and one equals 3-3: D2=3D_2 = 3; D3=detM=0D_3 = \det M = 0. Hence d1=1d_1 = 1, d2=3d_2 = 3, d3=0d_3 = 0: same answer. Two lessons: a zero invariant factor records the rank drop (the cokernel picks up a free Z\Z summand), and the divisibility chain 1301 \mid 3 \mid 0 is the Smith certificate — a diagonal reduction that violates the chain (say diag(2,3)\operatorname{diag}(2, 3), which the careless can produce from (2003)\bigl(\begin{smallmatrix}2 & 0\\ 0 & 3\end{smallmatrix}\bigr) by stopping too early: correct Smith form diag(1,6)\operatorname{diag}(1, 6), as D1=1D_1 = 1 here!) is not finished.

The sublattice L = ℤ(2,0) + ℤ(1,3) of ℤ2 (red points). The Smith normal form of ( smallmatrix 2 & 1\\ 0 & 3 smallmatrix ) is diag(1, 6): in the adapted basis f_1 = (1,3), f_2 = (0,1) of ℤ2, one has L = ℤ f_1 ℤ\,6f_2, so ℤ2/L ℤ/6ℤ — the index equals | | = 6, the area of the shaded fundamental domain.
The sublattice L=Z(2,0)+Z(1,3)L = \Z(2,0) + \Z(1,3) of Z2\Z^2 (red points). The Smith normal form of (2103)\bigl(\begin{smallmatrix} 2 & 1\\ 0 & 3\end{smallmatrix}\bigr) is diag(1,6)\operatorname{diag}(1, 6): in the adapted basis f1=(1,3)f_1 = (1,3), f2=(0,1)f_2 = (0,1) of Z2\Z^2, one has L=Zf1Z6f2L = \Z f_1 \oplus \Z\,6f_2, so Z2/LZ/6Z\Z^2/L \cong \Z/6\Z — the index equals det=6\abs{\det} = 6, the area of the shaded fundamental domain.

3.3 The structure theorem

Definition 3.11

Let AA be a domain and MM an AA-module. The torsion submodule is

T(M)={xM:ax=0 for some a0}T(M) = \{x \in M : ax = 0 \text{ for some } a \neq 0\}

(a submodule: if ax=by=0ax = by = 0 then ab(x+y)=0ab(x + y) = 0, ab0ab \ne 0). MM is torsion-free if T(M)=0T(M) = 0, a torsion module if T(M)=MT(M) = M.

Theorem 3.12 (Structure of finitely generated modules over a PID)

Let AA be a PID and MM a finitely generated AA-module. There exist a unique rNr \in \N and nonzero nonunits d1d2dsd_1 \mid d_2 \mid \cdots \mid d_s, unique up to associates, with

M    ArA/(d1)A/(ds).M \;\cong\; A^r \,\oplus\, A/(d_1) \oplus \cdots \oplus A/(d_s).

Moreover T(M)A/(d1)A/(ds)T(M) \cong A/(d_1)\oplus\dots\oplus A/(d_s) and M/T(M)ArM/T(M) \cong A^r: a finitely generated torsion-free module over a PID is free.

Proof. Existence. Present MAn/im(C)M \cong A^n/\operatorname{im}(C) (Remark 3.6) and put CC in Smith form: after the two base changes, MAn/(d1A××drA×0××0)A/(d1)A/(dr)AnrM \cong A^n / (d_1A \times \dots \times d_rA \times 0 \times \dots \times 0) \cong A/(d_1) \oplus \dots \oplus A/(d_r) \oplus A^{\,n - r}. Discard the factors where did_i is a unit (A/(di)=0A/(d_i) = 0); the divisibility chain survives.

The torsion identification. In the decomposition, ArA^r is torsion-free (a domain has no zero divisors) and each A/(di)A/(d_i) is torsion (killed by di0d_i \ne 0); a direct sum splits torsion accordingly: T(M)=iA/(di)T(M) = \bigoplus_i A/(d_i) and M/T(M)ArM/T(M) \cong A^r.

Uniqueness of rr: M/T(M)ArM/T(M) \cong A^r depends only on MM, and Proposition 3.4 pins rr down.

Uniqueness of the did_i: it suffices to treat the torsion module T=T(M)T = T(M). Decompose each did_i into primes and split by the Chinese remainder theorem (Theorem 2.9; distinct primes generate comaximal ideals):

A/(d)pdA/(pvp(d)):TpjA/(pkp,j),A/(d) \cong \bigoplus_{p \mid d} A/\bigl(p^{v_p(d)}\bigr): \qquad T \cong \bigoplus_{p} \bigoplus_{j} A/\bigl(p^{k_{p,j}}\bigr),

the elementary divisors pkp,jp^{k_{p,j}}. Conversely the did_i are reconstructed from the multiset of elementary divisors (dsd_s = product of the highest power of each prime, etc.), so it suffices to prove the multiset {kp,j}j\{k_{p,j}\}_j is determined by TT, for each prime pp. Fix pp; for j1j \geq 1 consider the A/(p)A/(p)-vector spaces pj1T/pjTp^{j-1}T/p^jT. On a cyclic factor A/(pk)A/(p^k):

pj1(A/(pk))/pj(A/(pk)){A/(p)if jk,0if j>k,p^{j-1}\bigl(A/(p^k)\bigr)\big/p^{j}\bigl(A/(p^k)\bigr) \cong \begin{cases} A/(p) & \text{if } j \leq k,\\ 0 & \text{if } j > k, \end{cases}

and on a factor A/(qk)A/(q^k), qpq \neq p: multiplication by pp is bijective there (pp invertible mod qkq^k: Bézout), so the quotient is 00. Direct sums pass through: dimA/(p)pj1T/pjT=#{i:kp,ij}\dim_{A/(p)} p^{j-1}T/p^jT = \#\{i : k_{p,i} \geq j\}. These intrinsic dimensions determine the multiset of exponents.

Corollary 3.13 (Finitely generated abelian groups)

Every finitely generated abelian group is Zr×Z/d1Z××Z/dsZ\Z^r \times \Z/d_1\Z\times\dots\times\Z/d_s\Z with d1dsd_1 \mid \dots \mid d_s, uniquely. Every finite abelian group is a product of cyclic groups of prime-power order, unique as a multiset.

Example 3.14

The abelian groups of order pnp^n correspond to partitions of nn: for p4p^4: Z/p4\Z/p^4, Z/p3×Z/p\Z/p^3\times\Z/p, (Z/p2)2(\Z/p^2)^2, Z/p2×(Z/p)2\Z/p^2 \times (\Z/p)^2, (Z/p)4(\Z/p)^4 — five groups, as 44 has five partitions. Mixed orders multiply the counts prime by prime (CRT): there are 5×25 \times 2 abelian groups of order 2432=1442^4 \cdot 3^2 = 144.

3.4 Application: canonical forms of endomorphisms

Let KK be a field, VV a KK-vector space of finite dimension nn, and uL(V)u \in \mathcal L(V); make VV a K[X]K[X]-module via Px=P(u)(x)P \cdot x = P(u)(x) (Example 3.2). This module is finitely generated (a KK-basis generates) and torsion: for each xx, the n+1n+1 vectors x,u(x),,un(x)x, u(x), \dots, u^n(x) are KK-dependent, providing a nonzero annihilating polynomial.

Definition 3.15

For P=Xm+am1Xm1++a0P = X^m + a_{m-1}X^{m-1} + \dots + a_0 monic, the companion matrix is

CP=(0a01a101am1):C_P = \begin{pmatrix} 0 & & & -a_0\\ 1 & \ddots & & -a_1\\ & \ddots & 0 & \vdots\\ & & 1 & -a_{m-1} \end{pmatrix} :

the matrix of “multiplication by XX” on K[X]/(P)K[X]/(P) in the basis 1,Xˉ,,Xˉm11, \bar X, \dots, \bar X^{m-1}.

Theorem 3.16 (Frobenius: rational canonical form)

There is a unique sequence of monic nonconstant polynomials P1P2PsP_1 \mid P_2 \mid \dots \mid P_s (the similarity invariants of uu) such that, as K[X]K[X]-modules,

VK[X]/(P1)K[X]/(Ps):V \cong K[X]/(P_1) \oplus \cdots \oplus K[X]/(P_s):

in a suitable basis, uu has block-diagonal matrix diag(CP1,,CPs)\operatorname{diag}(C_{P_1}, \dots, C_{P_s}). Moreover:

  1. Ps=μuP_s = \mu_u (minimal polynomial) and P1Ps=χuP_1\cdots P_s = \chi_u (characteristic polynomial); in particular μuχu\mu_u \mid \chi_u (Cayley–Hamilton re-proved) and χuμus\chi_u \mid \mu_u^{\,s}, so χu\chi_u and μu\mu_u have the same irreducible factors.
  2. Two endomorphisms (or square matrices) are similar iff they have the same similarity invariants.

Proof. The structure theorem (Theorem 3.12) applied to the PID K[X]K[X]: the torsion module VV decomposes with invariant factors PiP_i, normalized monic (units of K[X]K[X] are K×K^\times); no free part occurs (VV is torsion). On each cyclic factor K[X]/(Pi)K[X]/(P_i), multiplication by XX has matrix CPiC_{P_i} in the basis of powers of Xˉ\bar X: concatenating bases gives the block form.

(1) The annihilator of V=K[X]/(Pi)V = \bigoplus K[X]/(P_i) is (P1)(Ps)=(Ps)(P_1)\cap \dots\cap(P_s) = (P_s) (divisibility chain: PsP_s is a common multiple, and the class of 11 in the last factor is killed exactly by (Ps)(P_s)): μu=Ps\mu_u = P_s. For χu\chi_u: on a cyclic factor, χCP=P\chi_{C_P} = P, by induction on m=degPm = \deg P. Expanding det(XImCP)\det(XI_m - C_P) along the first row (whose entries are XX, then zeros, then a0a_0 in the last column):

det(XImCP)=Xdet(XIm1CP~)+(1)1+ma0detL,\det(XI_m - C_P) = X\,\det\bigl(XI_{m-1} - C_{\tilde P}\bigr) + (-1)^{1+m}\,a_0\,\det L ,

where P~=Xm1+am1Xm2++a1\tilde P = X^{m-1} + a_{m-1}X^{m-2} + \dots + a_1 (same shape, one size down) and LL is triangular with diagonal (1,,1)(-1, \dots, -1), so detL=(1)m1\det L = (-1)^{m-1}. By induction the first term is XP~X\tilde P, and the second is a0a_0: the total is XP~+a0=PX\tilde P + a_0 = P (base case m=1m=1: det(X+a0)=P\det(X + a_0) = P). Determinants multiply over blocks: χu=Pi\chi_u = \prod P_i. Cayley–Hamilton: χu(μu)\chi_u \in (\mu_u) since PsP_s \mid each… conversely each PiPsP_i \mid P_s, so χu=Pi\chi_u = \prod P_i divides Pss=μusP_s^{\,s} = \mu_u^s; and μu=Ps\mu_u = P_s divides χu\chi_u as one of its factors.

(2) Similar endomorphisms are conjugate module structures, hence have equal invariants (uniqueness in Theorem 3.12); conversely equal invariants give isomorphic K[X]K[X]-modules, and a module isomorphism is exactly a linear bijection intertwining the two endomorphisms: a similarity.

Corollary 3.17 (Similarity is insensitive to field extension)

Let KLK \subseteq L be fields and M,NMn(K)M, N \in M_n(K). If MM and NN are similar over LL, they are similar over KK.

Proof. The similarity invariants of MM are computed by Smith’s minor formula (Theorem 3.8) applied to the presentation matrix XInMXI_n - M over K[X]K[X] — indeed the K[X]K[X]-module VM=KnV_M = K^n has presentation XInMXI_n - M: the map K[X]nVMK[X]^n \to V_M, (Qi)Qi(M)ei(Q_i) \mapsto \sum Q_i(M)e_i, is onto with kernel generated by the columns of XInMXI_n - M (a direct verification: modulo those columns, every element of K[X]nK[X]^n reduces to a constant vector, and constant vectors map bijectively; the weekend problem spells this out). Gcds of polynomials do not change under field extension: if dd is the monic gcd in K[X]K[X] of a family (fj)(f_j), Bézout gives d=ujfjd = \sum u_jf_j with ujK[X]u_j \in K[X], so every common divisor of the fjf_j in L[X]L[X] divides dd; as dd is itself a common divisor, it is the gcd in L[X]L[X] too. Hence the invariant factors of XInMXI_n - M, quotients of successive minor gcds, are the same over KK and over LL: M,NM, N have the same similarity invariants over LL iff over KK; conclude by Theorem 3.16(2).

Theorem 3.18 (Jordan form, re-derived)

Suppose χu\chi_u splits over KK (e.g. K=CK = \C). Applying to VV the elementary-divisor decomposition (proof of Theorem 3.12) instead of invariant factors:

Vλ,jK[X]/((Xλ)kλ,j),V \cong \bigoplus_{\lambda, j} K[X]\big/\bigl((X - \lambda)^{k_{\lambda,j}}\bigr),

and in the basis ((Xλ)k1,,(Xλ),1ˉ)\bigl(\overline{(X-\lambda)^{k-1}}, \dots, \overline{(X - \lambda)}, \bar 1\bigr) of each factor, uu acts as the Jordan block Jk(λ)J_k(\lambda): every endomorphism with split characteristic polynomial has a Jordan basis, and the multiset of blocks (λ,k)(\lambda, k) is unique.

Proof. The elementary divisors of the torsion module VV are the (Xλ)k(X - \lambda)^k with XλX - \lambda ranging over the irreducible factors of μu\mu_u (which splits, since χu\chi_u does and both have the same irreducible factors, Theorem 3.16). In W=K[X]/((Xλ)k)W = K[X]/((X-\lambda)^k), put fj=(Xλ)kjf_j = \overline{(X - \lambda)^{k-j}} for j=1,,kj = 1, \dots, k: then (Xλ)fj=fj1(X - \lambda)f_j = f_{j-1} (with f0=0f_0 = 0), i.e. u(fj)=λfj+fj1u(f_j) = \lambda f_j + f_{j-1}: the matrix of uu on (f1,,fk)(f_1, \dots, f_k) is exactly Jk(λ)J_k(\lambda) (ones above the diagonal). Uniqueness of the multiset of elementary divisors is Theorem 3.12.

Remark 3.19

The hierarchy of canonical forms is now transparent: the rational form exists over every field and detects similarity absolutely (Corollary 3.17); the Jordan form is its refinement when χu\chi_u splits. Year 2’s dimension-counting proofs of Jordan’s theorem are subsumed: all the combinatorics was the arithmetic of the PID K[X]K[X].

3.5 Exercises

Exercise 3.1

(a) Show that Q\Q is not finitely generated as a Z\Z-module. (b) Show that Q\Q is torsion-free but not free. (c) Why does neither statement contradict Theorem 3.12?

Solution

Solution of Exercise 3.1.

(a) If Q=Zq1++Zqk\Q = \Z q_1 + \dots + \Z q_k, let dd be a common denominator of the qiq_i: every combination lies in 1dZ\frac1d\Z, but 12d1dZ\frac1{2d} \notin \frac1d\Z. Contradiction.

(b) Torsion-free: nq=0nq = 0 with n0n \neq 0 forces q=0q = 0 in Q\Q. Not free: any two nonzero rationals ab,cd\frac ab, \frac cd satisfy the nontrivial relation (bc)ab(ad)cd=0(bc)\frac ab - (ad)\frac cd = 0, so a basis has at most one element; QZ\Q \cong \Z would make Q=Zq\Q = \Z q cyclic, but q2Zq\frac q2 \notin \Z q. (And Q0\Q \neq 0.)

(c) Theorem 3.12 assumes finite generation, which (a) denies: no contradiction — rather, Q\Q shows the hypothesis is necessary in the statement “torsion-free \Rightarrow free”.

Exercise 3.2

List the abelian groups of order 360360 up to isomorphism, in both elementary-divisor and invariant-factor forms. How many abelian groups of order p5p^5 are there?

Solution

Solution of Exercise 3.2.

360=23325360 = 2^3\cdot3^2\cdot5. Partitions: of 33: (3),(2,1),(1,1,1)(3), (2,1), (1,1,1); of 22: (2),(1,1)(2), (1,1); of 11: (1)(1). Hence 3×2×1=63 \times 2 \times 1 = 6 groups. Elementary divisors \to invariant factors:

Z/8×Z/9×Z/5\Z/8 \times \Z/9 \times \Z/5Z/360\Z/360
Z/8×Z/3×Z/3×Z/5\Z/8 \times \Z/3 \times \Z/3 \times \Z/5Z/3×Z/120\Z/3 \times \Z/120
Z/4×Z/2×Z/9×Z/5\Z/4 \times \Z/2 \times \Z/9 \times \Z/5Z/2×Z/180\Z/2 \times \Z/180
Z/4×Z/2×Z/3×Z/3×Z/5\Z/4 \times \Z/2 \times \Z/3 \times \Z/3 \times \Z/5Z/6×Z/60\Z/6 \times \Z/60
(Z/2)3×Z/9×Z/5(\Z/2)^3 \times \Z/9 \times \Z/5Z/2×Z/2×Z/90\Z/2 \times \Z/2 \times \Z/90
(Z/2)3×Z/3×Z/3×Z/5(\Z/2)^3 \times \Z/3 \times \Z/3 \times \Z/5Z/2×Z/6×Z/30\Z/2 \times \Z/6 \times \Z/30

(To pass to invariant factors: the largest dsd_s collects the highest prime power of each prime, and so on down.) Of order p5p^5: as many as partitions of 55, namely 77.

Exercise 3.3

Compute the Smith normal form over Z\Z of

B=(2468),C=(2000300012),B = \begin{pmatrix} 2 & 4\\ 6 & 8 \end{pmatrix}, \qquad C = \begin{pmatrix} 2 & 0 & 0\\ 0 & 3 & 0\\ 0 & 0 & 12 \end{pmatrix},

and identify the abelian groups Z2/BZ2\Z^2/B\Z^2 and Z3/CZ3\Z^3/C\Z^3.

Solution

Solution of Exercise 3.3.

BB: D1=gcd(2,4,6,8)=2D_1 = \gcd(2,4,6,8) = 2; D2=detB=1624=8D_2 = \abs{\det B} = \abs{16 - 24} = 8. Invariant factors d1=2d_1 = 2, d2=8/2=4d_2 = 8/2 = 4: Smith form diag(2,4)\operatorname{diag}(2, 4), and Z2/BZ2Z/2Z×Z/4Z\Z^2/B\Z^2 \cong \Z/2\Z \times \Z/4\Z.

CC: diagonal but not Smith (232 \nmid 3). D1=gcd(2,3,12)=1D_1 = \gcd(2,3,12) = 1; D2=gcd(23,212,312)=gcd(6,24,36)=6D_2 = \gcd(2\cdot3,\, 2\cdot12,\, 3\cdot12) = \gcd(6, 24, 36) = 6; D3=72D_3 = 72. So d=(1,6,12)d = (1, 6, 12) and Z3/CZ3Z/6Z×Z/12Z\Z^3/C\Z^3 \cong \Z/6\Z\times\Z/12\Z — consistently with the CRT: Z/2×Z/3×Z/12Z/6×Z/12\Z/2\times\Z/3\times\Z/12 \cong \Z/6\times\Z/12.

Exercise 3.4 ★★

Let LZnL \subseteq \Z^n be a subgroup of rank nn with basis the columns of BMn(Z)B \in M_n(\Z), detB0\det B \neq 0. Show that Zn/L\Z^n/L is finite of cardinality detB\abs{\det B}, and that Zn/LiZ/diZ\Z^n/L \cong \prod_i \Z/d_i\Z for the invariant factors did_i of BB. Illustrate with L=Z(2,0)+Z(1,3)L = \Z(2,0) + \Z(1,3).

Solution

Solution of Exercise 3.4.

Write B=Qdiag(d1,,dn)PB = Q\,\operatorname{diag}(d_1, \dots, d_n)\,P with Q,PGLn(Z)Q, P \in GL_n(\Z) (Theorem 3.8; no zero did_i since detB0\det B \neq 0). Then Zn/BZnZn/diag(d)Zn=iZ/diZ\Z^n/B\Z^n \cong \Z^n/ \operatorname{diag}(d)\Z^n = \prod_i \Z/d_i\Z (the composed isomorphism xQ1xx \mapsto Q^{-1}x of Zn\Z^n maps BZnB\Z^n onto diag(d)PZn=diag(d)Zn\operatorname{diag}(d)P\Z^n = \operatorname{diag}(d)\Z^n). Its cardinality is di=detdiag(d)=detB\prod \abs{d_i} = \abs{\det \operatorname{diag}(d)} = \abs{\det B}, as detQ,detP=±1\det Q, \det P = \pm 1. For L=Z(2,0)+Z(1,3)L = \Z(2,0) + \Z(1,3): B=(2103)B = \bigl(\begin{smallmatrix}2 & 1\\ 0 & 3\end{smallmatrix}\bigr), D1=1D_1 = 1, D2=6D_2 = 6: Z2/LZ/6Z\Z^2/L \cong \Z/6\Z, of cardinality detB=6\abs{\det B} = 6.

Exercise 3.5 ★★

Let AA be a domain. (a) Verify that T(M)T(M) is a submodule and that M/T(M)M/T(M) is torsion-free. (b) Show that the ideal (X,Y)(X, Y) of K[X,Y]K[X,Y], as a K[X,Y]K[X,Y]-module, is torsion-free but not free: the structure theorem genuinely needs the PID hypothesis.

Solution

Solution of Exercise 3.5.

(a) Submodule: done in Definition 3.11. If a(x+T(M))=0a(x + T(M)) = 0 in M/T(M)M/T(M) with a0a \neq 0, then axT(M)ax \in T(M): bax=0bax = 0 for some b0b \neq 0, and ba0ba \neq 0 (domain), so xT(M)x \in T(M): the class is zero. M/T(M)M/T(M) is torsion-free.

(b) (X,Y)K[X,Y](X, Y) \subseteq K[X,Y] is torsion-free (a submodule of the domain K[X,Y]K[X,Y] acting on itself). Suppose it were free; any two elements P,QP, Q satisfy QPPQ=0Q\cdot P - P \cdot Q = 0, a nontrivial relation when PQP \ne Q are nonzero, so a basis has one element: (X,Y)=(P)(X, Y) = (P) principal — contradicting Exercise 2.6(a). Torsion-free and finitely generated (X,YX, Y generate), yet not free: over the non-PID K[X,Y]K[X,Y], the structure theorem fails.

Exercise 3.6 ★★

(a) Show that 2Z2\Z has no direct complement in the Z\Z-module Z\Z: submodules of free modules are free (Theorem 3.5), but direct summands they need not be. (b) Show that if MAnM \subseteq A^n (AA a PID) satisfies: An/MA^n/M is torsion-free, then MM is a direct summand.

Solution

Solution of Exercise 3.6.

(a) If Z=2ZC\Z = 2\Z \oplus C, the projection ZZ/2Z\Z \to \Z/2\Z restricts to an isomorphism CZ/2ZC \cong \Z/2\Z: CC would be a subgroup of Z\Z whose nonzero element xx satisfies 2xC2Z=02x \in C \cap 2\Z = 0. But Z\Z is torsion-free: C=0C = 0, forcing Z=2Z\Z = 2\Z — false.

(b) An/MA^n/M is finitely generated and torsion-free, hence free (Theorem 3.12): An/MArA^n/M \cong A^r with basis f1,,frf_1, \dots, f_r. Choose preimages yiAny_i \in A^n of the fif_i and set F=Ay1++AyrF = Ay_1 + \dots + Ay_r. Every xAnx \in A^n has π(x)=aifi\pi(x) = \sum a_if_i, so xaiyiMx - \sum a_iy_i \in M: An=M+FA^n = M + F. If aiyiM\sum a_iy_i \in M, applying π\pi gives aifi=0\sum a_if_i = 0, hence all ai=0a_i = 0 (basis): MF=0M \cap F = 0. So An=MFA^n = M \oplus F.

Exercise 3.7 ★★

Solve in Z2\Z^2 the system

{2x+4yb1(mod20),6x+8yb2(mod20),\begin{cases} 2x + 4y \equiv b_1 \pmod{20},\\ 6x + 8y \equiv b_2 \pmod{20}, \end{cases}

for which pairs (b1,b2)(b_1, b_2) solutions exist, using the Smith form of Exercise 3.3 (invertible changes of variables on both sides).

Solution

Solution of Exercise 3.7.

The reduction of Exercise 3.3 was effective: with

L=(1031),R=(1201),LBR=(2004)L = \begin{pmatrix} 1 & 0\\ -3 & 1\end{pmatrix}, \qquad R = \begin{pmatrix} 1 & 2\\ 0 & -1 \end{pmatrix}, \qquad LBR = \begin{pmatrix} 2 & 0\\ 0 & 4\end{pmatrix}

(row operation R2R23R1R_2 \leftarrow R_2 - 3R_1, column operations C2C22C1C_2 \leftarrow C_2 - 2C_1 then C2C2C_2 \leftarrow -C_2). Setting y=R1xy = R^{-1}x (a bijection of (Z/20Z)2(\Z/20\Z)^2, RR being invertible over Z\Z), the system Bxb(mod20)Bx \equiv b \pmod{20} is equivalent to

2y1b1,4y23b1+b2(mod20).2y_1 \equiv b_1, \qquad 4y_2 \equiv -3b_1 + b_2 \pmod{20}.

The congruence kyc(mod20)ky \equiv c \pmod{20} is solvable iff gcd(k,20)c\gcd(k, 20) \mid c: solutions exist iff 2b12 \mid b_1 and 4b23b14 \mid b_2 - 3b_1, i.e. b1b_1 even and b23b1(mod4)b_2 \equiv 3b_1 \pmod 4. When solvable there are 2×4=82 \times 4 = 8 solutions modulo 2020.

Exercise 3.8 ★★

(a) Determine all similarity invariants and possible Jordan forms of a nilpotent 4×44 \times 4 matrix, sorted by the partition of 44 they realize. (b) Exhibit two 4×44\times4 complex matrices with the same characteristic and minimal polynomials that are not similar, and prove that for n3n \leq 3 this cannot happen.

Solution

Solution of Exercise 3.8.

(a) A nilpotent uu has μu=Xk\mu_u = X^k; the elementary divisors are Xk1X^{k_1} \geq \dots, one Jordan block Jki(0)J_{k_i}(0) per part of a partition of 44:

partitionJordan forminvariant factors
(4)(4)J4J_4X4X^4
(3,1)(3,1)J3J1J_3 \oplus J_1X, X3X,\ X^3
(2,2)(2,2)J2J2J_2 \oplus J_2X2, X2X^2,\ X^2
(2,1,1)(2,1,1)J2J1J1J_2 \oplus J_1 \oplus J_1X, X, X2X,\ X,\ X^2
(1,1,1,1)(1,1,1,1)00X,X,X,XX, X, X, X

(b) Take u=J2J2u = J_2\oplus J_2 and v=J2J1J1v = J_2 \oplus J_1 \oplus J_1: both have χ=X4\chi = X^4, μ=X2\mu = X^2, but different invariant factors — not similar (Theorem 3.16); one can also compare ranks: rku=21=rkv\operatorname{rk} u = 2 \neq 1 = \operatorname{rk} v. For n3n \leq 3: χ\chi and μ\mu determine, for each eigenvalue λ\lambda (over a splitting field), the total size mλ3m_\lambda \leq 3 of the λ\lambda-blocks and the largest block rλr_\lambda; a partition of m3m \leq 3 is determined by its largest part (m=3,r=2m = 3, r = 2 forces (2,1)(2,1), etc.). So the elementary divisors coincide, and Corollary 3.17 descends the similarity to the base field.

Exercise 3.9 ★★★

Let uL(V)u \in \mathcal L(V), dimV=n\dim V = n. Show that the following are equivalent: (i) VV is a cyclic K[X]K[X]-module (there is xx with V=K[u]xV = K[u]x, a cyclic vector); (ii) μu=χu\mu_u = \chi_u; (iii) s=1s = 1 in Theorem 3.16. Deduce that a companion matrix has a cyclic vector, and determine when a diagonal matrix has one.

Solution

Solution of Exercise 3.9.

(i)\Rightarrow(ii): if V=K[u]xV = K[u]x, then VK[X]/Ann(x)V \cong K[X]/\operatorname{Ann}(x), and Ann(x)=(μu)\operatorname{Ann}(x) = (\mu_u) (a polynomial kills xx iff it kills all of V=K[u]xV = K[u]x, since P(u)Q(u)x=Q(u)P(u)xP(u)Q(u)x = Q(u)P(u)x). So n=dimV=degμun = \dim V = \deg \mu_u; as μuχu\mu_u \mid \chi_u and degχu=n\deg\chi_u = n, monicity gives μu=χu\mu_u = \chi_u.

(ii)\Rightarrow(iii): degχu=idegPi\deg\chi_u = \sum_i \deg P_i and μu=Ps\mu_u = P_s (Theorem 3.16); equality of degrees forces s=1s = 1.

(iii)\Rightarrow(i): VK[X]/(P1)V \cong K[X]/(P_1) is cyclic, generated by the preimage of 1ˉ\bar 1.

A companion matrix is the case V=K[X]/(P)V = K[X]/(P) itself: x=1ˉx = \bar 1, i.e. e1e_1, is cyclic. For a diagonal matrix diag(λ1,,λn)\operatorname{diag}(\lambda_1, \dots, \lambda_n): χ=(Xλi)\chi = \prod (X - \lambda_i), μ=λ distinct(Xλ)\mu = \prod_{\lambda \text{ distinct}} (X - \lambda); they agree iff the λi\lambda_i are pairwise distinct: a diagonal matrix has a cyclic vector iff its diagonal entries are pairwise distinct (then x=(1,,1)x = (1, \dots, 1) works: Vandermonde).

Exercise 3.10 ★★★

For MMn(Z)M \in M_n(\Z) viewed as an endomorphism of Zn\Z^n, prove the index formula: if detM0\det M \ne 0, then [Zn:MZn]=detM[\Z^n : M\Z^n] = \abs{\det M}, and deduce that MGLn(Z)M \in GL_n(\Z) iff detM=±1\det M = \pm 1. Application: the group Z2\Z^2 has exactly σ1(m)=dmd\sigma_1(m) = \sum_{d \mid m} d subgroups of index mm. (Count matrices in Hermite form (ab0d)\bigl(\begin{smallmatrix} a & b\\ 0 & d\end{smallmatrix}\bigr), ad=mad = m, 0b<d0 \leq b < d.)

Solution

Solution of Exercise 3.10.

Smith: M=Qdiag(d1,,dn)PM = Q\operatorname{diag}(d_1,\dots,d_n)P; Exercise 3.4 gives [Zn:MZn]=di=detM[\Z^n : M\Z^n] = \prod\abs{d_i} = \abs{\det M}. If detM=±1\det M = \pm 1: the adjugate formula M1=(detM)1t ⁣com(M)M^{-1} = (\det M)^{-1}\,{}^{t}\!\operatorname{com}(M) has integer entries, so MGLn(Z)M \in GL_n(\Z); conversely MM1=IMM^{-1} = I gives detMdetM1=1\det M \cdot \det M^{-1} = 1 in Z\Z, so detM=±1\det M = \pm1.

Subgroups of index mm in Z2\Z^2: such a subgroup LL has rank 22 (finite index) and a unique basis in Hermite normal form (ab0d)\bigl(\begin{smallmatrix} a & b\\ 0 & d\end{smallmatrix}\bigr): dd is characterized by L({0}×Z)={0}×dZL \cap (\{0\} \times \Z) = \{0\} \times d\Z, aa by π1(L)=aZ\pi_1(L) = a\Z (first coordinates), and bb is then unique modulo dd; normalize a,d>0a, d > 0 and 0b<d0 \leq b < d. The index is ad=mad = m. Counting: for each divisor dmd \mid m (a=m/da = m/d), there are dd choices of bb: total dmd=σ1(m)\sum_{d \mid m} d = \sigma_1(m).

Exercise 3.11 ★★

(Equations xk=ex^k = e in abelian groups) Let GG be a finite abelian group with invariant factors d1d2dsd_1 \mid d_2 \mid \dots \mid d_s. (a) Show that for every k1k \geq 1,

#{xG:xk=e}  =  i=1sgcd(k,di).\#\{x \in G : x^k = e\} \;=\; \prod_{i=1}^{s}\gcd(k, d_i) .

(b) Deduce: a finite abelian group is cyclic if and only if for every kk, the equation xk=ex^k = e has at most kk solutions. (c) Recover the cyclicity of finite subgroups of K×K^\times (KK a field, Chapter 4): why does the polynomial Xk1X^k - 1 guarantee the criterion of (b)?

Solution

Solution of Exercise 3.11.

(a) By the structure theorem, GiZ/diZG \cong \prod_i\Z/d_i\Z, and xk=ex^k = e decouples coordinatewise. In Z/dZ\Z/d\Z: kx0(modd)kx \equiv 0 \pmod d has exactly gcd(k,d)\gcd(k, d) solutions (xx must be a multiple of d/gcd(k,d)d/\gcd(k,d), and there are gcd(k,d)\gcd(k, d) of those). Multiply over the factors.

(b) If G=Z/dsZG = \Z/d_s\Z is cyclic (s=1s = 1), the count is gcd(k,ds)k\gcd(k, d_s) \leq k. If s2s \geq 2: take k=d1k = d_1; the count is igcd(d1,di)=d1s>d1\prod_i\gcd(d_1, d_i) = d_1^{\,s} > d_1 (each gcd\gcd equals d1d_1 by the divisibility chain): the equation xd1=ex^{d_1} = e has more than d1d_1 solutions.

(c) In a field, Xk1X^k - 1 has at most kk roots (Chapter 2: a nonzero polynomial of degree kk over a domain), so every finite subgroup GK×G \leq K^\times satisfies the criterion of (b): GG is cyclic — the one-line structural proof of the cyclicity of Fq×\mathbb F_q^\times, complementing the counting proof of Chapter 4.

Exercise 3.12 ★★★

(Elementary matrices generate) (a) Show that MMn(Z)M \in M_n(\Z) is invertible in Mn(Z)M_n(\Z) iff detM=±1\det M = \pm1. (b) Show that SL2(Z)SL_2(\Z) is generated by the two elementary matrices E=(1101)E = \bigl(\begin{smallmatrix}1 & 1\\ 0 & 1\end{smallmatrix}\bigr) and F=(1011)F = \bigl(\begin{smallmatrix}1 & 0\\ 1 & 1\end{smallmatrix}\bigr). (Run the Euclidean algorithm on the first column of MSL2(Z)M \in SL_2(\Z) by left multiplications by powers of E,FE, F, reaching ±(101)\pm\bigl(\begin{smallmatrix}1 & *\\ 0 & 1\end{smallmatrix}\bigr); finish by hand — note I=(EF1E)2-I = (EF^{-1}E)^2.) (c) Explain the connection with Smith reduction: over Z\Z, row and column operations of determinant 11 suffice to diagonalize, up to signs.

Solution

Solution of Exercise 3.12.

(a) If MN=IMN = I with integer NN: detMdetN=1\det M\det N = 1 with both integers, so detM=±1\det M = \pm1. Conversely if detM=±1\det M = \pm1, the cofactor formula M1=1detMt ⁣com(M)M^{-1} = \frac1{\det M}\,{}^t\!\operatorname{com}(M) has integer entries.

(b) Left multiplication by EkE^{-k} subtracts kk times row 22 from row 11; by FkF^{-k}, kk times row 11 from row 22. Given M=(ac)SL2(Z)M = \bigl(\begin{smallmatrix}a & *\\ c & *\end{smallmatrix}\bigr) \in SL_2(\Z), the first column (a,c)(a, c) is a unimodular vector (gcd(a,c)=1\gcd(a, c) = 1: it divides detM=1\det M = 1). Run Euclid on (a,c)(a, c) by these row operations: after finitely many steps the column becomes (±1,0)(\pm1, 0). The matrix is now ±(1b01)=±Eb\pm \bigl(\begin{smallmatrix}1 & b\\ 0 & 1\end{smallmatrix} \bigr) = \pm E^{b} (the determinant stayed 11). It remains to write I-I in the generators: (EF1E)2=(0110)2=I(EF^{-1}E)^2 = \bigl(\begin{smallmatrix}0 & 1\\ -1 & 0\end{smallmatrix}\bigr)^2 = -I (check the square of the rotation matrix). Unwinding, MM is a word in E±1,F±1E^{\pm1}, F^{\pm1}.

(c) The Smith algorithm (Method 3.9) uses exactly such row and column operations (plus swaps and sign changes, themselves products of elementary operations up to determinant sign): over Z\Z, every matrix is UDVU\,D\,V with U,VU, V products of elementary matrices and DD the Smith form — (b) is the 2×22\times2, determinant-11 instance of the general fact that EE-type matrices generate SLn(Z)SL_n(\Z).

3.6 Problem: the commutant and the double commutant

Problem 3.1

Weekend problem — rational form, commutant, bicommutant

Let KK be a field, VV a KK-vector space of dimension n1n \geq 1, and uL(V)u \in \mathcal L(V). We study the commutant

C(u)={vL(V):uv=vu},\mathcal C(u) = \{v \in \mathcal L(V) : uv = vu\},

a subalgebra of L(V)\mathcal L(V) containing K[u]={P(u):PK[X]}K[u] = \{P(u) : P \in K[X]\}, and we prove Frobenius’ dimension formula and the double commutant theorem: C(C(u))=K[u]\mathcal C(\mathcal C(u)) = K[u]. Throughout, VV is the K[X]K[X]-module defined by uu, with invariant factors P1PsP_1 \mid \cdots \mid P_s and cyclic decomposition V=i=1sViV = \bigoplus_{i=1}^s V_i, Vi=K[u]xiK[X]/(Pi)V_i = K[u]\,x_i \cong K[X]/(P_i), ni=degPin_i = \deg P_i (Theorem 3.16).

Part I — The presentation matrix XIMXI - M, and warm-ups.

  1. Let MMn(K)M \in M_n(K) and let φ ⁣:K[X]nVM=Kn\varphi \colon K[X]^n \to V_M = K^n send (Q1,,Qn)(Q_1, \dots, Q_n) to iQi(M)ei\sum_i Q_i(M)e_i. Show that φ\varphi is a surjective morphism of K[X]K[X]-modules and that every column of XInMXI_n - M lies in kerφ\ker\varphi.
  2. Show that, modulo the columns of XInMXI_n - M, every element of K[X]nK[X]^n is congruent to a constant vector (reduce degrees using XeijmjiejXe_i \equiv \sum_j m_{ji}e_j), and deduce kerφ=(XInM)K[X]n\ker\varphi = (XI_n - M)\,K[X]^n: the module VMV_M has presentation matrix XInMXI_n - M. Recover Corollary 3.17’s starting point: the similarity invariants of MM are the nonunit invariant factors of XInMXI_n - M.
  3. Compute the similarity invariants of: a scalar matrix λIn\lambda I_n; a diagonal matrix with distinct diagonal entries; the n×nn \times n Jordan block Jn(0)J_n(0); diag(J2(0),J1(0))\operatorname{diag}(J_2(0), J_1(0)) for n=3n = 3.
  4. Show that dimK[u]=degμu=ns\dim K[u] = \deg \mu_u = n_s.

Part II — Morphisms between cyclic modules.

  1. Let P,QP, Q be monic nonconstant. Show that a K[X]K[X]-morphism f ⁣:K[X]/(P)K[X]/(Q)f \colon K[X]/(P) \to K[X]/(Q) is determined by f(1ˉ)f(\bar 1), and that cˉK[X]/(Q)\bar c \in K[X]/(Q) can serve as f(1ˉ)f(\bar 1) iff Pcˉ=0P\bar c = 0 in K[X]/(Q)K[X]/(Q).
  2. Deduce

    HomK[X](K[X]/(P),K[X]/(Q))    K[X]/(gcd(P,Q)),\operatorname{Hom}_{K[X]}\bigl(K[X]/(P),\, K[X]/(Q)\bigr) \;\cong\; K[X]\big/\bigl(\gcd(P, Q)\bigr),

    of dimension deggcd(P,Q)\deg \gcd(P, Q) over KK. (Show that the solutions cˉ\bar c of Pcˉ=0P\bar c = 0 in K[X]/(Q)K[X]/(Q) form the cyclic submodule generated by Q/gcd(P,Q)\overline{Q/\gcd(P,Q)}.)

  3. Prove Frobenius’ formula:

    dimKC(u)=i,j=1sdeggcd(Pi,Pj)=i=1s(2s2i+1)ni.\dim_K \mathcal C(u) = \sum_{i,j=1}^{s} \deg\gcd(P_i, P_j) = \sum_{i=1}^{s} (2s - 2i + 1)\, n_i .

    (A commuting vv is exactly a K[X]K[X]-endomorphism of VV; decompose End(iVi)\operatorname{End}(\bigoplus_i V_i) as matrices of morphisms VjViV_j \to V_i and use the divisibility chain.)

  4. Deduce dimC(u)n\dim \mathcal C(u) \geq n, with equality iff uu is cyclic (s=1s = 1), and compute dimC(u)\dim\mathcal C(u) for u=λidu = \lambda\,\mathrm{id}: both extremes of the formula.
  5. Verify Frobenius’ formula directly for diag(J2(0),J1(0))\operatorname{diag}(J_2(0), J_1(0)) by computing the commutant explicitly as 3×33\times3 matrices.

Part III — The double commutant theorem. Let wC(C(u))w \in \mathcal C(\mathcal C(u)); we prove wK[u]w \in K[u].

  1. Show K[u]C(C(u))K[u] \subseteq \mathcal C(\mathcal C(u)), and that every wC(C(u))w \in \mathcal C(\mathcal C(u)) commutes with uu — so the inclusion to be proved, C(C(u))K[u]\mathcal C(\mathcal C(u)) \subseteq K[u], is a genuine sharpening of wC(u)w \in \mathcal C(u).
  2. Suppose first that uu is cyclic, V=K[u]xV = K[u]x. Show directly that C(u)=K[u]\mathcal C(u) = K[u] (evaluate a commuting vv on xx: v(x)=P(u)xv(x) = P(u)x for some PP, and compare vv with P(u)P(u) on the basis ukxu^k x), and conclude the theorem in this case.
  3. Back to the general case. For each ii, let πi ⁣:VVi\pi_i\colon V \to V_i be the projection along the other summands. Show πiC(u)\pi_i \in \mathcal C(u), and deduce that ww preserves each ViV_i and commutes with ui=uViu_i = u\restriction_{V_i}; conclude via question 11 applied to the cyclic uiu_i: there are polynomials QiQ_i with wVi=Qi(u)Viw\restriction_{V_i} = Q_i(u)\restriction_{V_i}.
  4. It remains to glue the QiQ_i into one polynomial. For iji \leq j (so PiPjP_i \mid P_j), show that ηij ⁣:VjVi\eta_{ij} \colon V_j \to V_i, R(u)xjR(u)xiR(u)x_j \mapsto R(u)x_i, is a well-defined K[X]K[X]-morphism (what must be checked is that R(u)xj=0R(u)x_j = 0 implies R(u)xi=0R(u)x_i = 0), and that η~ij=ηijπj\tilde\eta_{ij} = \eta_{ij}\circ\pi_j, extended by 00 on the other summands, lies in C(u)\mathcal C(u).
  5. Using wη~ij=η~ijww\tilde\eta_{ij} = \tilde\eta_{ij}w, show QiQj(modPi)Q_i \equiv Q_j \pmod{P_i} for iji \leq j. Deduce that Q=QsQ = Q_s satisfies QQi(modPi)Q \equiv Q_i \pmod {P_i} for all ii, hence w=Q(u)w = Q(u) on every ViV_i, hence on VV:

     C(C(u))=K[u]. \boxed{\ \mathcal C(\mathcal C(u)) = K[u].\ }
  6. (Coda) Deduce from the theorem: if vv commutes with every matrix commuting with uu, and uu is cyclic, then vv is a polynomial in uu; and give an example showing C(u)=K[u]\mathcal C(u) = K[u] fails for u=idu = \mathrm{id}, n2n \geq 2 — where exactly does cyclicity enter?

Part IV — Dividends of the similarity invariants. The rational canonical form is a machine; here are five of its classical outputs.

  1. (Transpose) Show that every MMn(K)M \in M_n(K) is similar to its transpose tM{}^tM. (The operations that bring XIMXI - M to Smith form, transposed, bring XItMXI - {}^tM to the same Smith form: equal similarity invariants.)
  2. (Descent of similarity) Let KLK \subseteq L be a field extension and M,NMn(K)M, N \in M_n(K). Show that if MM and NN are similar over LL, they are similar over KK. (The Smith form of XIMXI - M computed in K[X]K[X] is still a Smith form in L[X]L[X] — why do the invariant factors not change?) Consequence worth memorizing: two real matrices conjugate in GLn(C)GL_n(\C) are conjugate in GLn(R)GL_n(\R).
  3. (Nilpotent classification) Let uu be nilpotent. Show that the number of blocks of size k\geq k in its decomposition into nilpotent Jordan blocks equals rkuk1rkuk\operatorname{rk}u^{k-1} - \operatorname{rk}u^k, and deduce: nilpotent classes of Mn(K)M_n(K), for any field KK, are in bijection with the partitions of nn. How many nilpotent classes in M5(K)M_5(K)?
  4. (A concrete pair) Determine the similarity invariants of the derivation D ⁣:PPD\colon P \mapsto P' acting on the space Kn1[X]K_{n-1}[X] of polynomials of degree <n< n: (a) for K=QK = \Q; (b) for K=FpK = \mathbb F_p with p<np < n (in characteristic pp, (Xp)=0(X^p)' = 0: compute kerDk\ker D^k and use question 18).
  5. (Conjugacy classes of GL2(Fq)GL_2(\mathbb F_q)) Using invariant factors, show that every class of GL2(Fq)GL_2(\mathbb F_q) is of exactly one of four types: central aIaI; diagonalizable with two distinct eigenvalues aba \neq b in Fq×\mathbb F_q^\times; non-semisimple with minimal polynomial (Xa)2(X - a)^2; cyclic with irreducible characteristic polynomial.
  6. Count the classes of each type and conclude: GL2(Fq)GL_2(\mathbb F_q) has exactly q21q^2 - 1 conjugacy classes. (Count monic irreducible quadratics over Fq\mathbb F_q; unordered pairs {a,b}\{a, b\}; remember invertibility constrains constant terms.)
  7. (Cyclic is generic) Show that MM2(Fq)M \in M_2(\mathbb F_q) fails to be cyclic iff MM is scalar, and deduce that a uniformly random 2×22\times2 matrix over Fq\mathbb F_q is cyclic with probability 1q31 - q^{-3}. State the analogous heuristic for MnM_n and large qq (no proof required): non-cyclic matrices are rare — which is why Problem 3.1’s Part III needed real work only past the generic case.

Part V — Complements.

  1. (Center of the commutant) Show that the center of the algebra C(u)\mathcal C(u) is exactly K[u]K[u] (combine the two inclusions of Part III). Deduce that C(u)\mathcal C(u) is commutative iff uu is cyclic — recovering the equality case of question 8 by a purely structural route.
  2. (Which dimensions occur?) Deduce from Frobenius’ formula that dimC(u)n(mod2)\dim\mathcal C(u) \equiv n \pmod 2 for every uu. Then determine the exact set of values taken by dimC(u)\dim \mathcal C(u) as uu ranges over L(V)\mathcal L(V) with dimV=4\dim V = 4: show it is {4,6,8,10,16}\{4, 6, 8, 10, 16\} (enumerate the degree sequences n1nsn_1 \leq \dots \leq n_s summing to 44 and realize each by a nilpotent). In particular 1212 and 1414, though of the right parity, are not attained: the parity constraint is necessary but not sufficient.
  3. (Class equation of GL2(F3)GL_2(\mathbb F_3)) For q=3q = 3, compute the size of each conjugacy class of question 20 via orbit–stabilizer: the centralizer of a cyclic MM in GL2(Fq)GL_2(\mathbb F_q) is the unit group of K[M]K[M] (question 11). Identify K[M]K[M] in the three noncentral types, list the three monic irreducible quadratics over F3\mathbb F_3, and verify the class equation

    48=GL2(F3)=21+112+28+36,48 = \abs{GL_2(\mathbb F_3)} = 2\cdot1 + 1\cdot12 + 2\cdot8 + 3\cdot6,

    with 2+1+2+3=8=q212 + 1 + 2 + 3 = 8 = q^2 - 1 classes, as predicted by question 21.

Solution

Solution of Problem 3.1.

1. φ\varphi is additive, and K[X]K[X]-linear: φ(X(Qi)i)=i(XQi)(M)ei=MiQi(M)ei=Xφ((Qi)i)\varphi(X \cdot (Q_i)_i) = \sum_i (XQ_i)(M)e_i = M\sum_i Q_i(M)e_i = X \cdot \varphi\bigl((Q_i)_i\bigr), the module structure of VMV_M being Xv=MvX \cdot v = Mv. It is surjective: constant vectors give all of KnK^n. Column jj of XIMXI - M is XejimijeiXe_j - \sum_i m_{ij}e_i, whose image is Mejimijei=0Me_j - \sum_i m_{ij}e_i = 0.

2. Modulo the columns, XejimijeiXe_j \equiv \sum_i m_{ij}e_i: any vector of polynomials reduces, by induction on the top degree, to a constant vector cKnc \in K^n. If the original vector is in kerφ\ker\varphi, then φ(c)=c=0\varphi(c) = c = 0 (on constants, φ\varphi is the identification Kn=VMK^n = V_M), so the vector lies in the column span: kerφ=(XInM)K[X]n\ker\varphi = (XI_n - M)K[X]^n. Hence VMK[X]n/(XIM)K[X]nV_M \cong K[X]^n/(XI - M)K[X]^n, and Smith over K[X]K[X] (all invariant factors nonzero, their product being det(XIM)=χM\det(XI - M) = \chi_M) gives VMiK[X]/(fi)V_M \cong \bigoplus_i K[X]/(f_i): the nonconstant fif_i are the similarity invariants, computable as quotients of minor gcds (Theorem 3.8).

3. λIn\lambda I_n: XIλIXI - \lambda I is already Smith: invariants (Xλ,,Xλ)(X - \lambda, \dots, X - \lambda), nn of them. Distinct diagonal entries: ViK[X]/(Xλi)V \cong \bigoplus_i K[X]/(X - \lambda_i) with pairwise comaximal moduli, so CRT compresses to the single cyclic K[X]/(i(Xλi))K[X]/\bigl(\prod_i(X - \lambda_i)\bigr): one invariant, χ\chi. Jn(0)J_n(0): μ=Xn=χ\mu = X^n = \chi forces a single invariant XnX^n. diag(J2(0),J1(0))\operatorname{diag}(J_2(0), J_1(0)): elementary divisors X2,XX^2, X: invariants P1=XP2=X2P_1 = X \mid P_2 = X^2.

4. PP(u)P \mapsto P(u) maps K[X]K[X] onto K[u]K[u], with kernel (μu)(\mu_u) by definition of the minimal polynomial: K[u]K[X]/(μu)K[u] \cong K[X]/(\mu_u), of dimension degμu=degPs=ns\deg\mu_u = \deg P_s = n_s.

5. K[X]K[X]-linearity forces f(Qˉ)=f(Q1ˉ)=Qf(1ˉ)f(\bar Q) = f(Q\cdot\bar 1) = Q\,f(\bar 1). The class 1ˉ\bar 1 satisfies P1ˉ=0P \bar 1 = 0, so Pf(1ˉ)=0P f(\bar 1) = 0 is necessary. Conversely if Pcˉ=0P\bar c = 0, then f(Qˉ)=Qcˉf(\bar Q) = Q\bar c is well defined (QQmodP(QQ)cˉQ \equiv Q' \bmod P \Rightarrow (Q - Q')\bar c \in multiples of Pcˉ=0P\bar c = 0) and K[X]K[X]-linear.

6. Let g=gcd(P,Q)g = \gcd(P, Q), P=gPP = gP', Q=gQQ = gQ' with gcd(P,Q)=1\gcd(P', Q') = 1. In K[X]/(Q)K[X]/(Q): Pcˉ=0    QPc    QPc    QcP\bar c = 0 \iff Q \mid Pc \iff Q' \mid P'c \iff Q' \mid c (Euclid, gcd(P,Q)=1\gcd(P', Q') = 1). So the admissible cˉ\bar c form the submodule generated by Qˉ\bar{Q'}, whose annihilator is {R:QRQ}=(g)\{R : Q \mid RQ'\} = (g): that submodule is K[X]/(g)\cong K[X]/(g). With question 5, Hom(K[X]/(P),K[X]/(Q))K[X]/(gcd(P,Q))\operatorname{Hom}(K[X]/(P), K[X]/(Q)) \cong K[X]/(\gcd(P,Q)), of dimension deggcd(P,Q)\deg\gcd(P,Q).

7. vv commutes with uu iff vv commutes with every P(u)P(u), iff vv is K[X]K[X]-linear: C(u)=EndK[X](V)\mathcal C(u) = \operatorname{End}_{K[X]}(V). Writing morphisms of V=jVjV = \bigoplus_j V_j as matrices (fij)(f_{ij}), fijHom(Vj,Vi)f_{ij} \in \operatorname{Hom}(V_j, V_i) (compose with injections and projections), question 6 gives

dimC(u)=i,jdeggcd(Pi,Pj)=i,jnmin(i,j)=k=1s(2(sk)+1)nk,\dim \mathcal C(u) = \sum_{i,j} \deg\gcd(P_i, P_j) = \sum_{i,j} n_{\min(i,j)} = \sum_{k=1}^{s} \bigl(2(s - k) + 1\bigr)\,n_k ,

using the divisibility chain (gcd(Pi,Pj)=Pmin(i,j)\gcd(P_i, P_j) = P_{\min(i,j)}) and, for the last step, that min(i,j)=k\min(i,j) = k happens for exactly 2(sk)+12(s - k) + 1 pairs (i,j)(i, j).

8. Since 2(sk)+112(s-k)+1 \geq 1, dimC(u)knk=n\dim\mathcal C(u) \geq \sum_k n_k = n, with equality iff s=1s = 1, i.e. iff uu is cyclic (Exercise 3.9). For u=λidu = \lambda\,\mathrm{id}: s=ns = n, all nk=1n_k = 1: dim=k=1n(2(nk)+1)=n2\dim = \sum_{k=1}^n (2(n-k)+1) = n^2 — correct, since C(λid)=L(V)\mathcal C(\lambda\,\mathrm{id}) = \mathcal L(V).

9. Formula: invariants (X,X2)(X, X^2), so s=2s = 2, n1=1n_1 = 1, n2=2n_2 = 2: dim=31+12=5\dim = 3\cdot1 + 1\cdot2 = 5. Directly: in the basis (e1,e2,e3)(e_1, e_2, e_3) with ue2=e1u e_2 = e_1, ue1=ue3=0ue_1 = ue_3 = 0, writing Au=uAAu = uA for A=(aij)A = (a_{ij}) yields the conditions a21=a23=a31=0a_{21} = a_{23} = a_{31} = 0 and a11=a22a_{11} = a_{22}: five free parameters a11=a22,a12,a13,a32,a33a_{11}{=}a_{22}, a_{12}, a_{13}, a_{32}, a_{33}.

10. A polynomial P(u)P(u) commutes with anything that commutes with uu (it is a sum of powers of uu): K[u]C(C(u))K[u] \subseteq \mathcal C(\mathcal C(u)). And uC(u)u \in \mathcal C(u), so any wC(C(u))w \in \mathcal C(\mathcal C(u)) commutes with uu.

11. Let V=K[u]xV = K[u]x and vC(u)v \in \mathcal C(u). Write v(x)=P(u)xv(x) = P(u)x (cyclicity). For y=Q(u)xy = Q(u)x arbitrary: v(y)=vQ(u)x=Q(u)v(x)=Q(u)P(u)x=P(u)yv(y) = vQ(u)x = Q(u)v(x) = Q(u)P(u)x = P(u)y. So v=P(u)v = P(u): C(u)=K[u]\mathcal C(u) = K[u]. Then C(C(u))=C(K[u])=C(u)=K[u]\mathcal C(\mathcal C(u)) = \mathcal C(K[u]) = \mathcal C(u) = K[u] (commuting with all of K[u]K[u] is the same as commuting with uu). The theorem holds in the cyclic case.

12. πi\pi_i is K[X]K[X]-linear (the decomposition is a direct sum of submodules), so πiC(u)\pi_i \in \mathcal C(u), and ww commutes with it: w(Vi)=wπi(V)=πiw(V)Viw(V_i) = w\pi_i(V) = \pi_i w(V) \subseteq V_i. The restriction wi=wViw_i = w\restriction_{V_i} commutes with the cyclic ui=uViu_i = u\restriction_{V_i} (question 10), and wiC(ui)=K[ui]w_i \in \mathcal C(u_i) = K[u_i] (question 11): wi=Qi(ui)=Qi(u)Viw_i = Q_i(u_i) = Q_i(u)\restriction_{V_i} for some QiK[X]Q_i \in K[X].

13. Well-definedness of ηij(R(u)xj)=R(u)xi\eta_{ij}(R(u)x_j) = R(u)x_i: if R(u)xj=0R(u)x_j = 0 then PjRP_j \mid R, and PiPjP_i \mid P_j gives PiRP_i \mid R, so R(u)xi=0R(u)x_i = 0 (Ann(xi)=(Pi)\operatorname{Ann}(x_i) = (P_i)). ηij\eta_{ij} is then K[X]K[X]-linear by construction, and η~ij=ηijπj\tilde\eta_{ij} = \eta_{ij}\pi_j is a composition of K[X]K[X]-morphisms VVV \to V: η~ijC(u)\tilde\eta_{ij} \in \mathcal C(u).

14. Evaluate wη~ij=η~ijww\tilde\eta_{ij} = \tilde\eta_{ij}w at xjx_j: the left side is w(xi)=Qi(u)xiw(x_i) = Q_i(u)x_i; the right side is ηij(Qj(u)xj)=Qj(u)xi\eta_{ij}\bigl(Q_j(u)x_j\bigr) = Q_j(u)x_i. Hence (QiQj)(u)xi=0(Q_i - Q_j)(u)\,x_i = 0: PiQiQjP_i \mid Q_i - Q_j for all iji \leq j. In particular, with Q=QsQ = Q_s: QQi(modPi)Q \equiv Q_i \pmod{P_i}, so Q(u)Q(u) and Qi(u)Q_i(u) agree on ViV_i (which Pi(u)P_i(u) kills). Therefore w=Q(u)w = Q(u) on each ViV_i, hence on VV: C(C(u))K[u]\mathcal C(\mathcal C(u)) \subseteq K[u], and with question 10, C(C(u))=K[u]\mathcal C(\mathcal C(u)) = K[u].

15. The first assertion is questions 10–14 (or, for cyclic uu, question 11 alone). For u=idu = \mathrm{id}, n2n \geq 2: C(u)=L(V)\mathcal C(u) = \mathcal L(V) has dimension n2n^2, while dimK[u]=degμu=1\dim K[u] = \deg\mu_u = 1. So C(u)=K[u]\mathcal C(u) = K[u] fails badly; yet the bicommutant theorem holds (C(L(V))=Kid=K[u]\mathcal C(\mathcal L(V)) = K\,\mathrm{id} = K[u]: the center of the matrix algebra is the scalars). Cyclicity is what makes the single commutant already polynomial; the double commutant is polynomial always.

16. If P(XIM)Q=SP(XI - M)Q = S is a Smith reduction (P,QP, Q invertible over K[X]K[X]), transposing gives tQ(XItM)tP=tS=S{}^tQ\,(XI - {}^tM)\,{}^tP = {}^tS = S: same Smith form, so XIMXI - M and XItMXI - {}^tM have the same invariant factors, i.e. MM and tM{}^tM have the same similarity invariants (Corollary 3.17): they are similar.

17. The similarity invariants of MM over LL are the invariant factors of XIMXI - M in L[X]L[X]. A Smith reduction of XIMXI - M over K[X]K[X] — invertible P,QP, Q over K[X]K[X], diagonal with the divisibility chain — is also a valid Smith reduction over L[X]L[X] (P,QP, Q stay invertible: their determinants are nonzero constants), and monic invariant factors are unique: the invariant factors computed over KK and over LL coincide. So MLNM \sim_L N iff they have the same invariant factors iff MKNM \sim_K N. In particular C\C-conjugate real matrices are R\R-conjugate — a statement often proved analytically (specialize an invertible P+iQP + \iu Q), here structurally.

18. Decompose u=Jmt(0)u = \bigoplus J_{m_t}(0) into nilpotent Jordan blocks. In one block of size mm, rkJmk=max(mk,0)\operatorname{rk}J_m^k = \max(m - k, 0), so rkJmk1rkJmk=1\operatorname{rk}J_m^{k-1} - \operatorname{rk}J_m^k = 1 if mkm \geq k, 00 otherwise. Summing over blocks: rkuk1rkuk=#{t:mtk}\operatorname{rk}u^{k-1} - \operatorname{rk}u^k = \#\{t : m_t \geq k\}. The rank sequence therefore determines the multiset (mt)(m_t) — a partition of nn — and conversely each partition is realized: nilpotent classes \leftrightarrow partitions of nn, over every field. M5M_5: p(5)=7p(5) = 7 classes (55; 4+14{+}1; 3+23{+}2; 3+1+13{+}1{+}1; 2+2+12{+}2{+}1; 2+1+1+12{+}1{+}1{+}1; 151^5).

19. (a) Over Q\Q (or any characteristic-00 field), Dn=0D^n = 0, Dn1(Xn1)=(n1)!0D^{n-1}(X^{n-1}) = (n-1)!\, \neq 0: DD is nilpotent of index nn on an nn-dimensional space, hence cyclic with single invariant XnX^n (x=Xn1x = X^{n-1} generates: its iterated derivatives span). (b) Over Fp\mathbb F_p with p<np < n: Dp=0D^p = 0, because the pp-th derivative of every monomial XmX^m carries the factor m(m1)(mp+1)m(m-1)\cdots(m-p+1), a product of pp consecutive integers, hence 0modp\equiv 0 \bmod p. Write n=ap+rn = ap + r, 0r<p0 \leq r < p. Then kerDk\ker D^k is spanned by the monomials XmX^m with DkXm=0D^kX^m = 0; counting exponents m<nm < n by their residue mod pp: dimkerDk=ak+min(r,k)\dim\ker D^k = ak + \min(r, k) for 0kp0 \leq k \leq p, so rkDk1rkDk=dimkerDkdimkerDk1=a+1kr\operatorname{rk}D^{k-1} - \operatorname{rk}D^k = \dim\ker D^k - \dim\ker D^{k-1} = a + \mathbf 1_{k \leq r}. By question 18, the partition has aa blocks of size exactly pp and (if r>0r > 0) one block of size rr: similarity invariants XrXpXpX^r \mid X^p \mid \dots \mid X^p. Characteristic changes the canonical form of the most familiar operator in mathematics.

20. MGL2M \in GL_2 has s{1,2}s \in \{1, 2\} invariant factors. If s=2s = 2: P1=P2=XaP_1 = P_2 = X - a (a0a \neq 0: invertibility), i.e. M=aIM = aI, central. If s=1s = 1: MM is cyclic with characteristic == minimal polynomial χ\chi of degree 22, and the classes correspond to the possible χ\chi with χ(0)0\chi(0) \neq 0: χ\chi split with distinct roots aba \neq b (companion \sim diagonal); χ=(Xa)2\chi = (X - a)^2 (companion, non-semisimple); χ\chi irreducible. Exactly one type each — the invariant factors are a complete invariant.

21. Central: q1q - 1 choices of aa. Distinct split eigenvalues: unordered pairs {a,b}Fq×\{a, b\} \subseteq \mathbb F_q^\times, aba \neq b: (q12)\binom{q-1}2 classes. Minimal (Xa)2(X-a)^2: q1q - 1 classes. Irreducible quadratics with nonzero constant term: all irreducible quadratics qualify (their roots are nonzero), and there are q2q2\frac{q^2 - q}2 monic irreducible quadratics (the q2q^2 monic quadratics minus the (q2)+q=q2+q2\binom q2 + q = \frac{q^2+q}2 split ones). Total:

(q1)+(q1)(q2)2+(q1)+q2q2=q21.(q - 1) + \frac{(q-1)(q-2)}2 + (q - 1) + \frac{q^2 - q}2 = q^2 - 1 .

22. If MM is not cyclic, s=2s = 2 and MM is scalar (question 20’s dichotomy holds in M2M_2, invertible or not: two invariant factors of degree 11 with P1P2P_1 \mid P_2 and P1=P2P_1 = P_2 forces M=aIM = aI). Scalars number qq among the q4q^4 matrices: cyclic probability 1q31 - q^{-3}. In general the non-cyclic locus of MnM_n is where the (n1)×(n1)(n-1)\times(n-1) minors of XIMXI - M share a factor — a proper algebraic condition — so its proportion is O(1/q)O(1/q)-small for large qq: matrices with χ=μ\chi = \mu are the rule, and Part III’s gluing argument is the price paid for the exceptions.

23. An element of the center of C(u)\mathcal C(u) lies in C(u)\mathcal C(u) and commutes with every element of C(u)\mathcal C(u), i.e. lies in C(C(u))=K[u]\mathcal C(\mathcal C(u)) = K[u] (question 14). Conversely K[u]C(u)K[u] \subseteq \mathcal C(u), and every P(u)P(u) commutes with every vC(u)v \in \mathcal C(u) (such a vv commutes with uu, hence with each power of uu): K[u]K[u] is central in C(u)\mathcal C(u). Hence Z(C(u))=K[u]Z(\mathcal C(u)) = K[u]. Consequently C(u)\mathcal C(u) is commutative iff C(u)=Z(C(u))=K[u]\mathcal C(u) = Z(\mathcal C(u)) = K[u]; in that case dimC(u)=dimK[u]=nsn\dim\mathcal C(u) = \dim K[u] = n_s \leq n, while question 8 gives dimC(u)n\dim\mathcal C(u) \geq n: so ns=nn_s = n and s=1s = 1, i.e. uu is cyclic. Conversely, for uu cyclic question 11 gives C(u)=K[u]\mathcal C(u) = K[u], commutative. Structurally: a matrix algebra equal to its own center is exactly a polynomial algebra K[u]K[u] of a cyclic uu.

24. Each coefficient 2s2i+12s - 2i + 1 in Frobenius’ formula is odd, so

dimC(u)=i=1s(2s2i+1)nii=1sni=n(mod2).\dim\mathcal C(u) = \sum_{i=1}^s(2s - 2i + 1)\,n_i \equiv \sum_{i=1}^s n_i = n \pmod 2 .

For n=4n = 4, the possible degree sequences n1nsn_1 \leq \dots \leq n_s of the invariant factors, summing to 44, are (4)(4), (1,3)(1, 3), (2,2)(2, 2), (1,1,2)(1, 1, 2), (1,1,1,1)(1, 1, 1, 1); every one is realized, e.g. by the nilpotent with Pi=XniP_i = X^{n_i} (the divisibility chain holds automatically). The formula gives, respectively,

14=4,3+3=6,6+2=8,5+3+2=10,7+5+3+1=16.1\cdot4 = 4, \quad 3 + 3 = 6, \quad 6 + 2 = 8, \quad 5 + 3 + 2 = 10, \quad 7 + 5 + 3 + 1 = 16 .

So the value set is {4,6,8,10,16}\{4, 6, 8, 10, 16\}: even numbers of the right parity, but 1212 and 1414 never occur — between the almost-cyclic sequences and the scalar’s n2n^2 there is a gap.

25. GL2(F3)=(q21)(q2q)=86=48\abs{GL_2(\mathbb F_3)} = (q^2 - 1)(q^2 - q) = 8 \cdot 6 = 48. Central type: II and 2I2I, two classes of size 11. In the three other types MM is cyclic (question 20), so its centralizer in GL2GL_2 is the group of invertible elements of C(M)=K[M]\mathcal C(M) = K[M] (question 11), and class size =48/K[M]×=48/\abs{K[M]^\times} by orbit–stabilizer. Distinct split eigenvalues: only the pair {1,2}\{1, 2\}, one class; K[M]F3×F3K[M] \cong \mathbb F_3 \times \mathbb F_3 (CRT on χ=(X1)(X2)\chi = (X-1) (X-2)), units 22=42 \cdot 2 = 4, size 48/4=1248/4 = 12. Minimal (Xa)2(X - a)^2, a{1,2}a \in \{1, 2\}: two classes; K[M]F3[X]/((Xa)2)K[M] \cong \mathbb F_3[X]/((X-a)^2), units q2q=6q^2 - q = 6 (constant term of the unit 0\neq 0 after centering), size 48/6=848/6 = 8. Irreducible χ\chi: the monic irreducible quadratics over F3\mathbb F_3 number (93)/2=3(9 - 3)/2 = 3, namely

X2+1,X2+X+2,X2+2X+2,X^2 + 1, \qquad X^2 + X + 2, \qquad X^2 + 2X + 2,

(no roots in F3\mathbb F_3: check 0,1,20, 1, 2); three classes, K[M]F9K[M] \cong \mathbb F_9, units q21=8q^2 - 1 = 8, size 48/8=648/8 = 6. Class equation: 21+112+28+36=2+12+16+18=482\cdot1 + 1\cdot12 + 2\cdot8 + 3\cdot6 = 2 + 12 + 16 + 18 = 48; and 2+1+2+3=8=q212 + 1 + 2 + 3 = 8 = q^2 - 1 classes, matching question 21.