---
title: "Molecular Evolution and Phylogenomics"
book: "University Biology — Year 3"
subject: biology
language: en
chapter: 25
exercises: 12
source: https://one-course.com/books/biology/5/en/chapter/25-molecular-evolution-and-phylogenomics
---

# Chapter 25 — Molecular Evolution and Phylogenomics

In 1968 Motoo Kimura did a sum that unsettled a field. Comparing the haemoglobins, cytochromes and other proteins of mammals whose common ancestors were dated by fossils, he found that amino acids had been replaced at a rate of roughly one substitution per site per billion years — which, over a mammalian genome and a generation, came to a new substitution fixed in the population every two years or so. Haldane had shown a decade earlier that natural selection can fix at most about one substitution per three hundred generations before the cost in lost offspring becomes unpayable. The arithmetic left one way out: most of the changes that accumulate in DNA are not driven by selection at all, but are neutral, fixed by chance in finite populations, at a rate that — as this chapter’s first theorem shows — is simply the mutation rate. The [neutral theory](#thm-b3-molecular-evolution-neutral) became the null hypothesis of molecular evolution, the baseline against which selection is detected; the [molecular clock](#prop-b3-molecular-evolution-clock) became a way of dating what fossils could not; and the genomes sequenced since have given the tools to see, gene by gene, where selection has acted, how new genes are born from old, and why the tree of a gene is not always the tree of its species.

## 25.1 Drift, mutation and the neutral rate

**Theorem 25.1 (The neutral substitution rate).**

In a diploid population of $N$ individuals, a new mutation with no effect on fitness has probability $1/2N$ of eventually replacing all other alleles (*fixation*) and $1 - 1/2N$ of being lost. If each of the $2N$ gene copies mutates at rate $\mu$ per generation, $2N\mu$ new neutral mutations arise per generation and, in the long run, the rate at which substitutions accumulate is

$$
k = 2N\mu\times\frac{1}{2N} = \mu :
$$

the neutral *[substitution rate](#thm-b3-molecular-evolution-neutral)* equals the mutation rate, independent of population size. A mutation destined to fix takes on average $4N$ generations to do so, so large populations hold more variation in transit but fix it no faster. For a mutation with selective advantage $s$, Kimura’s formula gives the [fixation probability](#thm-b3-molecular-evolution-neutral)

$$
P(s) = \frac{1 - \mathrm{e}^{-2s}}{1 - \mathrm{e}^{-4Ns}} \approx 2s
\quad (4Ns \gg 1),
$$

so an advantage of $1\,\%$ is fixed with probability about $2\,\%$ — forty times a neutral mutation’s chance in a population of two thousand, but still lost forty-nine times in fifty; and a disadvantage $s < 0$ with $4N|s| \gg 1$ is fixed essentially never, while one with $4N|s| \ll 1$ behaves as neutral: selection sees only what drift does not swamp, and the boundary is $|s| \sim
1/4N$ — *effectively neutral*.

**Proof.** Neutrality: some copy present today will be the ancestor of the whole population in the distant future, and by symmetry each of the $2N$ copies is equally likely to be it; the new mutant is one copy, hence $1/2N$. Rate: substitutions per generation $=$ (mutations arising per generation) $\times$ (probability each fixes) $= 2N\mu/2N$. Kimura’s formula follows from the diffusion approximation to the change of allele frequency under drift and selection, admitted here; its limits are checked directly: as $s \to 0$, $(2s)/(4Ns) = 1/2N$, the neutral value; for $4Ns \gg 1$ the denominator is $1$ and the numerator $\approx 2s$. Fixation time: the frequency of a neutral allele performs a random walk with variance $p(1-p)/2N$ per generation, and the expected time to reach $1$ from $1/2N$, conditional on doing so, is $4N$ generations (Kimura and Ohta, 1969), admitted. ∎

![Kimura’s fixation probability against the scaled selection coefficient, for N = 1000. Near zero the curve passes through the neutral value; to the right it rises toward 2s; to the left it falls off a cliff. Selection acts only on what drift cannot hide.](https://one-course.com/images/onecourse/chapters/biology-5/b3-molecular-evolution/fig-90cf12a4f1e1.svg)

*Kimura’s [fixation probability](#thm-b3-molecular-evolution-neutral) against the scaled selection coefficient, for $N = 1000$. Near zero the curve passes through the neutral value; to the right it rises toward $2s$; to the left it falls off a cliff. Selection acts only on what drift cannot hide.*

**Evidence.** Kimura (1968) and King and Jukes (1969) made the argument from rates: the observed rate of amino-acid substitution across mammalian proteins implied more substitutions per generation than Haldane’s cost of selection allowed, so most had to be neutral. The theory’s predictions then held: rates are highest where function is least constrained — synonymous sites, introns, [pseudogenes](#def-b3-molecular-evolution-duplication), the third codon position — and lowest in [histones](https://one-course.com/books/biology/5/en/chapter/1-chromatin-and-epigenetics#def-b3-chromatin-epigenetics-nucleosome) and ubiquitin, which change one residue in a hundred million years; the level of variation within species tracks mutation and population size; and the rate of substitution per year is roughly constant across lineages with very different population sizes, as the formula $k = \mu$ requires and a selectionist account does not. The controversy that followed did not overturn it: it fixed the neutral rate as the null against which selection is measured. ∎

![Left: Motoo Kimura (1924–1994), who showed that most molecular change is fixed by chance and at the mutation rate (photograph, 1986, CC BY 4.0). Right: an Antarctic icefish, whose blood is kept from freezing by an antifreeze glycoprotein assembled, some ten million years ago, from a duplicated digestive-enzyme gene and a run of repeats.](https://one-course.com/images/onecourse/chapters/biology-5/b3-molecular-evolution/img-155e3db867ae.jpg)

![Left: Motoo Kimura (1924–1994), who showed that most molecular change is fixed by chance and at the mutation rate (photograph, 1986, CC BY 4.0). Right: an Antarctic icefish, whose blood is kept from freezing by an antifreeze glycoprotein assembled, some ten million years ago, from a duplicated digestive-enzyme gene and a run of repeats.](https://one-course.com/images/onecourse/chapters/biology-5/b3-molecular-evolution/img-2002e8221dec.jpg)

*Left: Motoo Kimura (1924–1994), who showed that most molecular change is fixed by chance and at the mutation rate (photograph, 1986, CC BY 4.0). Right: an Antarctic icefish, whose blood is kept from freezing by an antifreeze glycoprotein assembled, some ten million years ago, from a duplicated digestive-enzyme gene and a run of repeats.*

## 25.2 The molecular clock

**Proposition 25.2 (The clock and its irregularities).**

If substitutions accumulate at a constant rate $k$ per site per year in two lineages that split $T$ years ago, the fraction of sites at which they differ is, for small values, $d \approx 2kT$: the divergence of two sequences is a *[molecular clock](#prop-b3-molecular-evolution-clock)* (Zuckerkandl and Pauling, 1965), which can be *calibrated* by one fossil-dated split and then read for splits with no fossils. The clock is stochastic — a Poisson process, so $d$ has variance about equal to its mean over a fixed number of sites — and it is irregular in three known ways. Each protein has its own rate, set by the fraction of its sites that function permits to change: fibrinopeptides at $8$, haemoglobin $1$, cytochrome $c$ $0.3$, [histone](https://one-course.com/books/biology/5/en/chapter/1-chromatin-and-epigenetics#def-b3-chromatin-epigenetics-nucleosome) H4 $0.01$ substitutions per site per billion years. Rates differ between lineages: since the neutral rate is $\mu$ per generation, animals with short generations (rodents) accumulate more change per year than those with long (primates, whales) — the *[generation-time effect](#prop-b3-molecular-evolution-clock)*, partly offset because per-generation mutation rates rise with generation length. And $d$ saturates: once many sites have changed, further changes hit sites already changed, so the raw difference must be corrected for multiple hits, as the [alignment](https://one-course.com/books/biology/5/en/chapter/5-bioinformatics-and-sequence-analysis#def-b3-bioinformatics-alignment) chapter’s distances do. Modern *[relaxed clocks](#prop-b3-molecular-evolution-clock)* let the rate vary between branches within a statistical model and are calibrated by many fossils at once; they date the split of humans and chimpanzees at six to seven million years, of placental mammal orders near the end of the Cretaceous, and of animals and fungi in the Precambrian, with uncertainties of ten to twenty per cent that come mostly from the fossils.

**Proof.** Two lineages each accumulate $kT$ substitutions per site, so the difference between them is $2kT$ where sites have not been hit twice; a fossil-dated pair gives $k = d/2T$. The Poisson variance and the per-protein rates are Zuckerkandl and Pauling’s, Dickerson’s (1971) and Kimura’s fits to vertebrate protein sets; the saturation correction is the Jukes–Cantor formula of the [alignment](https://one-course.com/books/biology/5/en/chapter/5-bioinformatics-and-sequence-analysis#def-b3-bioinformatics-alignment) chapter, $d_{\text{true}} =
-\tfrac{3}{4}\ln(1 - \tfrac{4}{3}p)$. Rate variation between lineages was shown by relative-rate tests: with an outgroup $O$ and two species $A$, $B$, the distances $d_{AO} - d_{BO}$ should be zero under a strict clock, and are not for rodents against primates. ∎

![The molecular clock, protein by protein. Each runs at its own rate, set by how much of the molecule function leaves free to change: a fibrinopeptide, cut away and discarded when blood clots, changes hundreds of times faster than the histone that packs DNA.](https://one-course.com/images/onecourse/chapters/biology-5/b3-molecular-evolution/fig-88f03e2887d1.svg)

*The [molecular clock](#prop-b3-molecular-evolution-clock), protein by protein. Each runs at its own rate, set by how much of the molecule function leaves free to change: a fibrinopeptide, cut away and discarded when blood clots, changes hundreds of times faster than the [histone](https://one-course.com/books/biology/5/en/chapter/1-chromatin-and-epigenetics#def-b3-chromatin-epigenetics-nucleosome) that packs DNA.*

## 25.3 Reading selection in sequences

**Definition 25.3 (dN/dS).**

In a protein-coding gene a substitution is *synonymous* if it leaves the amino acid unchanged and *nonsynonymous* if it changes it. Because the genetic code lets most third positions and some first positions change silently, about a quarter of the possible mutations in a typical gene are synonymous. Let $d_{S}$ be the number of synonymous substitutions per synonymous site and $d_{N}$ the nonsynonymous substitutions per nonsynonymous site, each corrected for multiple hits. Synonymous changes are close to neutral, so $d_{S}$ estimates $2\mu T$ — the clock; and the ratio

$$
\omega = \frac{d_{N}}{d_{S}}
$$

measures what selection has done to the protein: $\omega = 1$ if amino-acid changes are as free as silent ones (no constraint, as in a [pseudogene](#def-b3-molecular-evolution-duplication)); $\omega < 1$ under *purifying selection*, most changes being removed — the typical gene has $\omega \approx 0.1$ to $0.2$, [histones](https://one-course.com/books/biology/5/en/chapter/1-chromatin-and-epigenetics#def-b3-chromatin-epigenetics-nucleosome) $0.001$; $\omega > 1$ under *positive selection*, amino-acid changes being favoured faster than drift alone allows — the antigen-binding groove of MHC molecules, the surface proteins of viruses in an arms race with antibodies, the sperm and egg proteins of fertilisation, and the lysozyme of ruminants, which was recruited to digest bacteria in the stomach. Computed along one branch of a tree or site by site, $\omega$ maps where and when a protein was changed by selection rather than by chance.

![Orders of magnitude of . Almost every gene sits well below one, its amino acids guarded by purifying selection; a pseudogene drifts at one; the few above one are the proteins in arms races or newly recruited to a job.](https://one-course.com/images/onecourse/chapters/biology-5/b3-molecular-evolution/fig-5beff5caeedd.svg)

*Orders of magnitude of $\omega$. Almost every gene sits well below one, its amino acids guarded by [purifying selection](#def-b3-molecular-evolution-dnds); a [pseudogene](#def-b3-molecular-evolution-duplication) drifts at one; the few above one are the proteins in arms races or newly recruited to a job.*

**Method 25.4 (Estimating dN/dSd_{N}/d_{S}dN​/dS​).**

(1) Align the two coding sequences codon by codon (align the proteins, then map back to the DNA, so that no gap breaks the reading frame). (2) For each codon count the synonymous and nonsynonymous *sites*: each of its three positions is scored by the fraction of possible changes that are silent — a fourfold-degenerate third position counts as one synonymous site, a position where every change alters the amino acid as one nonsynonymous site, a twofold position as a third of one and two thirds of the other. Sum over codons: $S$ synonymous and $N$ nonsynonymous sites, with $S + N = 3 \times$ codons. (3) Count the synonymous and nonsynonymous *differences* between the sequences, resolving codons that differ at two positions by averaging over the possible paths. (4) $p_{S} =$ differences$/S$, $p_{N} =$ differences$/N$; correct each for multiple hits with Jukes–Cantor. (5) $\omega =
d_{N}/d_{S}$; test $\omega \ne 1$ against the sampling variance, which is large when $d_{S}$ is small — two closely related sequences give an unreliable ratio, and two very distant ones have a saturated $d_{S}$. (6) For where and when: fit $\omega$ per branch of a tree, or per site with a mixture of site classes, and test whether a class with $\omega > 1$ improves the fit.

## 25.4 New genes from old

**Definition 25.5 (Duplication, families and transfer).**

Most new genes are copies of old ones. A *gene duplication* — by unequal crossing-over, retrotransposition, or the doubling of a whole genome, as happened twice at the origin of vertebrates and again in the ancestor of teleost fish and in many plant lineages — leaves two [paralogues](https://one-course.com/books/biology/5/en/chapter/4-genomics-and-sequencing#def-b3-genomics-comparative) where there was one; genes in different species descended from a single ancestral gene are [orthologues](https://one-course.com/books/biology/5/en/chapter/4-genomics-and-sequencing#def-b3-genomics-comparative). The copy’s usual fate is decay: freed from constraint ($\omega \to 1$) it accumulates a stop codon or a frameshift and becomes a *pseudogene*, of which the human genome carries some twenty thousand. Sometimes both survive: by *neofunctionalisation*, one copy acquiring a new function while the other keeps the old (the antifreeze glycoprotein of icefish, from a trypsinogen gene; the red and green opsins of primates, from a duplication forty million years ago that gave the trichromatic vision the Year 1 volume described; the lens crystallins, from metabolic enzymes); or by *subfunctionalisation*, each copy keeping part of the ancestor’s expression or activity so that both are needed. Repeated duplication builds *gene families* — the globins, the Hox clusters of the development chapter, the thousand [olfactory receptor](https://one-course.com/books/biology/5/en/chapter/18-sensory-systems#def-b3-sensory-systems-chemical) genes of a mouse, a third of them pseudogenes in humans. The other source of new genes is *horizontal transfer*: bacteria acquire genes from unrelated bacteria by [plasmids](https://one-course.com/books/biology/5/en/chapter/6-genetic-engineering-and-biotechnology#def-b3-genetic-engineering-tools), phages and free DNA, which is how [antibiotic resistance](https://one-course.com/books/biology/5/en/chapter/12-bacteriology-growth-physiology-and-genetics#def-b3-bacteriology-resistance) crosses species in a hospital and why a bacterium’s “species” is a core genome surrounded by a shifting cloud; in eukaryotes transfer is rarer but real — the [T-DNA](https://one-course.com/books/biology/5/en/chapter/6-genetic-engineering-and-biotechnology#def-b3-genetic-engineering-plants) *[Agrobacterium](https://one-course.com/books/biology/5/en/chapter/6-genetic-engineering-and-biotechnology#def-b3-genetic-engineering-plants)* inserts into plants, bacterial genes in bdelloid rotifers, and the ancient wholesale transfers that were the mitochondrion and the chloroplast. Where transfer is common the history of life is a network, and the tree drawn from any one gene is the tree of that gene.

![The fates of a duplicated gene. Freed from constraint, the extra copy usually decays; occasionally it is caught by selection for a new function, or the two copies divide the old one between them.](https://one-course.com/images/onecourse/chapters/biology-5/b3-molecular-evolution/fig-3daa02e49445.svg)

*The fates of a duplicated gene. Freed from constraint, the extra copy usually decays; occasionally it is caught by selection for a new function, or the two copies divide the old one between them.*

**Evidence.** Ohno (1970) proposed duplication as the main source of new genes before a single family was sequenced; the globins confirmed it, their tree of $\alpha$, $\beta$, $\gamma$, $\delta$, myoglobin and the plant and bacterial globins matching duplication dates to the vertebrate radiation. Chen, DeVries and Cheng (1997) showed the icefish antifreeze gene shares its signal sequence and flanking sequences with trypsinogen and that the antifreeze repeats arose by expansion of a nine-nucleotide piece spanning an intron–exon junction: a new protein from a digestive enzyme’s spare parts, dated by the clock to the freezing of the Southern Ocean. Thornton and colleagues (2006) reconstructed the ancestral sequence of the steroid receptors, synthesised the 450-million-year-old protein, and found it responded to oestrogens only; two later substitutions, identified and tested, switched the duplicated copy’s preference to cortisol — a resurrected ancestor, and the path between it and its descendants walked in the laboratory. ∎

## 25.5 Gene trees and species trees

**Proposition 25.6 (Incomplete lineage sorting).**

The tree of a gene need not match the tree of the species that carry it. Consider three species with [species tree](#prop-b3-molecular-evolution-ils) $((A,B),C)$, the $A$–$B$ split preceded by an ancestral population of size $N$ that persisted for $T$ generations before the split from $C$. Two gene copies, one from $A$ and one from $B$, traced backward, *coalesce* in a common ancestor at a rate $1/2N$ per generation; the probability that they have not yet coalesced when they enter the population ancestral to all three is $\mathrm{e}^{-T/2N}$, and then the three lineages coalesce in random order, so two thirds of the time the [gene tree](#prop-b3-molecular-evolution-ils) groups $A$ or $B$ with $C$. The probability that the [gene tree](#prop-b3-molecular-evolution-ils) disagrees with the [species tree](#prop-b3-molecular-evolution-ils) is therefore

$$
P_{\text{discord}} = \tfrac{2}{3}\,\mathrm{e}^{-T/2N} .
$$

For humans, chimpanzees and gorillas, with $T$ of a few hundred thousand generations and $N$ near $50\,000$, about a third of the genome supports a tree in which humans are closer to gorillas or chimpanzees to gorillas than the [species tree](#prop-b3-molecular-evolution-ils) — observed. This *[incomplete lineage sorting](#prop-b3-molecular-evolution-ils)*, together with duplication and loss and horizontal transfer, is why *phylogenomics* infers the [species tree](#prop-b3-molecular-evolution-ils) from thousands of [gene trees](#prop-b3-molecular-evolution-ils) under a model of the coalescent, rather than concatenating them and drawing one; it also lets the ancestral population sizes themselves be estimated from the fraction of discordant trees, and ancient hybridisations (Neanderthal DNA in modern humans, the gene flow between bears, between butterflies) be read from trees that disagree in a direction drift alone would not produce.

**Proof.** Two lineages in a diploid population of $N$ pick parents from $2N$ copies each generation and share one with probability $1/2N$; over $T$ generations the chance of never doing so is $(1 - 1/2N)^{T} \approx
\mathrm{e}^{-T/2N}$. Three lineages entering the common ancestral population are exchangeable, so the first pair to coalesce is each of the three pairs with probability $1/3$; only the pair $(A,B)$ matches the [species tree](#prop-b3-molecular-evolution-ils), hence $2/3$ of these cases are discordant. Neither the coalescent’s rate nor the exchangeability depends on the gene, so the formula holds for every neutral locus independently, and the fraction of discordant trees across a genome estimates $\mathrm{e}^{-T/2N}$. ∎

![Gene trees inside a species tree. The grey tubes are populations; the lines are the ancestry of one gene. Lineages that fail to meet in the short ancestral population enter the deeper one together and pair off at random — so a third of the genome of three closely spaced species tells a different story from the species themselves.](https://one-course.com/images/onecourse/chapters/biology-5/b3-molecular-evolution/fig-d07705063bc9.svg)

*[Gene trees](#prop-b3-molecular-evolution-ils) inside a [species tree](#prop-b3-molecular-evolution-ils). The grey tubes are populations; the lines are the ancestry of one gene. Lineages that fail to meet in the short ancestral population enter the deeper one together and pair off at random — so a third of the genome of three closely spaced species tells a different story from the species themselves.*

**Example 25.7 (Reading a genome’s history).**

The five hundred cichlid species of Lake Victoria arose in the fifteen thousand years since the lake refilled — too fast for mutation to supply the differences between them. Their genomes show how: the variants that distinguish a bottom-feeder from an algae-scraper were present as standing variation in the river fish that colonised the lake, sorted and recombined, and augmented by hybridisation between lineages; the opsin genes were tuned to clear or turbid water by substitutions that the $\omega$ of their branches shows to have been selected; and the [gene trees](#prop-b3-molecular-evolution-ils) disagree with one another across the genome in exactly the pattern [incomplete lineage sorting](#prop-b3-molecular-evolution-ils) and gene flow predict. A radiation is not the origin of new genes but the redistribution of old alleles among new combinations under new selection — which the sequences record, allele by allele, if one knows how to read a tree that is really a forest.

![Cichlids of an African lake: hundreds of species in a few thousand years, their differences drawn largely from variation their common ancestors already carried.](https://one-course.com/images/onecourse/chapters/biology-5/b3-molecular-evolution/img-b1df65c42eed.jpg)

*Cichlids of an African lake: hundreds of species in a few thousand years, their differences drawn largely from variation their common ancestors already carried.*

**Remark 25.8 (Chance as the baseline).**

The [neutral theory](#thm-b3-molecular-evolution-neutral) is not a claim that selection is unimportant; it is a statement of what happens when selection is absent, precise enough to be tested against. Because the neutral rate is the mutation rate, divergence measures time; because neutral change fills the sites that selection leaves free, $\omega$ measures constraint; because neutral lineages coalesce at a known rate, the disagreement of [gene trees](#prop-b3-molecular-evolution-ils) measures the size of populations no one saw. Selection is then what remains when chance has been subtracted — a gene with $\omega$ above one, a sweep that has emptied variation from a region, a tree that disagrees in a direction drift cannot explain. The Antarctic icefish’s antifreeze, the ruminant’s stomach lysozyme, the primate’s third opsin are the exceptions found this way, and each is a story of a copied gene and a new job. The rest of the genome is a clock.

## 25.6 Exercises

**Exercise 25.1 ★.**

State the neutral [substitution rate](#thm-b3-molecular-evolution-neutral) and explain in words why it does not depend on population size.

**Solution of Exercise 25.1.**

$k = \mu$: the neutral [substitution rate](#thm-b3-molecular-evolution-neutral) per site per generation equals the mutation rate. Each generation $2N\mu$ new neutral mutations arise and each has probability $1/2N$ of fixing; a larger population makes more mutations and gives each a proportionally smaller chance, and the two effects cancel exactly.

**Exercise 25.2 ★.**

Define [orthologue](https://one-course.com/books/biology/5/en/chapter/4-genomics-and-sequencing#def-b3-genomics-comparative), [paralogue](https://one-course.com/books/biology/5/en/chapter/4-genomics-and-sequencing#def-b3-genomics-comparative) and [pseudogene](#def-b3-molecular-evolution-duplication), with one example each from the globin family.

**Solution of Exercise 25.2.**

[Orthologues](https://one-course.com/books/biology/5/en/chapter/4-genomics-and-sequencing#def-b3-genomics-comparative): genes in two species descended from one gene in their common ancestor — human $\beta$-globin and mouse $\beta$-globin. [Paralogues](https://one-course.com/books/biology/5/en/chapter/4-genomics-and-sequencing#def-b3-genomics-comparative): genes in one genome descended from a duplication — human $\alpha$- and $\beta$-globin, or $\beta$ and $\gamma$. [Pseudogene](#def-b3-molecular-evolution-duplication): a decayed copy that no longer encodes a protein — the $\psi\beta$ gene in the human $\beta$-globin cluster, complete with stop codons.

**Exercise 25.3 ★.**

What does $\omega = d_{N}/d_{S}$ measure? Interpret $\omega = 0.02$, $\omega = 1.0$ and $\omega = 2.5$, and name a gene of each kind.

**Solution of Exercise 25.3.**

$\omega$ compares the rate of amino-acid-changing substitution with the rate of silent substitution, which is the neutral baseline. $0.02$: strong [purifying selection](#def-b3-molecular-evolution-dnds), $98\,\%$ of amino-acid changes removed — actin, [histone](https://one-course.com/books/biology/5/en/chapter/1-chromatin-and-epigenetics#def-b3-chromatin-epigenetics-nucleosome) H3. $1.0$: no constraint — a [pseudogene](#def-b3-molecular-evolution-duplication). $2.5$: [positive selection](#def-b3-molecular-evolution-dnds), amino-acid changes favoured — the peptide-binding groove of MHC, the envelope of HIV, ruminant lysozyme.

**Exercise 25.4 ★.**

Why did Kimura conclude that most substitutions are neutral? Restate the argument with the numbers in the chapter’s opening.

**Solution of Exercise 25.4.**

Proteins compared across mammals change at about one substitution per site per billion years; over a genome of a billion coding sites and a generation of a few years that is of the order of one substitution fixed per generation, or one every two years. Haldane’s cost of selection allows about one selected substitution per three hundred generations. The rates differ by a factor of a hundred or more, so nearly all fixed changes must be neutral.

**Exercise 25.5 ★★.**

$N = 5000$. Compute the [fixation probability](#thm-b3-molecular-evolution-neutral) of a new mutation with $s
= 0$, $s = 0.001$, $s = 0.01$ and $s = -0.001$, using Kimura’s formula, and comment on which are effectively neutral.

**Solution of Exercise 25.5.**

Neutral: $1/2N = 10^{-4}$. $s = 0.001$, $4Ns = 20$: $(1 - \mathrm{e}^{-0.002})/
(1 - \mathrm{e}^{-20}) = 0.002$. $s = 0.01$: $0.0198$. $s = -0.001$: $(1 -
\mathrm{e}^{0.002})/(1 - \mathrm{e}^{20}) = 0.002\,\mathrm{e}^{-20} =
4\times 10^{-12}$. None is effectively neutral: that would need $|4Ns|
\lesssim 1$, i.e. $|s| \lesssim 5\times 10^{-5}$; a selection coefficient of a tenth of a per cent is already decisive in a population of five thousand.

**Exercise 25.6 ★★.**

Two species differ at $12\,\%$ of the sites of a gene whose neutral rate is $2 \times 10^{-9}\,$ per site per year. Estimate their divergence time with and without the Jukes–Cantor correction. At what raw difference does the correction reach a factor of two?

**Solution of Exercise 25.6.**

Uncorrected: $T = 0.12/(2\times 2\times 10^{-9}) = 30\,\mathrm{Myr}$. Jukes–Cantor: $d = -\tfrac{3}{4}\ln(1 - 0.16) = 0.131$, $T =
33\,\mathrm{Myr}$. The correction reaches a factor of two when $-\tfrac{3}{4}
\ln(1 - 4p/3) = 2p$, at $p \approx 0.6$: by then the raw difference has lost most of its information.

**Exercise 25.7 ★★.**

A gene of 300 codons has 700 nonsynonymous and 200 synonymous sites. Between two species it shows 7 nonsynonymous and 10 synonymous differences. Compute $p_{N}$, $p_{S}$ and $\omega$ (uncorrected). Which selection regime? What if the counts were 35 and 10?

**Solution of Exercise 25.7.**

$p_{N} = 7/700 = 0.010$, $p_{S} = 10/200 = 0.050$, $\omega = 0.2$: [purifying selection](#def-b3-molecular-evolution-dnds) removing four fifths of amino-acid changes. With $35$ and $10$: $p_{N} = 0.05$, $\omega = 1$: no detectable constraint — a [pseudogene](#def-b3-molecular-evolution-duplication), or a gene in which [positive selection](#def-b3-molecular-evolution-dnds) at some sites balances [purifying selection](#def-b3-molecular-evolution-dnds) at others.

**Exercise 25.8 ★★.**

Human and chimpanzee neutral sequences differ at $1.2\,\%$; the split was $6.5\,\mathrm{Myr}$ ago, generation time $25\,\mathrm{yr}$. Estimate the mutation rate per site per generation, and the number of new mutations in a diploid genome of $6\times 10^{9}$ bases each generation.

**Solution of Exercise 25.8.**

$k = d/2T = 0.012/(1.3\times 10^{7}) = 9.2\times 10^{-10}$ per site per year; $\mu = 25k = 2.3\times 10^{-8}$ per site per generation; $6\times
10^{9}\times 2.3\times 10^{-8} \approx 140$ new mutations per diploid genome per generation. Direct sequencing of parents and children finds about seventy: the divergence estimate is inflated by polymorphism already present in the ancestral population, which adds a few hundred thousand generations to the apparent split.

**Exercise 25.9 ★★.**

With $N = 50\,000$ and $T = 100\,000$ generations between the two splits, compute the fraction of [gene trees](#prop-b3-molecular-evolution-ils) discordant with the [species tree](#prop-b3-molecular-evolution-ils) for human, chimpanzee and gorilla. What $T$ would make it $10\,\%$?

**Solution of Exercise 25.9.**

$\mathrm{e}^{-T/2N} = \mathrm{e}^{-1} = 0.37$; discordant fraction $\tfrac{2}{3}
\times 0.37 = 0.25$. For $10\,\%$: $\mathrm{e}^{-T/2N} = 0.15$, $T =
2N\ln(1/0.15) = 190\,000$ generations.

**Exercise 25.10 ★★★.**

Show that Kimura’s formula reduces to $1/2N$ as $s \to 0$ and to $2s$ for $4Ns \gg 1$, and that for $s < 0$ with $4N|s| \gg 1$ it behaves as $2|s|\,\mathrm{e}^{-4N|s|}$. Hence estimate how strongly deleterious a mutation must be, in a population of $10^{4}$, to have its [fixation probability](#thm-b3-molecular-evolution-neutral) reduced a hundredfold below neutral.

**Solution of Exercise 25.10.**

As $s \to 0$: $1 - \mathrm{e}^{-2s} \approx 2s$ and $1 - \mathrm{e}^{-4Ns}
\approx 4Ns$, ratio $1/2N$. For $4Ns \gg 1$ the denominator is $1$ and the numerator $\approx 2s$. For $s < 0$ with $4N|s| \gg 1$: numerator $1 -
\mathrm{e}^{2|s|} \approx -2|s|$, denominator $1 - \mathrm{e}^{4N|s|} \approx
-\mathrm{e}^{4N|s|}$, so $P \approx 2|s|\,\mathrm{e}^{-4N|s|}$. A hundredfold below neutral in $N = 10^{4}$: $2|s|\,\mathrm{e}^{-4\times 10^{4}|s|} =
5\times 10^{-7}$; $|s| = 1.6\times 10^{-4}$ gives $3.2\times 10^{-4}\times
\mathrm{e}^{-6.4} = 5.3\times 10^{-7}$. So $|s| \approx 1.6\times 10^{-4}$, $4N|s| \approx 6.5$: a disadvantage of a hundredth of a per cent is enough, in ten thousand individuals, to make fixation a hundred times rarer than chance.

**Exercise 25.11 ★★★.**

A relative-rate test: an outgroup $O$ is at corrected distance $0.30$ from mouse and $0.24$ from human, and mouse and human are $0.20$ apart. Compute the mouse and human branch lengths since their split (assume the split point is equidistant from $O$ on the two paths), and their ratio. With a mouse generation of $0.5\,\mathrm{yr}$ and a human one of $25\,\mathrm{yr}$, what would a strict per-generation clock predict for the ratio, and what does the observation say about mutation rates per generation?

**Solution of Exercise 25.11.**

With the split point at distance $a$ from $O$ on both paths, $a + m =
0.30$, $a + h = 0.24$, $m + h = 0.20$: $m - h = 0.06$, so $m = 0.13$, $h
= 0.07$, ratio $1.9$. A strict per-generation clock would give the ratio of generation counts, $25/0.5 = 50$. The observed $1.9$ means the mouse has accumulated only twice the human change per year despite fifty times the generations: the mutation rate per generation must be some $25$ times higher in humans — more germ-line cell divisions per generation — so that per year the two lineages differ by only a factor of two.

**Exercise 25.12 ★★★.**

A duplicated gene is released from constraint ($\omega = 1$) at neutral rate $k = 2 \times 10^{-9}\,$ per site per year. In a 1000-base coding sequence, roughly how many years until the first stop codon or frameshift is expected (assume about $4\,\%$ of random point mutations in a coding sequence create a stop codon, and ignore indels)? Compare with the time selection has to find a new function, and explain why most duplicates die.

**Solution of Exercise 25.12.**

Stop-creating substitutions arrive at $1000\times 2\times 10^{-9}\times
0.04 = 8\times 10^{-8}$ per year: the first is expected after about $12\,\mathrm{Myr}$ (frameshifts roughly halve this). In that window a beneficial mutation, one in $10^{3}$ to $10^{5}$ of all mutations, is far less likely to arrive than a stop codon, and even when it arrives it fixes with probability only $2s$. So the copy is usually dead before selection can find it a use — the observed half-life of duplicates in vertebrates is a few million years.

## 25.7 Problem: The Rate of Change

**Problem 25.1.**

Weekend problem — a genome’s history in numbers: the mutation rate from the human–chimpanzee divergence, Kimura’s substitution load, the fixation odds of selected mutations, a gene’s $\omega$ from codon counts, the clock read for a duplication, and the discordance of gene trees among the great apes, ending on the mutation rate, the $\omega$ of the gene and the discordant fraction

Data: human–chimpanzee neutral divergence $d = 0.012$; split $6.5\,\mathrm{Myr}$ ago; generation $25\,\mathrm{yr}$; diploid genome $6\times
10^{9}$ bases, $1.5\,\%$ coding. Ancestral population $N =
50\,000$. Gene: 400 codons, 920 nonsynonymous and 280 synonymous sites; human–mouse differences 46 nonsynonymous and 84 synonymous. Human–mouse split $90\,\mathrm{Myr}$. Globin: $\alpha$ and $\beta$ differ at $55\,\%$ of amino-acid sites (raw); haemoglobin rate $1 \times 10^{-9}$ per site per year. Apes: $T$ between the gorilla and chimpanzee splits $80\,000$ generations. Haldane’s limit: one selected substitution per 300 generations.

**Part I — The rate.**

1. From $d = 2kT$ , compute $k$ per site per year, then $\mu$ per site per generation.
2. New mutations per diploid genome per generation, and how many of them fall in coding sequence.
3. Substitutions fixed per generation in the whole genome at the neutral rate (the population fixes $\mu$ per site per generation). Compare with Haldane’s limit and draw Kimura’s conclusion.
4. In a population of $N = 50\,000$ , how many new neutral mutations arise per site per generation in the whole population, and what fraction of them will ever fix?
5. How long does a neutral mutation destined to fix take, in generations and years? Compare with the time since the human–chimpanzee split.
6. Explain why, despite question 5, the divergence between two species measures the time since their split and not the time since their alleles fixed.

**Part II — Selection’s odds.**

7. [Fixation probability](#thm-b3-molecular-evolution-neutral) of a neutral mutation in $N = 50\,000$ , and of one with $s = 0.001$ , $s = 0.01$ , $s = 0.1$ .
8. For which $s$ is $4Ns = 1$ ? Interpret the boundary.
9. A deleterious mutation with $s = -0.001$ : [fixation probability](#thm-b3-molecular-evolution-neutral) relative to neutral. And $s = -10^{-5}$ ?
10. If a fraction $10^{-5}$ of new mutations in a gene are beneficial with $s = 0.01$ and the rest neutral, what fraction of the substitutions in that gene are adaptive? Comment on how rarely selection shows in the sequence even when it acts.
11. Compute $p_{N}$ and $p_{S}$ for the human–mouse gene, correct each with Jukes–Cantor, and give $\omega$ .
12. Interpret $\omega$ ; and estimate the neutral rate the gene’s $d_{S}$ implies (per site per year) using the split time. Is it consistent with question 1?

**Part III — Clocks and copies.**

13. Correct the $\alpha$ – $\beta$ globin difference for multiple hits (use $d = -\ln(1 - p)$ , the protein version with many states), and date the duplication with the haemoglobin rate.
14. The date falls near the origin of jawed vertebrates ( $450\text{ to }500\,\mathrm{Myr}$ ). What does having separate $\alpha$ and $\beta$ chains allow that a single chain does not?
15. Fibrinopeptides change at $8\times 10^{-9}$ per site per year. Two species split $40\,\mathrm{Myr}$ ago: expected raw difference? Why is this protein useless for dating splits of $500\,\mathrm{Myr}$ ?
16. [Histone](https://one-course.com/books/biology/5/en/chapter/1-chromatin-and-epigenetics#def-b3-chromatin-epigenetics-nucleosome) H4 changes at $10^{-11}$ : how many substitutions in 100 sites over $1000\,\mathrm{Myr}$ ? Why is it useless for dating recent splits?
17. A duplicate gene drifts with $\omega = 1$ from the moment of duplication. If $4\,\%$ of substitutions in coding sequence create stops, and the gene has 1200 sites at rate $2\times  10^{-9}$ per site per year, estimate the expected time to the first stop codon.
18. Why does [whole-genome duplication](https://one-course.com/books/biology/5/en/chapter/4-genomics-and-sequencing#def-b3-genomics-comparative) give duplicates a better chance of survival than a single tandem duplication?

**Part IV — Trees within trees.**

19. Probability that a human and a chimpanzee lineage fail to coalesce in the $80\,000$ generations before the gorilla split.
20. Fraction of [gene trees](#prop-b3-molecular-evolution-ils) discordant with the [species tree](#prop-b3-molecular-evolution-ils) . Observed: about $30\,\%$ . Consistent?
21. Of the discordant trees, what fraction group human with gorilla, and what fraction chimpanzee with gorilla? What would a strong departure from equality indicate?
22. If the ancestral population had been $N = 10\,000$ , what discordant fraction would you expect? What does the observed fraction therefore say about our ancestors’ numbers?
23. Why does concatenating a thousand genes into one [alignment](https://one-course.com/books/biology/5/en/chapter/5-bioinformatics-and-sequence-analysis#def-b3-bioinformatics-alignment) give a confident but possibly wrong tree, and what does a coalescent method do instead?
24. Neanderthal DNA in non-African humans is about $2\,\%$ and the trees carrying it group some Europeans with Neanderthals more often than Africans. Why is this not [incomplete lineage sorting](#prop-b3-molecular-evolution-ils) ?
25. Summarise: $\mu$ (question 1), $\omega$ of the gene (question 11), and the discordant fraction (question 20).

**Solution of Problem 25.1.**

**1.** $k = 0.012/(2\times 6.5\times 10^{6}) = 9.2\times 10^{-10}$ per site per year; $\mu = 25k = 2.3\times 10^{-8}$ per site per generation. **2.** $6\times 10^{9}\times 2.3\times 10^{-8} \approx 140$ new mutations per diploid genome; $1.5\,\%$ of them, about $2$, in coding sequence. **3.** Per haploid genome, $3\times 10^{9}\times 2.3\times 10^{-8}
\approx 70$ substitutions fixed per generation, against Haldane’s $1/300$: twenty thousand times more than selection could drive, so the overwhelming majority are neutral. **4.** $2N\mu = 10^{5}\times 2.3\times 10^{-8} = 2.3\times 10^{-3}$ new mutations per site per generation in the population; $1/2N =
10^{-5}$ of them will fix. **5.** $4N = 200\,000$ generations, $5\,\mathrm{Myr}$ — nearly as long as the time since the split. **6.** Divergence counts the mutations that arose along the two separate lineages since their common ancestor; each fixes within its own lineage, and the long-run rate is $\mu$ whatever the time each takes, since mutations arise steadily and the lag only delays, without changing, the flow. (The only correction is for polymorphism present at the split, which makes the common ancestor of two alleles older than the species split.) **7.** Neutral $10^{-5}$; $s = 0.001$: $0.002$; $s = 0.01$: $0.0198$; $s = 0.1$: $1 - \mathrm{e}^{-0.2} = 0.18$. **8.** $s = 1/4N = 5\times 10^{-6}$: below this drift swamps selection and the mutation behaves as neutral; above it selection governs the odds. **9.** $s = -0.001$: $P = 0.002\,\mathrm{e}^{-200}$, zero for all purposes — $10^{-85}$ of neutral. $s = -10^{-5}$, $4N|s| = 2$: $P =
2\times 10^{-5}/(\mathrm{e}^{2} - 1) = 3.1\times 10^{-6}$, $31\,\%$ of neutral: mildly deleterious mutations do fix. **10.** Beneficial substitutions per site per generation: $2N\times
10^{-5}\mu\times 2s = 10^{5}\times 10^{-5}\times 0.02\,\mu = 0.02\mu$; neutral: $\approx\mu$. Adaptive fraction $0.02/1.02 \approx 2\,\%$. Though each beneficial mutation is two thousand times likelier to fix than a neutral one, they are so rare that only one substitution in fifty is adaptive: selection acts and barely shows. **11.** $p_{N} = 46/920 = 0.050$, $p_{S} = 84/280 = 0.30$. Corrected: $d_{N} = -\tfrac{3}{4}\ln(1 - 0.0667) = 0.052$, $d_{S} = -\tfrac{3}{4}
\ln(1 - 0.40) = 0.38$. $\omega = 0.14$. **12.** [Purifying selection](#def-b3-molecular-evolution-dnds) removing about $86\,\%$ of amino-acid changes — a typical gene. $k = d_{S}/2T = 0.38/(1.8\times
10^{8}) = 2.1\times 10^{-9}$ per site per year: twice the human–chimp value, as expected for a comparison that includes the fast-ticking rodent lineage; the same order of magnitude. **13.** $d = -\ln(1 - 0.55) = 0.80$ per site; $T = 0.80/(2\times
10^{-9}) = 400\,\mathrm{Myr}$. **14.** Two different chains make the $\alpha_{2}\beta_{2}$ tetramer, whose cooperative oxygen binding, [Bohr effect](https://one-course.com/books/biology/5/en/chapter/7-structural-biology-of-proteins#prop-b3-structural-biology-mechanism) and allosteric control by 2,3-bisphosphoglycerate depend on the interface between unlike subunits; and the $\beta$ cluster could then diversify into embryonic, fetal and adult chains with different affinities. **15.** $2kT = 2\times 8\times 10^{-9}\times 4\times 10^{7} = 0.64$ substitutions per site; with saturation among twenty amino acids the raw difference is about $0.95(1 - \mathrm{e}^{-0.64/0.95}) \approx 0.47$. At $500\,\mathrm{Myr}$, $d = 8$: every site has changed many times and the raw difference sits at its ceiling of $95\,\%$, carrying no information about time. **16.** $10^{-11}\times 10^{9}\times 100\times 2 = 2$ substitutions between two lineages in 100 sites over a billion years. For a $10\,\mathrm{Myr}$ split the expectation is $0.02$: almost always zero differences, no resolution. **17.** $1200\times 2\times 10^{-9}\times 0.04 = 9.6\times 10^{-8}$ per year: the first stop is expected after about $10\,\mathrm{Myr}$. **18.** Doubling the whole genome doubles every gene at once, so the stoichiometry of interacting proteins is preserved and no dosage imbalance selects against the copies; a lone tandem duplicate changes one gene’s dose and is often deleterious, or immediately redundant and free to decay. **19.** $\mathrm{e}^{-80\,000/100\,000} = \mathrm{e}^{-0.8} = 0.45$. **20.** $\tfrac{2}{3}\times 0.45 = 0.30$: consistent with the observed $30\,\%$. **21.** Half each, $15\,\%$ of all trees grouping human with gorilla and $15\,\%$ chimpanzee with gorilla. A clear excess of one would indicate gene flow between those two species after their split, which drift cannot produce. **22.** $N = 10\,000$: $\mathrm{e}^{-4} = 0.018$, discordance $1.2\,\%$. The observed $30\,\%$ therefore requires an ancestral population of some fifty thousand — several times the effective size of humans today. **23.** Concatenation assumes every gene has the [species tree](#prop-b3-molecular-evolution-ils); with [incomplete lineage sorting](#prop-b3-molecular-evolution-ils) the genes have many trees, and for short internal branches the commonest [gene tree](#prop-b3-molecular-evolution-ils) can differ from the [species tree](#prop-b3-molecular-evolution-ils), so more data make the wrong answer more confident. A coalescent method computes, for each candidate [species tree](#prop-b3-molecular-evolution-ils), the expected distribution of [gene trees](#prop-b3-molecular-evolution-ils), and picks the [species tree](#prop-b3-molecular-evolution-ils) that best explains the frequencies observed. **24.** [Incomplete lineage sorting](#prop-b3-molecular-evolution-ils) happened in the common ancestral population of all modern humans, so it would make every human population equally related to Neanderthals. An excess in non-Africans means Neanderthal alleles entered after the ancestors of non-Africans left Africa: admixture, some fifty thousand years ago. **25.** $\mu \approx 2.3\times 10^{-8}$ per site per generation; $\omega \approx 0.14$; about $30\,\%$ of [gene trees](#prop-b3-molecular-evolution-ils) discordant with the [species tree](#prop-b3-molecular-evolution-ils).
