Biology · Book 5 · Bachelor Year 3

University Biology — Year 3

University Biology — Year 3 · Bachelor Year 3

25Molecular Evolution and Phylogenomics

In 1968 Motoo Kimura did a sum that unsettled a field. Comparing the haemoglobins, cytochromes and other proteins of mammals whose common ancestors were dated by fossils, he found that amino acids had been replaced at a rate of roughly one substitution per site per billion years — which, over a mammalian genome and a generation, came to a new substitution fixed in the population every two years or so. Haldane had shown a decade earlier that natural selection can fix at most about one substitution per three hundred generations before the cost in lost offspring becomes unpayable. The arithmetic left one way out: most of the changes that accumulate in DNA are not driven by selection at all, but are neutral, fixed by chance in finite populations, at a rate that — as this chapter’s first theorem shows — is simply the mutation rate. The neutral theory became the null hypothesis of molecular evolution, the baseline against which selection is detected; the molecular clock became a way of dating what fossils could not; and the genomes sequenced since have given the tools to see, gene by gene, where selection has acted, how new genes are born from old, and why the tree of a gene is not always the tree of its species.

25.1 Drift, mutation and the neutral rate

Theorem 25.1 (The neutral substitution rate)

In a diploid population of NN individuals, a new mutation with no effect on fitness has probability 1/2N1/2N of eventually replacing all other alleles (fixation) and 11/2N1 - 1/2N of being lost. If each of the 2N2N gene copies mutates at rate μ\mu per generation, 2Nμ2N\mu new neutral mutations arise per generation and, in the long run, the rate at which substitutions accumulate is

k=2Nμ×12N=μ:k = 2N\mu\times\frac{1}{2N} = \mu :

the neutral substitution rate equals the mutation rate, independent of population size. A mutation destined to fix takes on average 4N4N generations to do so, so large populations hold more variation in transit but fix it no faster. For a mutation with selective advantage ss, Kimura’s formula gives the fixation probability

P(s)=1e2s1e4Ns2s(4Ns1),P(s) = \frac{1 - \mathrm{e}^{-2s}}{1 - \mathrm{e}^{-4Ns}} \approx 2s \quad (4Ns \gg 1),

so an advantage of 1%1\,\% is fixed with probability about 2%2\,\% — forty times a neutral mutation’s chance in a population of two thousand, but still lost forty-nine times in fifty; and a disadvantage s<0s < 0 with 4Ns14N|s| \gg 1 is fixed essentially never, while one with 4Ns14N|s| \ll 1 behaves as neutral: selection sees only what drift does not swamp, and the boundary is s1/4N|s| \sim 1/4Neffectively neutral.

Proof. Neutrality: some copy present today will be the ancestor of the whole population in the distant future, and by symmetry each of the 2N2N copies is equally likely to be it; the new mutant is one copy, hence 1/2N1/2N. Rate: substitutions per generation == (mutations arising per generation) ×\times (probability each fixes) =2Nμ/2N= 2N\mu/2N. Kimura’s formula follows from the diffusion approximation to the change of allele frequency under drift and selection, admitted here; its limits are checked directly: as s0s \to 0, (2s)/(4Ns)=1/2N(2s)/(4Ns) = 1/2N, the neutral value; for 4Ns14Ns \gg 1 the denominator is 11 and the numerator 2s\approx 2s. Fixation time: the frequency of a neutral allele performs a random walk with variance p(1p)/2Np(1-p)/2N per generation, and the expected time to reach 11 from 1/2N1/2N, conditional on doing so, is 4N4N generations (Kimura and Ohta, 1969), admitted.

Kimura’s fixation probability against the scaled selection coefficient, for N = 1000. Near zero the curve passes through the neutral value; to the right it rises toward 2s; to the left it falls off a cliff. Selection acts only on what drift cannot hide.
Kimura’s fixation probability against the scaled selection coefficient, for N=1000N = 1000. Near zero the curve passes through the neutral value; to the right it rises toward 2s2s; to the left it falls off a cliff. Selection acts only on what drift cannot hide.

Evidence. Kimura (1968) and King and Jukes (1969) made the argument from rates: the observed rate of amino-acid substitution across mammalian proteins implied more substitutions per generation than Haldane’s cost of selection allowed, so most had to be neutral. The theory’s predictions then held: rates are highest where function is least constrained — synonymous sites, introns, pseudogenes, the third codon position — and lowest in histones and ubiquitin, which change one residue in a hundred million years; the level of variation within species tracks mutation and population size; and the rate of substitution per year is roughly constant across lineages with very different population sizes, as the formula k=μk = \mu requires and a selectionist account does not. The controversy that followed did not overturn it: it fixed the neutral rate as the null against which selection is measured.

Left: Motoo Kimura (1924–1994), who showed that most molecular change is fixed by chance and at the mutation rate (photograph, 1986, CC BY 4.0). Right: an Antarctic icefish, whose blood is kept from freezing by an antifreeze glycoprotein assembled, some ten million years ago, from a duplicated digestive-enzyme gene and a run of repeats. Left: Motoo Kimura (1924–1994), who showed that most molecular change is fixed by chance and at the mutation rate (photograph, 1986, CC BY 4.0). Right: an Antarctic icefish, whose blood is kept from freezing by an antifreeze glycoprotein assembled, some ten million years ago, from a duplicated digestive-enzyme gene and a run of repeats.
Left: Motoo Kimura (1924–1994), who showed that most molecular change is fixed by chance and at the mutation rate (photograph, 1986, CC BY 4.0). Right: an Antarctic icefish, whose blood is kept from freezing by an antifreeze glycoprotein assembled, some ten million years ago, from a duplicated digestive-enzyme gene and a run of repeats.

25.2 The molecular clock

Proposition 25.2 (The clock and its irregularities)

If substitutions accumulate at a constant rate kk per site per year in two lineages that split TT years ago, the fraction of sites at which they differ is, for small values, d2kTd \approx 2kT: the divergence of two sequences is a molecular clock (Zuckerkandl and Pauling, 1965), which can be calibrated by one fossil-dated split and then read for splits with no fossils. The clock is stochastic — a Poisson process, so dd has variance about equal to its mean over a fixed number of sites — and it is irregular in three known ways. Each protein has its own rate, set by the fraction of its sites that function permits to change: fibrinopeptides at 88, haemoglobin 11, cytochrome cc 0.30.3, histone H4 0.010.01 substitutions per site per billion years. Rates differ between lineages: since the neutral rate is μ\mu per generation, animals with short generations (rodents) accumulate more change per year than those with long (primates, whales) — the generation-time effect, partly offset because per-generation mutation rates rise with generation length. And dd saturates: once many sites have changed, further changes hit sites already changed, so the raw difference must be corrected for multiple hits, as the alignment chapter’s distances do. Modern relaxed clocks let the rate vary between branches within a statistical model and are calibrated by many fossils at once; they date the split of humans and chimpanzees at six to seven million years, of placental mammal orders near the end of the Cretaceous, and of animals and fungi in the Precambrian, with uncertainties of ten to twenty per cent that come mostly from the fossils.

Proof. Two lineages each accumulate kTkT substitutions per site, so the difference between them is 2kT2kT where sites have not been hit twice; a fossil-dated pair gives k=d/2Tk = d/2T. The Poisson variance and the per-protein rates are Zuckerkandl and Pauling’s, Dickerson’s (1971) and Kimura’s fits to vertebrate protein sets; the saturation correction is the Jukes–Cantor formula of the alignment chapter, dtrue=34ln(143p)d_{\text{true}} = -\tfrac{3}{4}\ln(1 - \tfrac{4}{3}p). Rate variation between lineages was shown by relative-rate tests: with an outgroup OO and two species AA, BB, the distances dAOdBOd_{AO} - d_{BO} should be zero under a strict clock, and are not for rodents against primates.

The molecular clock, protein by protein. Each runs at its own rate, set by how much of the molecule function leaves free to change: a fibrinopeptide, cut away and discarded when blood clots, changes hundreds of times faster than the histone that packs DNA.
The molecular clock, protein by protein. Each runs at its own rate, set by how much of the molecule function leaves free to change: a fibrinopeptide, cut away and discarded when blood clots, changes hundreds of times faster than the histone that packs DNA.

25.3 Reading selection in sequences

Definition 25.3 (dN/dS)

In a protein-coding gene a substitution is synonymous if it leaves the amino acid unchanged and nonsynonymous if it changes it. Because the genetic code lets most third positions and some first positions change silently, about a quarter of the possible mutations in a typical gene are synonymous. Let dSd_{S} be the number of synonymous substitutions per synonymous site and dNd_{N} the nonsynonymous substitutions per nonsynonymous site, each corrected for multiple hits. Synonymous changes are close to neutral, so dSd_{S} estimates 2μT2\mu T — the clock; and the ratio

ω=dNdS\omega = \frac{d_{N}}{d_{S}}

measures what selection has done to the protein: ω=1\omega = 1 if amino-acid changes are as free as silent ones (no constraint, as in a pseudogene); ω<1\omega < 1 under purifying selection, most changes being removed — the typical gene has ω0.1\omega \approx 0.1 to 0.20.2, histones 0.0010.001; ω>1\omega > 1 under positive selection, amino-acid changes being favoured faster than drift alone allows — the antigen-binding groove of MHC molecules, the surface proteins of viruses in an arms race with antibodies, the sperm and egg proteins of fertilisation, and the lysozyme of ruminants, which was recruited to digest bacteria in the stomach. Computed along one branch of a tree or site by site, ω\omega maps where and when a protein was changed by selection rather than by chance.

Orders of magnitude of . Almost every gene sits well below one, its amino acids guarded by purifying selection; a pseudogene drifts at one; the few above one are the proteins in arms races or newly recruited to a job.
Orders of magnitude of ω\omega. Almost every gene sits well below one, its amino acids guarded by purifying selection; a pseudogene drifts at one; the few above one are the proteins in arms races or newly recruited to a job.

Method 25.4 (Estimating dN/dSd_{N}/d_{S})

(1) Align the two coding sequences codon by codon (align the proteins, then map back to the DNA, so that no gap breaks the reading frame). (2) For each codon count the synonymous and nonsynonymous sites: each of its three positions is scored by the fraction of possible changes that are silent — a fourfold-degenerate third position counts as one synonymous site, a position where every change alters the amino acid as one nonsynonymous site, a twofold position as a third of one and two thirds of the other. Sum over codons: SS synonymous and NN nonsynonymous sites, with S+N=3×S + N = 3 \times codons. (3) Count the synonymous and nonsynonymous differences between the sequences, resolving codons that differ at two positions by averaging over the possible paths. (4) pS=p_{S} = differences/S/S, pN=p_{N} = differences/N/N; correct each for multiple hits with Jukes–Cantor. (5) ω=dN/dS\omega = d_{N}/d_{S}; test ω1\omega \ne 1 against the sampling variance, which is large when dSd_{S} is small — two closely related sequences give an unreliable ratio, and two very distant ones have a saturated dSd_{S}. (6) For where and when: fit ω\omega per branch of a tree, or per site with a mixture of site classes, and test whether a class with ω>1\omega > 1 improves the fit.

25.4 New genes from old

Definition 25.5 (Duplication, families and transfer)

Most new genes are copies of old ones. A gene duplication — by unequal crossing-over, retrotransposition, or the doubling of a whole genome, as happened twice at the origin of vertebrates and again in the ancestor of teleost fish and in many plant lineages — leaves two paralogues where there was one; genes in different species descended from a single ancestral gene are orthologues. The copy’s usual fate is decay: freed from constraint (ω1\omega \to 1) it accumulates a stop codon or a frameshift and becomes a pseudogene, of which the human genome carries some twenty thousand. Sometimes both survive: by neofunctionalisation, one copy acquiring a new function while the other keeps the old (the antifreeze glycoprotein of icefish, from a trypsinogen gene; the red and green opsins of primates, from a duplication forty million years ago that gave the trichromatic vision the Year 1 volume described; the lens crystallins, from metabolic enzymes); or by subfunctionalisation, each copy keeping part of the ancestor’s expression or activity so that both are needed. Repeated duplication builds gene families — the globins, the Hox clusters of the development chapter, the thousand olfactory receptor genes of a mouse, a third of them pseudogenes in humans. The other source of new genes is horizontal transfer: bacteria acquire genes from unrelated bacteria by plasmids, phages and free DNA, which is how antibiotic resistance crosses species in a hospital and why a bacterium’s “species” is a core genome surrounded by a shifting cloud; in eukaryotes transfer is rarer but real — the T-DNA Agrobacterium inserts into plants, bacterial genes in bdelloid rotifers, and the ancient wholesale transfers that were the mitochondrion and the chloroplast. Where transfer is common the history of life is a network, and the tree drawn from any one gene is the tree of that gene.

The fates of a duplicated gene. Freed from constraint, the extra copy usually decays; occasionally it is caught by selection for a new function, or the two copies divide the old one between them.
The fates of a duplicated gene. Freed from constraint, the extra copy usually decays; occasionally it is caught by selection for a new function, or the two copies divide the old one between them.

Evidence. Ohno (1970) proposed duplication as the main source of new genes before a single family was sequenced; the globins confirmed it, their tree of α\alpha, β\beta, γ\gamma, δ\delta, myoglobin and the plant and bacterial globins matching duplication dates to the vertebrate radiation. Chen, DeVries and Cheng (1997) showed the icefish antifreeze gene shares its signal sequence and flanking sequences with trypsinogen and that the antifreeze repeats arose by expansion of a nine-nucleotide piece spanning an intron–exon junction: a new protein from a digestive enzyme’s spare parts, dated by the clock to the freezing of the Southern Ocean. Thornton and colleagues (2006) reconstructed the ancestral sequence of the steroid receptors, synthesised the 450-million-year-old protein, and found it responded to oestrogens only; two later substitutions, identified and tested, switched the duplicated copy’s preference to cortisol — a resurrected ancestor, and the path between it and its descendants walked in the laboratory.

25.5 Gene trees and species trees

Proposition 25.6 (Incomplete lineage sorting)

The tree of a gene need not match the tree of the species that carry it. Consider three species with species tree ((A,B),C)((A,B),C), the AABB split preceded by an ancestral population of size NN that persisted for TT generations before the split from CC. Two gene copies, one from AA and one from BB, traced backward, coalesce in a common ancestor at a rate 1/2N1/2N per generation; the probability that they have not yet coalesced when they enter the population ancestral to all three is eT/2N\mathrm{e}^{-T/2N}, and then the three lineages coalesce in random order, so two thirds of the time the gene tree groups AA or BB with CC. The probability that the gene tree disagrees with the species tree is therefore

Pdiscord=23eT/2N.P_{\text{discord}} = \tfrac{2}{3}\,\mathrm{e}^{-T/2N} .

For humans, chimpanzees and gorillas, with TT of a few hundred thousand generations and NN near 5000050\,000, about a third of the genome supports a tree in which humans are closer to gorillas or chimpanzees to gorillas than the species tree — observed. This incomplete lineage sorting, together with duplication and loss and horizontal transfer, is why phylogenomics infers the species tree from thousands of gene trees under a model of the coalescent, rather than concatenating them and drawing one; it also lets the ancestral population sizes themselves be estimated from the fraction of discordant trees, and ancient hybridisations (Neanderthal DNA in modern humans, the gene flow between bears, between butterflies) be read from trees that disagree in a direction drift alone would not produce.

Proof. Two lineages in a diploid population of NN pick parents from 2N2N copies each generation and share one with probability 1/2N1/2N; over TT generations the chance of never doing so is (11/2N)TeT/2N(1 - 1/2N)^{T} \approx \mathrm{e}^{-T/2N}. Three lineages entering the common ancestral population are exchangeable, so the first pair to coalesce is each of the three pairs with probability 1/31/3; only the pair (A,B)(A,B) matches the species tree, hence 2/32/3 of these cases are discordant. Neither the coalescent’s rate nor the exchangeability depends on the gene, so the formula holds for every neutral locus independently, and the fraction of discordant trees across a genome estimates eT/2N\mathrm{e}^{-T/2N}.

Gene trees inside a species tree. The grey tubes are populations; the lines are the ancestry of one gene. Lineages that fail to meet in the short ancestral population enter the deeper one together and pair off at random — so a third of the genome of three closely spaced species tells a different story from the species themselves.
Gene trees inside a species tree. The grey tubes are populations; the lines are the ancestry of one gene. Lineages that fail to meet in the short ancestral population enter the deeper one together and pair off at random — so a third of the genome of three closely spaced species tells a different story from the species themselves.

Example 25.7 (Reading a genome’s history)

The five hundred cichlid species of Lake Victoria arose in the fifteen thousand years since the lake refilled — too fast for mutation to supply the differences between them. Their genomes show how: the variants that distinguish a bottom-feeder from an algae-scraper were present as standing variation in the river fish that colonised the lake, sorted and recombined, and augmented by hybridisation between lineages; the opsin genes were tuned to clear or turbid water by substitutions that the ω\omega of their branches shows to have been selected; and the gene trees disagree with one another across the genome in exactly the pattern incomplete lineage sorting and gene flow predict. A radiation is not the origin of new genes but the redistribution of old alleles among new combinations under new selection — which the sequences record, allele by allele, if one knows how to read a tree that is really a forest.

Cichlids of an African lake: hundreds of species in a few thousand years, their differences drawn largely from variation their common ancestors already carried.
Cichlids of an African lake: hundreds of species in a few thousand years, their differences drawn largely from variation their common ancestors already carried.

Remark 25.8 (Chance as the baseline)

The neutral theory is not a claim that selection is unimportant; it is a statement of what happens when selection is absent, precise enough to be tested against. Because the neutral rate is the mutation rate, divergence measures time; because neutral change fills the sites that selection leaves free, ω\omega measures constraint; because neutral lineages coalesce at a known rate, the disagreement of gene trees measures the size of populations no one saw. Selection is then what remains when chance has been subtracted — a gene with ω\omega above one, a sweep that has emptied variation from a region, a tree that disagrees in a direction drift cannot explain. The Antarctic icefish’s antifreeze, the ruminant’s stomach lysozyme, the primate’s third opsin are the exceptions found this way, and each is a story of a copied gene and a new job. The rest of the genome is a clock.

25.6 Exercises

Exercise 25.1

State the neutral substitution rate and explain in words why it does not depend on population size.

Solution

Solution of Exercise 25.1.

k=μk = \mu: the neutral substitution rate per site per generation equals the mutation rate. Each generation 2Nμ2N\mu new neutral mutations arise and each has probability 1/2N1/2N of fixing; a larger population makes more mutations and gives each a proportionally smaller chance, and the two effects cancel exactly.

Exercise 25.2

Define orthologue, paralogue and pseudogene, with one example each from the globin family.

Solution

Solution of Exercise 25.2.

Orthologues: genes in two species descended from one gene in their common ancestor — human β\beta-globin and mouse β\beta-globin. Paralogues: genes in one genome descended from a duplication — human α\alpha- and β\beta-globin, or β\beta and γ\gamma. Pseudogene: a decayed copy that no longer encodes a protein — the ψβ\psi\beta gene in the human β\beta-globin cluster, complete with stop codons.

Exercise 25.3

What does ω=dN/dS\omega = d_{N}/d_{S} measure? Interpret ω=0.02\omega = 0.02, ω=1.0\omega = 1.0 and ω=2.5\omega = 2.5, and name a gene of each kind.

Solution

Solution of Exercise 25.3.

ω\omega compares the rate of amino-acid-changing substitution with the rate of silent substitution, which is the neutral baseline. 0.020.02: strong purifying selection, 98%98\,\% of amino-acid changes removed — actin, histone H3. 1.01.0: no constraint — a pseudogene. 2.52.5: positive selection, amino-acid changes favoured — the peptide-binding groove of MHC, the envelope of HIV, ruminant lysozyme.

Exercise 25.4

Why did Kimura conclude that most substitutions are neutral? Restate the argument with the numbers in the chapter’s opening.

Solution

Solution of Exercise 25.4.

Proteins compared across mammals change at about one substitution per site per billion years; over a genome of a billion coding sites and a generation of a few years that is of the order of one substitution fixed per generation, or one every two years. Haldane’s cost of selection allows about one selected substitution per three hundred generations. The rates differ by a factor of a hundred or more, so nearly all fixed changes must be neutral.

Exercise 25.5 ★★

N=5000N = 5000. Compute the fixation probability of a new mutation with s=0s = 0, s=0.001s = 0.001, s=0.01s = 0.01 and s=0.001s = -0.001, using Kimura’s formula, and comment on which are effectively neutral.

Solution

Solution of Exercise 25.5.

Neutral: 1/2N=1041/2N = 10^{-4}. s=0.001s = 0.001, 4Ns=204Ns = 20: (1e0.002)/(1e20)=0.002(1 - \mathrm{e}^{-0.002})/ (1 - \mathrm{e}^{-20}) = 0.002. s=0.01s = 0.01: 0.01980.0198. s=0.001s = -0.001: (1e0.002)/(1e20)=0.002e20=4×1012(1 - \mathrm{e}^{0.002})/(1 - \mathrm{e}^{20}) = 0.002\,\mathrm{e}^{-20} = 4\times 10^{-12}. None is effectively neutral: that would need 4Ns1|4Ns| \lesssim 1, i.e. s5×105|s| \lesssim 5\times 10^{-5}; a selection coefficient of a tenth of a per cent is already decisive in a population of five thousand.

Exercise 25.6 ★★

Two species differ at 12%12\,\% of the sites of a gene whose neutral rate is 2×1092 \times 10^{-9}\, per site per year. Estimate their divergence time with and without the Jukes–Cantor correction. At what raw difference does the correction reach a factor of two?

Solution

Solution of Exercise 25.6.

Uncorrected: T=0.12/(2×2×109)=30MyrT = 0.12/(2\times 2\times 10^{-9}) = 30\,\mathrm{Myr}. Jukes–Cantor: d=34ln(10.16)=0.131d = -\tfrac{3}{4}\ln(1 - 0.16) = 0.131, T=33MyrT = 33\,\mathrm{Myr}. The correction reaches a factor of two when 34ln(14p/3)=2p-\tfrac{3}{4} \ln(1 - 4p/3) = 2p, at p0.6p \approx 0.6: by then the raw difference has lost most of its information.

Exercise 25.7 ★★

A gene of 300 codons has 700 nonsynonymous and 200 synonymous sites. Between two species it shows 7 nonsynonymous and 10 synonymous differences. Compute pNp_{N}, pSp_{S} and ω\omega (uncorrected). Which selection regime? What if the counts were 35 and 10?

Solution

Solution of Exercise 25.7.

pN=7/700=0.010p_{N} = 7/700 = 0.010, pS=10/200=0.050p_{S} = 10/200 = 0.050, ω=0.2\omega = 0.2: purifying selection removing four fifths of amino-acid changes. With 3535 and 1010: pN=0.05p_{N} = 0.05, ω=1\omega = 1: no detectable constraint — a pseudogene, or a gene in which positive selection at some sites balances purifying selection at others.

Exercise 25.8 ★★

Human and chimpanzee neutral sequences differ at 1.2%1.2\,\%; the split was 6.5Myr6.5\,\mathrm{Myr} ago, generation time 25yr25\,\mathrm{yr}. Estimate the mutation rate per site per generation, and the number of new mutations in a diploid genome of 6×1096\times 10^{9} bases each generation.

Solution

Solution of Exercise 25.8.

k=d/2T=0.012/(1.3×107)=9.2×1010k = d/2T = 0.012/(1.3\times 10^{7}) = 9.2\times 10^{-10} per site per year; μ=25k=2.3×108\mu = 25k = 2.3\times 10^{-8} per site per generation; 6×109×2.3×1081406\times 10^{9}\times 2.3\times 10^{-8} \approx 140 new mutations per diploid genome per generation. Direct sequencing of parents and children finds about seventy: the divergence estimate is inflated by polymorphism already present in the ancestral population, which adds a few hundred thousand generations to the apparent split.

Exercise 25.9 ★★

With N=50000N = 50\,000 and T=100000T = 100\,000 generations between the two splits, compute the fraction of gene trees discordant with the species tree for human, chimpanzee and gorilla. What TT would make it 10%10\,\%?

Solution

Solution of Exercise 25.9.

eT/2N=e1=0.37\mathrm{e}^{-T/2N} = \mathrm{e}^{-1} = 0.37; discordant fraction 23×0.37=0.25\tfrac{2}{3} \times 0.37 = 0.25. For 10%10\,\%: eT/2N=0.15\mathrm{e}^{-T/2N} = 0.15, T=2Nln(1/0.15)=190000T = 2N\ln(1/0.15) = 190\,000 generations.

Exercise 25.10 ★★★

Show that Kimura’s formula reduces to 1/2N1/2N as s0s \to 0 and to 2s2s for 4Ns14Ns \gg 1, and that for s<0s < 0 with 4Ns14N|s| \gg 1 it behaves as 2se4Ns2|s|\,\mathrm{e}^{-4N|s|}. Hence estimate how strongly deleterious a mutation must be, in a population of 10410^{4}, to have its fixation probability reduced a hundredfold below neutral.

Solution

Solution of Exercise 25.10.

As s0s \to 0: 1e2s2s1 - \mathrm{e}^{-2s} \approx 2s and 1e4Ns4Ns1 - \mathrm{e}^{-4Ns} \approx 4Ns, ratio 1/2N1/2N. For 4Ns14Ns \gg 1 the denominator is 11 and the numerator 2s\approx 2s. For s<0s < 0 with 4Ns14N|s| \gg 1: numerator 1e2s2s1 - \mathrm{e}^{2|s|} \approx -2|s|, denominator 1e4Nse4Ns1 - \mathrm{e}^{4N|s|} \approx -\mathrm{e}^{4N|s|}, so P2se4NsP \approx 2|s|\,\mathrm{e}^{-4N|s|}. A hundredfold below neutral in N=104N = 10^{4}: 2se4×104s=5×1072|s|\,\mathrm{e}^{-4\times 10^{4}|s|} = 5\times 10^{-7}; s=1.6×104|s| = 1.6\times 10^{-4} gives 3.2×104×e6.4=5.3×1073.2\times 10^{-4}\times \mathrm{e}^{-6.4} = 5.3\times 10^{-7}. So s1.6×104|s| \approx 1.6\times 10^{-4}, 4Ns6.54N|s| \approx 6.5: a disadvantage of a hundredth of a per cent is enough, in ten thousand individuals, to make fixation a hundred times rarer than chance.

Exercise 25.11 ★★★

A relative-rate test: an outgroup OO is at corrected distance 0.300.30 from mouse and 0.240.24 from human, and mouse and human are 0.200.20 apart. Compute the mouse and human branch lengths since their split (assume the split point is equidistant from OO on the two paths), and their ratio. With a mouse generation of 0.5yr0.5\,\mathrm{yr} and a human one of 25yr25\,\mathrm{yr}, what would a strict per-generation clock predict for the ratio, and what does the observation say about mutation rates per generation?

Solution

Solution of Exercise 25.11.

With the split point at distance aa from OO on both paths, a+m=0.30a + m = 0.30, a+h=0.24a + h = 0.24, m+h=0.20m + h = 0.20: mh=0.06m - h = 0.06, so m=0.13m = 0.13, h=0.07h = 0.07, ratio 1.91.9. A strict per-generation clock would give the ratio of generation counts, 25/0.5=5025/0.5 = 50. The observed 1.91.9 means the mouse has accumulated only twice the human change per year despite fifty times the generations: the mutation rate per generation must be some 2525 times higher in humans — more germ-line cell divisions per generation — so that per year the two lineages differ by only a factor of two.

Exercise 25.12 ★★★

A duplicated gene is released from constraint (ω=1\omega = 1) at neutral rate k=2×109k = 2 \times 10^{-9}\, per site per year. In a 1000-base coding sequence, roughly how many years until the first stop codon or frameshift is expected (assume about 4%4\,\% of random point mutations in a coding sequence create a stop codon, and ignore indels)? Compare with the time selection has to find a new function, and explain why most duplicates die.

Solution

Solution of Exercise 25.12.

Stop-creating substitutions arrive at 1000×2×109×0.04=8×1081000\times 2\times 10^{-9}\times 0.04 = 8\times 10^{-8} per year: the first is expected after about 12Myr12\,\mathrm{Myr} (frameshifts roughly halve this). In that window a beneficial mutation, one in 10310^{3} to 10510^{5} of all mutations, is far less likely to arrive than a stop codon, and even when it arrives it fixes with probability only 2s2s. So the copy is usually dead before selection can find it a use — the observed half-life of duplicates in vertebrates is a few million years.

25.7 Problem: The Rate of Change

Problem 25.1

Weekend problem — a genome’s history in numbers: the mutation rate from the human–chimpanzee divergence, Kimura’s substitution load, the fixation odds of selected mutations, a gene’s ω\omega from codon counts, the clock read for a duplication, and the discordance of gene trees among the great apes, ending on the mutation rate, the ω\omega of the gene and the discordant fraction

Data: human–chimpanzee neutral divergence d=0.012d = 0.012; split 6.5Myr6.5\,\mathrm{Myr} ago; generation 25yr25\,\mathrm{yr}; diploid genome 6×1096\times 10^{9} bases, 1.5%1.5\,\% coding. Ancestral population N=50000N = 50\,000. Gene: 400 codons, 920 nonsynonymous and 280 synonymous sites; human–mouse differences 46 nonsynonymous and 84 synonymous. Human–mouse split 90Myr90\,\mathrm{Myr}. Globin: α\alpha and β\beta differ at 55%55\,\% of amino-acid sites (raw); haemoglobin rate 1×1091 \times 10^{-9} per site per year. Apes: TT between the gorilla and chimpanzee splits 8000080\,000 generations. Haldane’s limit: one selected substitution per 300 generations.

Part I — The rate.

  1. From d=2kTd = 2kT, compute kk per site per year, then μ\mu per site per generation.
  2. New mutations per diploid genome per generation, and how many of them fall in coding sequence.
  3. Substitutions fixed per generation in the whole genome at the neutral rate (the population fixes μ\mu per site per generation). Compare with Haldane’s limit and draw Kimura’s conclusion.
  4. In a population of N=50000N = 50\,000, how many new neutral mutations arise per site per generation in the whole population, and what fraction of them will ever fix?
  5. How long does a neutral mutation destined to fix take, in generations and years? Compare with the time since the human–chimpanzee split.
  6. Explain why, despite question 5, the divergence between two species measures the time since their split and not the time since their alleles fixed.

Part II — Selection’s odds.

  1. Fixation probability of a neutral mutation in N=50000N = 50\,000, and of one with s=0.001s = 0.001, s=0.01s = 0.01, s=0.1s = 0.1.
  2. For which ss is 4Ns=14Ns = 1? Interpret the boundary.
  3. A deleterious mutation with s=0.001s = -0.001: fixation probability relative to neutral. And s=105s = -10^{-5}?
  4. If a fraction 10510^{-5} of new mutations in a gene are beneficial with s=0.01s = 0.01 and the rest neutral, what fraction of the substitutions in that gene are adaptive? Comment on how rarely selection shows in the sequence even when it acts.
  5. Compute pNp_{N} and pSp_{S} for the human–mouse gene, correct each with Jukes–Cantor, and give ω\omega.
  6. Interpret ω\omega; and estimate the neutral rate the gene’s dSd_{S} implies (per site per year) using the split time. Is it consistent with question 1?

Part III — Clocks and copies.

  1. Correct the α\alphaβ\beta globin difference for multiple hits (use d=ln(1p)d = -\ln(1 - p), the protein version with many states), and date the duplication with the haemoglobin rate.
  2. The date falls near the origin of jawed vertebrates (450 to 500Myr450\text{ to }500\,\mathrm{Myr}). What does having separate α\alpha and β\beta chains allow that a single chain does not?
  3. Fibrinopeptides change at 8×1098\times 10^{-9} per site per year. Two species split 40Myr40\,\mathrm{Myr} ago: expected raw difference? Why is this protein useless for dating splits of 500Myr500\,\mathrm{Myr}?
  4. Histone H4 changes at 101110^{-11}: how many substitutions in 100 sites over 1000Myr1000\,\mathrm{Myr}? Why is it useless for dating recent splits?
  5. A duplicate gene drifts with ω=1\omega = 1 from the moment of duplication. If 4%4\,\% of substitutions in coding sequence create stops, and the gene has 1200 sites at rate 2×1092\times 10^{-9} per site per year, estimate the expected time to the first stop codon.
  6. Why does whole-genome duplication give duplicates a better chance of survival than a single tandem duplication?

Part IV — Trees within trees.

  1. Probability that a human and a chimpanzee lineage fail to coalesce in the 8000080\,000 generations before the gorilla split.
  2. Fraction of gene trees discordant with the species tree. Observed: about 30%30\,\%. Consistent?
  3. Of the discordant trees, what fraction group human with gorilla, and what fraction chimpanzee with gorilla? What would a strong departure from equality indicate?
  4. If the ancestral population had been N=10000N = 10\,000, what discordant fraction would you expect? What does the observed fraction therefore say about our ancestors’ numbers?
  5. Why does concatenating a thousand genes into one alignment give a confident but possibly wrong tree, and what does a coalescent method do instead?
  6. Neanderthal DNA in non-African humans is about 2%2\,\% and the trees carrying it group some Europeans with Neanderthals more often than Africans. Why is this not incomplete lineage sorting?
  7. Summarise: μ\mu (question 1), ω\omega of the gene (question 11), and the discordant fraction (question 20).
Solution

Solution of Problem 25.1.

1. k=0.012/(2×6.5×106)=9.2×1010k = 0.012/(2\times 6.5\times 10^{6}) = 9.2\times 10^{-10} per site per year; μ=25k=2.3×108\mu = 25k = 2.3\times 10^{-8} per site per generation. 2. 6×109×2.3×1081406\times 10^{9}\times 2.3\times 10^{-8} \approx 140 new mutations per diploid genome; 1.5%1.5\,\% of them, about 22, in coding sequence. 3. Per haploid genome, 3×109×2.3×108703\times 10^{9}\times 2.3\times 10^{-8} \approx 70 substitutions fixed per generation, against Haldane’s 1/3001/300: twenty thousand times more than selection could drive, so the overwhelming majority are neutral. 4. 2Nμ=105×2.3×108=2.3×1032N\mu = 10^{5}\times 2.3\times 10^{-8} = 2.3\times 10^{-3} new mutations per site per generation in the population; 1/2N=1051/2N = 10^{-5} of them will fix. 5. 4N=2000004N = 200\,000 generations, 5Myr5\,\mathrm{Myr} — nearly as long as the time since the split. 6. Divergence counts the mutations that arose along the two separate lineages since their common ancestor; each fixes within its own lineage, and the long-run rate is μ\mu whatever the time each takes, since mutations arise steadily and the lag only delays, without changing, the flow. (The only correction is for polymorphism present at the split, which makes the common ancestor of two alleles older than the species split.) 7. Neutral 10510^{-5}; s=0.001s = 0.001: 0.0020.002; s=0.01s = 0.01: 0.01980.0198; s=0.1s = 0.1: 1e0.2=0.181 - \mathrm{e}^{-0.2} = 0.18. 8. s=1/4N=5×106s = 1/4N = 5\times 10^{-6}: below this drift swamps selection and the mutation behaves as neutral; above it selection governs the odds. 9. s=0.001s = -0.001: P=0.002e200P = 0.002\,\mathrm{e}^{-200}, zero for all purposes — 108510^{-85} of neutral. s=105s = -10^{-5}, 4Ns=24N|s| = 2: P=2×105/(e21)=3.1×106P = 2\times 10^{-5}/(\mathrm{e}^{2} - 1) = 3.1\times 10^{-6}, 31%31\,\% of neutral: mildly deleterious mutations do fix. 10. Beneficial substitutions per site per generation: 2N×105μ×2s=105×105×0.02μ=0.02μ2N\times 10^{-5}\mu\times 2s = 10^{5}\times 10^{-5}\times 0.02\,\mu = 0.02\mu; neutral: μ\approx\mu. Adaptive fraction 0.02/1.022%0.02/1.02 \approx 2\,\%. Though each beneficial mutation is two thousand times likelier to fix than a neutral one, they are so rare that only one substitution in fifty is adaptive: selection acts and barely shows. 11. pN=46/920=0.050p_{N} = 46/920 = 0.050, pS=84/280=0.30p_{S} = 84/280 = 0.30. Corrected: dN=34ln(10.0667)=0.052d_{N} = -\tfrac{3}{4}\ln(1 - 0.0667) = 0.052, dS=34ln(10.40)=0.38d_{S} = -\tfrac{3}{4} \ln(1 - 0.40) = 0.38. ω=0.14\omega = 0.14. 12. Purifying selection removing about 86%86\,\% of amino-acid changes — a typical gene. k=dS/2T=0.38/(1.8×108)=2.1×109k = d_{S}/2T = 0.38/(1.8\times 10^{8}) = 2.1\times 10^{-9} per site per year: twice the human–chimp value, as expected for a comparison that includes the fast-ticking rodent lineage; the same order of magnitude. 13. d=ln(10.55)=0.80d = -\ln(1 - 0.55) = 0.80 per site; T=0.80/(2×109)=400MyrT = 0.80/(2\times 10^{-9}) = 400\,\mathrm{Myr}. 14. Two different chains make the α2β2\alpha_{2}\beta_{2} tetramer, whose cooperative oxygen binding, Bohr effect and allosteric control by 2,3-bisphosphoglycerate depend on the interface between unlike subunits; and the β\beta cluster could then diversify into embryonic, fetal and adult chains with different affinities. 15. 2kT=2×8×109×4×107=0.642kT = 2\times 8\times 10^{-9}\times 4\times 10^{7} = 0.64 substitutions per site; with saturation among twenty amino acids the raw difference is about 0.95(1e0.64/0.95)0.470.95(1 - \mathrm{e}^{-0.64/0.95}) \approx 0.47. At 500Myr500\,\mathrm{Myr}, d=8d = 8: every site has changed many times and the raw difference sits at its ceiling of 95%95\,\%, carrying no information about time. 16. 1011×109×100×2=210^{-11}\times 10^{9}\times 100\times 2 = 2 substitutions between two lineages in 100 sites over a billion years. For a 10Myr10\,\mathrm{Myr} split the expectation is 0.020.02: almost always zero differences, no resolution. 17. 1200×2×109×0.04=9.6×1081200\times 2\times 10^{-9}\times 0.04 = 9.6\times 10^{-8} per year: the first stop is expected after about 10Myr10\,\mathrm{Myr}. 18. Doubling the whole genome doubles every gene at once, so the stoichiometry of interacting proteins is preserved and no dosage imbalance selects against the copies; a lone tandem duplicate changes one gene’s dose and is often deleterious, or immediately redundant and free to decay. 19. e80000/100000=e0.8=0.45\mathrm{e}^{-80\,000/100\,000} = \mathrm{e}^{-0.8} = 0.45. 20. 23×0.45=0.30\tfrac{2}{3}\times 0.45 = 0.30: consistent with the observed 30%30\,\%. 21. Half each, 15%15\,\% of all trees grouping human with gorilla and 15%15\,\% chimpanzee with gorilla. A clear excess of one would indicate gene flow between those two species after their split, which drift cannot produce. 22. N=10000N = 10\,000: e4=0.018\mathrm{e}^{-4} = 0.018, discordance 1.2%1.2\,\%. The observed 30%30\,\% therefore requires an ancestral population of some fifty thousand — several times the effective size of humans today. 23. Concatenation assumes every gene has the species tree; with incomplete lineage sorting the genes have many trees, and for short internal branches the commonest gene tree can differ from the species tree, so more data make the wrong answer more confident. A coalescent method computes, for each candidate species tree, the expected distribution of gene trees, and picks the species tree that best explains the frequencies observed. 24. Incomplete lineage sorting happened in the common ancestral population of all modern humans, so it would make every human population equally related to Neanderthals. An excess in non-Africans means Neanderthal alleles entered after the ancestors of non-Africans left Africa: admixture, some fifty thousand years ago. 25. μ2.3×108\mu \approx 2.3\times 10^{-8} per site per generation; ω0.14\omega \approx 0.14; about 30%30\,\% of gene trees discordant with the species tree.

Terms defined in this chapter

See all 479 terms in the glossary