Biology · Book 3 · Bachelor Year 1

University Biology — Year 1

University Biology — Year 1 · Bachelor Year 1

19Gene Expression: Transcription and Translation

A gene of thirty thousand base pairs is read in a quarter of an hour into a working copy, cut down to two thousand letters, shipped out of the nucleus, and translated by twenty-five ribosomes at once into a protein of five hundred amino acids — one every four seconds, for as long as the copy lasts. The cell spends more energy on this than on anything else it does. This chapter follows the information from DNA to protein: the copying of a gene into RNA, the editing of that RNA in eukaryotes, the code that maps triplets of bases onto amino acids, the ribosome that reads it, and what happens to a protein once it is made.

19.1 From gene to protein

Proposition 19.1 (The flow of information)

Genetic information flows from DNA to RNA to protein. Transcription copies one strand of a gene into an RNA of the same sequence as the other strand (with U for T); translation reads the messenger RNA three bases at a time and assembles the corresponding amino acids into a polypeptide. Every step uses the pairing of bases: DNA with RNA in transcription, messenger with transfer RNA in translation. The sequence of a protein is thus a transcript of the sequence of its gene, and a change of one base can change one amino acid — the link between mutation and phenotype. Information does not flow back from protein to nucleic acid; the one reverse step, RNA copied into DNA by retroviruses, does not touch protein.

19.2 Transcription

Definition 19.2 (RNA polymerase, promoter, transcription)

RNA polymerase synthesises RNA on a DNA template, 535' \to 3', from the four ribonucleoside triphosphates, without a primer. It starts at a promoter, a sequence just upstream of the gene that it recognises and binds; in bacteria a subunit called σ\sigma (sigma) does the recognising, at two conserved stretches ten and thirty-five pairs before the start. The enzyme opens about fifteen pairs of the helix into a transcription bubble, copies the template strand (read 353' \to 5') into an RNA identical in sequence to the other, coding strand, and moves along at about fifty nucleotides a second, re-closing the helix behind it; the RNA peels off as it is made. It stops at a terminator: in bacteria a self-complementary sequence whose RNA folds into a hairpin that pulls the transcript free. Many polymerases can follow one another along a gene, so a transcript can be started every second.

Transcription. The polymerase opens a bubble, pairs ribonucleotides to the template strand and joins them 5' 3'; the RNA leaves through a channel as the enzyme advances.
Transcription. The polymerase opens a bubble, pairs ribonucleotides to the template strand and joins them 535' \to 3'; the RNA leaves through a channel as the enzyme advances.

Proposition 19.3 (Eukaryotic transcription and RNA processing)

Eukaryotes have three polymerases: I for the large ribosomal RNAs, II for messenger RNAs, III for transfer and small RNAs. Polymerase II does not recognise its promoter alone: a set of general transcription factors assembles on it (at a TATA sequence about thirty pairs upstream in many genes) and recruits the enzyme, and regulatory proteins bound near or far modulate the rate (Chapter 20). The primary transcript is processed in the nucleus: a modified guanine cap is added to the 55' end, a tail of some two hundred adenines (poly-A) to the 33' end, and the introns are removed by splicing — a complex of small RNAs and proteins, the spliceosome, recognises the GU at each intron’s start and the AG at its end, cuts, and joins the exons. Only then is the mature messenger exported to the cytosol. Many genes are spliced in more than one way (alternative splicing), so that one gene yields several proteins; the twenty thousand human genes make perhaps a hundred thousand.

Evidence. Sharp and Roberts (1977) hybridised a viral messenger RNA to the DNA of its gene and looked at the hybrids in the electron microscope: the RNA paired with the DNA in several stretches, and between them the DNA looped out unpaired — the introns, present in the gene and absent from the message. Bacterial genes, treated the same way, gave no loops. The size of a gene’s primary transcript, measured in the nucleus, matches the gene; the size of the cytoplasmic messenger matches the exons alone.

From gene to messenger in a eukaryote: the whole gene is transcribed, then the introns are cut out and the exons joined, a cap and a poly-A tail are added, and the mature message leaves the nucleus.
From gene to messenger in a eukaryote: the whole gene is transcribed, then the introns are cut out and the exons joined, a cap and a poly-A tail are added, and the mature message leaves the nucleus.

19.3 The genetic code

Definition 19.4 (The genetic code)

The genetic code maps triplets of messenger bases, codons, onto amino acids. Of the 64 codons, 61 specify the twenty amino acids and three (UAA, UAG, UGA) are stop signals; AUG specifies methionine and is also the start codon, so every new polypeptide begins with methionine. The code is degenerate — most amino acids have several codons, differing mostly at the third position — non-overlapping, read without punctuation in a reading frame fixed by the start codon, and almost universal, the same in bacteria, plants and animals with minor variants in mitochondria: a human gene put into a bacterium is read correctly.

firstsecond basethird
baseUCAGbase
UPhe Phe Leu LeuSer Ser Ser SerTyr Tyr stop stopCys Cys stop TrpU C A G
CLeu Leu Leu LeuPro Pro Pro ProHis His Gln GlnArg Arg Arg ArgU C A G
AIle Ile Ile MetThr Thr Thr ThrAsn Asn Lys LysSer Ser Arg ArgU C A G
GVal Val Val ValAla Ala Ala AlaAsp Asp Glu GluGly Gly Gly GlyU C A G
The genetic code. Each cell lists the four codons with the given first and second bases and third base U, C, A, G in order. AUG is methionine and the start signal; UAA, UAG and UGA are stops. The third base often does not matter: the code is degenerate.

Proposition 19.5 (How the code was read)

The code is a non-overlapping triplet code, and each codon’s meaning can be determined by chemistry.

Evidence. Crick and Brenner (1961) made mutations in a phage gene that added or removed one base: one such change destroyed the gene’s function, as did two, but three insertions (or three deletions) close together restored it — the message is read in threes, from a fixed starting point, and an insertion shifts the frame of everything downstream. Nirenberg and Matthaei (1961) added a synthetic RNA of uracils only to a cell-free extract of E. coli with the twenty amino acids: it made a chain of phenylalanine only, so UUU means Phe; other synthetic messengers, and then binding of single trinucleotides to ribosomes with their transfer RNAs, assigned all sixty-four by 1966.

Example 19.6 (Reading a message)

The messenger 55'-…GCAUGGCUUUCGGAUAA…-33' is read from the AUG: AUG GCU UUC GGA UAA — Met-Ala-Phe-Gly-stop: a peptide of four residues. Delete the first G after AUG and the frame shifts: AUG CUU UCG GAU AA… — Met-Leu-Ser-Asp…, a different protein of a different length. Change UUC to UUU and nothing changes: both are Phe, a silent mutation.

19.4 Translation

Definition 19.7 (Transfer RNA and its synthetases)

A transfer RNA (Chapter 11) is the adaptor between codon and amino acid: its anticodon, three bases in a loop, pairs antiparallel with a codon of the message, and its 33' end carries the corresponding amino acid. The pairing at the third codon position is loose (wobble), so about forty tRNAs suffice for sixty-one codons. Each tRNA is loaded by its own aminoacyl-tRNA synthetase, an enzyme that recognises both the amino acid and the tRNA and joins them in two steps at the cost of one ATP (split to AMP: two high-energy bonds), proofreading the amino acid as it does so. These twenty enzymes are where the code is actually implemented: the ribosome does not check which amino acid a tRNA carries, only that its anticodon matches.

Definition 19.8 (The ribosome)

The ribosome is a particle of two subunits, each of ribosomal RNA and proteins (in bacteria: a small 30S subunit and a large 50S subunit, together 70S, 2.5MDa2.5\,\mathrm{MDa}; eukaryotic ribosomes are larger, 80S). The small subunit binds the messenger and decodes it; the large subunit holds the tRNAs and forms the peptide bond — the catalyst is the ribosomal RNA itself. Three tRNA sites span both subunits: A (aminoacyl, where the next charged tRNA enters), P (peptidyl, holding the growing chain), E (exit).

The ribosome: two subunits clamped on the messenger, two transfer RNAs in the cleft, the growing chain leaving through a tunnel in the large subunit.
The ribosome: two subunits clamped on the messenger, two transfer RNAs in the cleft, the growing chain leaving through a tunnel in the large subunit.

Proposition 19.9 (The cycle of elongation)

Translation starts when the small subunit finds the start codon — in bacteria by pairing a ribosomal RNA sequence with a site just upstream of the AUG, in eukaryotes by binding the cap and scanning to the first AUG — and the initiator tRNA (methionine) settles in the P site; the large subunit then joins. Each round of elongation adds one residue: a charged tRNA whose anticodon matches the A-site codon is delivered by an elongation factor and checked (one GTP); the ribosomal RNA of the large subunit transfers the growing chain from the P-site tRNA onto the amino acid of the A-site tRNA, forming the peptide bond; the ribosome moves one codon along (a second GTP), shifting the tRNAs to P and E, and the empty one leaves. At a stop codon a release factor enters the A site and the chain is hydrolysed free. Bacterial ribosomes add fifteen to twenty residues a second, eukaryotic ones two to five; several ribosomes read one message at once, forming a polysome.

One round of elongation. A charged tRNA is admitted to the A site when its anticodon matches; the ribosomal RNA joins the chain to its amino acid; the ribosome moves on by one codon, and the empty tRNA leaves. Two GTP per residue, plus the two high-energy bonds spent in charging the tRNA.
One round of elongation. A charged tRNA is admitted to the A site when its anticodon matches; the ribosomal RNA joins the chain to its amino acid; the ribosome moves on by one codon, and the empty tRNA leaves. Two GTP per residue, plus the two high-energy bonds spent in charging the tRNA.
Polysomes under the electron microscope: ribosomes strung along a messenger, each making its own copy of the protein.
Polysomes under the electron microscope: ribosomes strung along a messenger, each making its own copy of the protein.

Method 19.10 (Reckoning a protein’s synthesis)

  1. Length: a protein of nn residues needs a coding sequence of 3n+33n + 3 nucleotides (the stop codon included), inside a messenger longer by its untranslated ends and, in the gene, by its introns.
  2. Time: divide nn by the elongation rate (15 to 2015\text{ to }20\, per second in bacteria, 2 to 52\text{ to }5\, in eukaryotes); the message is being read by one ribosome every 80 to 10080\text{ to }100\, nucleotides, so the output per message is one chain every (spacing / rate) seconds.
  3. Energy: four high-energy phosphate bonds per residue (two to charge the tRNA, two GTP on the ribosome), plus the transcription of the message at two per nucleotide, shared among all the chains that message yields.
  4. Fidelity: about one wrong amino acid in 10410^4 codons; a protein of 500500\, residues is wrong somewhere in one copy out of twenty — tolerable because proteins are replaceable and errors are not inherited.

19.5 After translation

Proposition 19.11 (What happens to a new polypeptide)

The chain folds as it emerges, helped by chaperones (Chapter 12). Its destination is written in its sequence: a signal peptide at the N-terminus binds a signal recognition particle that halts translation, docks the ribosome on the endoplasmic reticulum, and threads the chain into its lumen (Chapter 6); other sequences direct proteins into mitochondria, chloroplasts, the nucleus or peroxisomes after synthesis; a protein with no signal stays in the cytosol. Many proteins are modified: the initial methionine removed, sugars attached in the ER and Golgi, phosphates added and removed by kinases and phosphatases, lipids attached, pieces cut out (insulin is made as one chain and cut into two; zymogens are cut to activate them). And every protein is eventually degraded: tagged with the small protein ubiquitin and unfolded and digested by the proteasome, after a life of minutes (regulatory proteins) to months (haemoglobin) — so that the proteins present are those the cell is currently making.

Example 19.12 (The numbers of a cell)

A growing E. coli holds some 2000020\,000 ribosomes, each adding twenty residues a second: 4×1054 \times 10^{5} residues a second, a million proteins of 300300\, residues in an hour — its own content, which is what doubling every hour requires. Half its energy goes to protein synthesis. A liver cell holds ten million ribosomes and makes proteins for weeks of use rather than for division; a plasma cell, secreting antibody, devotes nearly all of them to one protein and pours out two thousand molecules a second.

19.6 Exercises

Exercise 19.1

Give the RNA transcribed from the template strand 33'-TACGGATTC-55', and the coding strand.

Solution

Solution of Exercise 19.1.

RNA 55'-AUGCCUAAG-33'; coding strand 55'-ATGCCTAAG-33'.

Exercise 19.2

List the three modifications a eukaryotic primary transcript undergoes before it leaves the nucleus.

Solution

Solution of Exercise 19.2.

A 55' cap (modified guanine), a 33' poly-A tail, and the removal of the introns by splicing.

Exercise 19.3

Using the code table, translate 55'-AUGCCGAAAGUUUGA-33'.

Solution

Solution of Exercise 19.3.

AUG CCG AAA GUU UGA: Met-Pro-Lys-Val, then stop.

Exercise 19.4

Name the three tRNA sites of the ribosome and what happens in each.

Solution

Solution of Exercise 19.4.

A: the incoming charged tRNA is admitted and checked. P: the tRNA carrying the growing chain; the bond forms between its chain and the A-site amino acid. E: the emptied tRNA on its way out.

Exercise 19.5 ★★

A messenger of 24002400\, nucleotides has 150150\, of 55' untranslated region and 450450\, of 33' untranslated region and poly-A. How many residues has the protein? How long does one ribosome take to make it at four residues a second, and how many ribosomes can read the message at once at one per 9090\, nucleotides?

Solution

Solution of Exercise 19.5.

Coding 2400600=18002400 - 600 = 1800 nucleotides: 599 residues (600 codons, one a stop). Time 599/4=150s599/4 = 150\,\mathrm{s}. Ribosomes 1800/90=201800/90 = 20 at once.

Exercise 19.6 ★★

Explain why three insertions restore a gene’s reading frame while one or two do not, and what the protein made from the triple insertion looks like.

Solution

Solution of Exercise 19.6.

The message is read in threes from a fixed start; one or two extra bases shift every codon downstream and the rest of the protein is gibberish, usually ending at a premature stop. Three extra bases add one codon and restore the frame: the protein has one extra residue and a few wrong ones between the insertions, and often still works.

Exercise 19.7 ★★

A mutation changes the codon CAG to UAG in the middle of a gene of 300300\, codons; another changes CAG to CAA; a third changes it to CGG. Give the effect of each on the protein.

Solution

Solution of Exercise 19.7.

CAG (Gln) to UAG (stop): the chain ends at residue 150, a truncated, almost certainly inactive protein (nonsense mutation). CAG to CAA: still Gln, silent. CAG to CGG: Arg for Gln, a substitution (missense) whose effect depends on the position.

Exercise 19.8 ★★

Explain why the fidelity of translation rests on the aminoacyl-tRNA synthetases rather than on the ribosome, and describe an experiment that showed it (a cysteine attached to its tRNA is chemically converted to alanine; where does the alanine end up?).

Solution

Solution of Exercise 19.8.

The ribosome checks only the codon–anticodon pairing; it cannot see the amino acid. If cysteine on its tRNA is converted chemically to alanine, the ribosome inserts alanine wherever the message says cysteine: the tRNA, not the amino acid, is read. Hence the synthetase, which pairs each amino acid with the right tRNA, is the true translator.

Exercise 19.9 ★★

Compute the energy cost, in high-energy bonds, of a protein of 400400\, residues, and the fraction of it spent on the ribosome.

Solution

Solution of Exercise 19.9.

400×4=1600400\times 4 = 1600 bonds; the ribosome’s two GTP per residue are half of it, the synthetases’ two the other half.

Exercise 19.10 ★★★

In bacteria ribosomes begin translating a message while it is still being transcribed; in eukaryotes they cannot. Explain why (two reasons), and what this difference makes possible in eukaryotes.

Solution

Solution of Exercise 19.10.

In eukaryotes the transcript is made in the nucleus and the ribosomes are in the cytosol, separated by the envelope; and the transcript is not a messenger until it has been spliced and capped. The separation makes possible the processing itself — alternative splicing, the control of export, and a check that the message is complete before it is read.

Exercise 19.11 ★★★

A gene of five exons can be spliced to include or skip exon 3 (9090\, nucleotides) and exon 4 (100100\, nucleotides). List the possible messengers and say which ones keep the reading frame of exon 5. What does this show about alternative splicing?

Solution

Solution of Exercise 19.11.

Four messengers: with both exons (190 nucleotides added), with 3 only (90), with 4 only (100), with neither (0). A frame is kept if the added length is a multiple of 3: 0 and 90 keep it; 100 and 190 shift it, so the messengers with exon 4 alone or both exons read exon 5 out of frame and truncate the protein. Alternative splicing must respect the frame, and exons are often multiples of three for that reason.

Exercise 19.12 ★★★

“The code is a frozen accident: arbitrary, but too costly to change.” Discuss in a paragraph: what in the code looks arbitrary, what looks optimised (the third position, similar codons for similar amino acids), and why a mutation changing a synthetase’s specificity is almost always lethal.

Solution

Solution of Exercise 19.12.

Arbitrary: nothing in chemistry says that UUU must mean Phe; the assignments are conventions held by the synthetases. Optimised: the third position is the most degenerate, so the errors and mutations that fall there are often silent, and codons that differ by one base tend to encode similar amino acids, so a substitution is often mild — the code minimises the damage of error. Frozen: a synthetase that changed its specificity would alter every protein of the cell at once, thousands of them, in the same instant; almost no cell survives that, so the code cannot drift and has been fixed since the common ancestor of all living things.

19.7 Problem: From a Gene to a Protein

Problem 19.1

Weekend problem — a thirty-kilobase gene followed into a five-hundred-residue protein: transcription timed, splicing weighed, ribosomes counted, energy and errors reckoned, ending on the cost of one protein in ATP

A human gene spans 30kb30\,\mathrm{kb} with eight exons totalling 20002000\, nucleotides of mature messenger, of which 15031503\, code (500 residues plus the stop). Polymerase II transcribes at 3030\, nucleotides per second; ribosomes elongate at 44\, residues per second and space themselves one per 8080\, nucleotides; the messenger’s half-life is 2h2\,\mathrm{h}. Costs: two high-energy bonds per nucleotide transcribed, four per residue translated; take one high-energy bond as one ATP. Error rates: 10510^{-5} per nucleotide transcribed, 10410^{-4} per codon translated.

Part I — Transcription and splicing.

  1. How long does one polymerase take to transcribe the gene?
  2. What fraction of the primary transcript is removed by splicing?
  3. How many nucleotides are transcribed for every nucleotide of mature messenger?
  4. Compute the ATP spent transcribing one primary transcript.
  5. If a polymerase starts every 10s10\,\mathrm{s}, how many polymerases are on the gene at once, and how many messengers does the gene produce per hour?
  6. Compute the number of introns and their mean length.
  7. The spliceosome removes each intron in about a minute, in parallel. Does splicing or transcription set the time from gene to messenger?
  8. Compute the physical length of the gene and of the mature messenger (0.34nm0.34\,\mathrm{nm} per nucleotide).

Part II — Translation.

  1. How long does one ribosome take to translate the protein?
  2. How many ribosomes read one messenger at once?
  3. How many protein molecules does one messenger yield per hour?
  4. Over its 2h2\,\mathrm{h} half-life, how many does it yield in all (a message with half-life t1/2t_{1/2} yields, on average, t1/2/ln2t_{1/2}/\ln 2 hours of full production)?
  5. Compute the ATP spent translating one protein.
  6. Compute the ATP spent on transcription per protein, sharing the transcript’s cost among the proteins of question 12.
  7. Compute the total cost of one protein and the fraction due to translation.

Part III — Errors.

  1. Compute the probability that a given messenger carries at least one transcription error in its 15031503\, coding nucleotides.
  2. Compute the probability that a given protein molecule carries at least one translation error.
  3. About a quarter of nucleotide changes are silent and a third of amino-acid substitutions are harmless. What fraction of the protein molecules made are defective?
  4. A defective messenger yields defective proteins for two hours; a defective protein is one molecule. Explain why the cell can afford a translation error rate ten times the transcription rate, and both far above the replication rate.

Part IV — The cell’s budget. The cell makes 30003000 copies of this protein per hour, and in all 3×1093 \times 10^{9} residues of protein per day.

  1. How many messengers of this gene must be present at once (each yielding the number of question 11)?
  2. How many ribosomes are occupied by this protein?
  3. Compute the daily ATP the cell spends on all its protein synthesis (four per residue) and, at 50kJ/mol50\,\mathrm{kJ}/\mathrm{mol}, the power in watts for a cell of 2000µm32000\,\text{µ}\mathrm{m}^{3}.
  4. Compare with the cell’s total power if it consumes 1×1091 \times 10^{9}\, ATP per second.
  5. A drug blocks the spliceosome. Predict its effect on this protein and on a bacterial protein.
  6. State the result: the ATP cost of one molecule of the protein, split between translation and transcription, and the time from the start of transcription to the first finished protein.
Solution

Solution of Problem 19.1.

1. 30000/30=1000s30\,000/30 = 1000\,\mathrm{s}, about 17min17\,\mathrm{min}. 2. 28000/30000=93%28\,000/30\,000 = 93\,\%. 3. 30000/2000=1530\,000/2000 = 15. 4. 2×30000=600002\times 30\,000 = 60\,000 ATP. 5. 1000/10=1001000/10 = 100 polymerases on the gene; 360 messengers per hour. 6. Seven introns, mean 28000/7=400028\,000/7 = 4000\, nucleotides. 7. Transcription (17min17\,\mathrm{min}); the introns are spliced as they are made, and the last one adds only a minute. 8. Gene 30000×0.34nm=10µm30\,000\times 0.34\,\mathrm{nm} = 10\,\text{µ}\mathrm{m}; messenger 0.68µm0.68\,\text{µ}\mathrm{m}. 9. 500/4=125s500/4 = 125\,\mathrm{s}. 10. 2000/80=252000/80 = 25. 11. One chain finishes every 80/4=20s80/4 = 20\,\mathrm{s}: 180 per hour. 12. 2/ln2=2.9h2/\ln 2 = 2.9\,\mathrm{h} of full production: about 520 proteins. 13. 4×500=20004\times 500 = 2000 ATP. 14. 60000/520=11560\,000/520 = 115 ATP per protein. 15. About 2100 ATP, 95%95\,\% of it in translation. 16. 1(1105)15031503×105=1.5%1 - (1 - 10^{-5})^{1503} \approx 1503\times 10^{-5} = 1.5\,\%. 17. 1(1104)5005%1 - (1 - 10^{-4})^{500} \approx 5\,\%. 18. Transcription: 1.5%×0.75×0.670.75%1.5\%\times 0.75\times 0.67 \approx 0.75\% of messengers defective, hence of proteins; translation: 5%×0.67=3.3%5\%\times 0.67 = 3.3\%; about 4%4\,\% of the molecules. 19. A protein error is confined to one molecule, which is soon degraded; a messenger error is copied into hundreds of proteins; a replication error is inherited by every descendant for ever. The cost of an error, and hence the accuracy worth paying for, rises at each step back. 20. 3000/180=173000/180 = 17 messengers present. 21. 17×25=42017\times 25 = 420 ribosomes. 22. 3×109×4=1.2×10103 \times 10^{9}\times 4 = 1.2 \times 10^{10} ATP per day =2×1014mol= 2 \times 10^{-14}\,\mathrm{mol}, 1×109J1 \times 10^{-9}\,\mathrm{J} per day, i.e. 1.2×1014W1.2 \times 10^{-14}\,\mathrm{W}6fW6\,\mathrm{fW} per cubic micrometre. 23. 10910^9 ATP per second is 8.6×10138.6 \times 10^{13} per day: protein synthesis is a hundredth of it in this slowly renewing cell (in a dividing bacterium it is half). 24. No mature messenger: the primary transcripts accumulate in the nucleus and the protein disappears as its messengers decay, within hours. The bacterial protein, whose gene has no introns, is unaffected. 25. About 2100 ATP per molecule — 2000 for translation, a hundred for its share of the transcript; first protein after 1000s1000\,\mathrm{s} of transcription, a minute of processing and export, and 125s125\,\mathrm{s} of translation: about twenty minutes.

Terms defined in this chapter

See all 479 terms in the glossary