University Biology — Year 1 · Bachelor Year 1
19Gene Expression: Transcription and Translation
A gene of thirty thousand base pairs is read in a quarter of an hour into a working copy, cut down to two thousand letters, shipped out of the nucleus, and translated by twenty-five ribosomes at once into a protein of five hundred amino acids — one every four seconds, for as long as the copy lasts. The cell spends more energy on this than on anything else it does. This chapter follows the information from DNA to protein: the copying of a gene into RNA, the editing of that RNA in eukaryotes, the code that maps triplets of bases onto amino acids, the ribosome that reads it, and what happens to a protein once it is made.
19.1 From gene to protein
Proposition 19.1 (The flow of information)
Genetic information flows from DNA to RNA to protein. Transcription copies one strand of a gene into an RNA of the same sequence as the other strand (with U for T); translation reads the messenger RNA three bases at a time and assembles the corresponding amino acids into a polypeptide. Every step uses the pairing of bases: DNA with RNA in transcription, messenger with transfer RNA in translation. The sequence of a protein is thus a transcript of the sequence of its gene, and a change of one base can change one amino acid — the link between mutation and phenotype. Information does not flow back from protein to nucleic acid; the one reverse step, RNA copied into DNA by retroviruses, does not touch protein.
19.2 Transcription
Definition 19.2 (RNA polymerase, promoter, transcription)
RNA polymerase synthesises RNA on a DNA template, , from the four ribonucleoside triphosphates, without a primer. It starts at a promoter, a sequence just upstream of the gene that it recognises and binds; in bacteria a subunit called (sigma) does the recognising, at two conserved stretches ten and thirty-five pairs before the start. The enzyme opens about fifteen pairs of the helix into a transcription bubble, copies the template strand (read ) into an RNA identical in sequence to the other, coding strand, and moves along at about fifty nucleotides a second, re-closing the helix behind it; the RNA peels off as it is made. It stops at a terminator: in bacteria a self-complementary sequence whose RNA folds into a hairpin that pulls the transcript free. Many polymerases can follow one another along a gene, so a transcript can be started every second.
Proposition 19.3 (Eukaryotic transcription and RNA processing)
Eukaryotes have three polymerases: I for the large ribosomal RNAs, II for messenger RNAs, III for transfer and small RNAs. Polymerase II does not recognise its promoter alone: a set of general transcription factors assembles on it (at a TATA sequence about thirty pairs upstream in many genes) and recruits the enzyme, and regulatory proteins bound near or far modulate the rate (Chapter 20). The primary transcript is processed in the nucleus: a modified guanine cap is added to the end, a tail of some two hundred adenines (poly-A) to the end, and the introns are removed by splicing — a complex of small RNAs and proteins, the spliceosome, recognises the GU at each intron’s start and the AG at its end, cuts, and joins the exons. Only then is the mature messenger exported to the cytosol. Many genes are spliced in more than one way (alternative splicing), so that one gene yields several proteins; the twenty thousand human genes make perhaps a hundred thousand.
Evidence. Sharp and Roberts (1977) hybridised a viral messenger RNA to the DNA of its gene and looked at the hybrids in the electron microscope: the RNA paired with the DNA in several stretches, and between them the DNA looped out unpaired — the introns, present in the gene and absent from the message. Bacterial genes, treated the same way, gave no loops. The size of a gene’s primary transcript, measured in the nucleus, matches the gene; the size of the cytoplasmic messenger matches the exons alone. ∎
19.3 The genetic code
Definition 19.4 (The genetic code)
The genetic code maps triplets of messenger bases, codons, onto amino acids. Of the 64 codons, 61 specify the twenty amino acids and three (UAA, UAG, UGA) are stop signals; AUG specifies methionine and is also the start codon, so every new polypeptide begins with methionine. The code is degenerate — most amino acids have several codons, differing mostly at the third position — non-overlapping, read without punctuation in a reading frame fixed by the start codon, and almost universal, the same in bacteria, plants and animals with minor variants in mitochondria: a human gene put into a bacterium is read correctly.
| first | second base | third | |||
| base | U | C | A | G | base |
| U | Phe Phe Leu Leu | Ser Ser Ser Ser | Tyr Tyr stop stop | Cys Cys stop Trp | U C A G |
| C | Leu Leu Leu Leu | Pro Pro Pro Pro | His His Gln Gln | Arg Arg Arg Arg | U C A G |
| A | Ile Ile Ile Met | Thr Thr Thr Thr | Asn Asn Lys Lys | Ser Ser Arg Arg | U C A G |
| G | Val Val Val Val | Ala Ala Ala Ala | Asp Asp Glu Glu | Gly Gly Gly Gly | U C A G |
Proposition 19.5 (How the code was read)
The code is a non-overlapping triplet code, and each codon’s meaning can be determined by chemistry.
Evidence. Crick and Brenner (1961) made mutations in a phage gene that added or removed one base: one such change destroyed the gene’s function, as did two, but three insertions (or three deletions) close together restored it — the message is read in threes, from a fixed starting point, and an insertion shifts the frame of everything downstream. Nirenberg and Matthaei (1961) added a synthetic RNA of uracils only to a cell-free extract of E. coli with the twenty amino acids: it made a chain of phenylalanine only, so UUU means Phe; other synthetic messengers, and then binding of single trinucleotides to ribosomes with their transfer RNAs, assigned all sixty-four by 1966. ∎
Example 19.6 (Reading a message)
The messenger -…GCAUGGCUUUCGGAUAA…- is read from the AUG: AUG GCU UUC GGA UAA — Met-Ala-Phe-Gly-stop: a peptide of four residues. Delete the first G after AUG and the frame shifts: AUG CUU UCG GAU AA… — Met-Leu-Ser-Asp…, a different protein of a different length. Change UUC to UUU and nothing changes: both are Phe, a silent mutation.
19.4 Translation
Definition 19.7 (Transfer RNA and its synthetases)
A transfer RNA (Chapter 11) is the adaptor between codon and amino acid: its anticodon, three bases in a loop, pairs antiparallel with a codon of the message, and its end carries the corresponding amino acid. The pairing at the third codon position is loose (wobble), so about forty tRNAs suffice for sixty-one codons. Each tRNA is loaded by its own aminoacyl-tRNA synthetase, an enzyme that recognises both the amino acid and the tRNA and joins them in two steps at the cost of one ATP (split to AMP: two high-energy bonds), proofreading the amino acid as it does so. These twenty enzymes are where the code is actually implemented: the ribosome does not check which amino acid a tRNA carries, only that its anticodon matches.
Definition 19.8 (The ribosome)
The ribosome is a particle of two subunits, each of ribosomal RNA and proteins (in bacteria: a small 30S subunit and a large 50S subunit, together 70S, ; eukaryotic ribosomes are larger, 80S). The small subunit binds the messenger and decodes it; the large subunit holds the tRNAs and forms the peptide bond — the catalyst is the ribosomal RNA itself. Three tRNA sites span both subunits: A (aminoacyl, where the next charged tRNA enters), P (peptidyl, holding the growing chain), E (exit).
Proposition 19.9 (The cycle of elongation)
Translation starts when the small subunit finds the start codon — in bacteria by pairing a ribosomal RNA sequence with a site just upstream of the AUG, in eukaryotes by binding the cap and scanning to the first AUG — and the initiator tRNA (methionine) settles in the P site; the large subunit then joins. Each round of elongation adds one residue: a charged tRNA whose anticodon matches the A-site codon is delivered by an elongation factor and checked (one GTP); the ribosomal RNA of the large subunit transfers the growing chain from the P-site tRNA onto the amino acid of the A-site tRNA, forming the peptide bond; the ribosome moves one codon along (a second GTP), shifting the tRNAs to P and E, and the empty one leaves. At a stop codon a release factor enters the A site and the chain is hydrolysed free. Bacterial ribosomes add fifteen to twenty residues a second, eukaryotic ones two to five; several ribosomes read one message at once, forming a polysome.
Method 19.10 (Reckoning a protein’s synthesis)
- Length: a protein of residues needs a coding sequence of nucleotides (the stop codon included), inside a messenger longer by its untranslated ends and, in the gene, by its introns.
- Time: divide by the elongation rate ( per second in bacteria, in eukaryotes); the message is being read by one ribosome every nucleotides, so the output per message is one chain every (spacing / rate) seconds.
- Energy: four high-energy phosphate bonds per residue (two to charge the tRNA, two GTP on the ribosome), plus the transcription of the message at two per nucleotide, shared among all the chains that message yields.
- Fidelity: about one wrong amino acid in codons; a protein of residues is wrong somewhere in one copy out of twenty — tolerable because proteins are replaceable and errors are not inherited.
19.5 After translation
Proposition 19.11 (What happens to a new polypeptide)
The chain folds as it emerges, helped by chaperones (Chapter 12). Its destination is written in its sequence: a signal peptide at the N-terminus binds a signal recognition particle that halts translation, docks the ribosome on the endoplasmic reticulum, and threads the chain into its lumen (Chapter 6); other sequences direct proteins into mitochondria, chloroplasts, the nucleus or peroxisomes after synthesis; a protein with no signal stays in the cytosol. Many proteins are modified: the initial methionine removed, sugars attached in the ER and Golgi, phosphates added and removed by kinases and phosphatases, lipids attached, pieces cut out (insulin is made as one chain and cut into two; zymogens are cut to activate them). And every protein is eventually degraded: tagged with the small protein ubiquitin and unfolded and digested by the proteasome, after a life of minutes (regulatory proteins) to months (haemoglobin) — so that the proteins present are those the cell is currently making.
Example 19.12 (The numbers of a cell)
A growing E. coli holds some ribosomes, each adding twenty residues a second: residues a second, a million proteins of residues in an hour — its own content, which is what doubling every hour requires. Half its energy goes to protein synthesis. A liver cell holds ten million ribosomes and makes proteins for weeks of use rather than for division; a plasma cell, secreting antibody, devotes nearly all of them to one protein and pours out two thousand molecules a second.
19.6 Exercises
Exercise 19.1 ★
Give the RNA transcribed from the template strand -TACGGATTC-, and the coding strand.
Exercise 19.2 ★
List the three modifications a eukaryotic primary transcript undergoes before it leaves the nucleus.
Solution
Solution of Exercise 19.2.
A cap (modified guanine), a poly-A tail, and the removal of the introns by splicing.
Exercise 19.3 ★
Using the code table, translate -AUGCCGAAAGUUUGA-.
Solution
Solution of Exercise 19.3.
AUG CCG AAA GUU UGA: Met-Pro-Lys-Val, then stop.
Exercise 19.4 ★
Name the three tRNA sites of the ribosome and what happens in each.
Solution
Solution of Exercise 19.4.
A: the incoming charged tRNA is admitted and checked. P: the tRNA carrying the growing chain; the bond forms between its chain and the A-site amino acid. E: the emptied tRNA on its way out.
Exercise 19.5 ★★
A messenger of nucleotides has of untranslated region and of untranslated region and poly-A. How many residues has the protein? How long does one ribosome take to make it at four residues a second, and how many ribosomes can read the message at once at one per nucleotides?
Solution
Solution of Exercise 19.5.
Coding nucleotides: 599 residues (600 codons, one a stop). Time . Ribosomes at once.
Exercise 19.6 ★★
Explain why three insertions restore a gene’s reading frame while one or two do not, and what the protein made from the triple insertion looks like.
Solution
Solution of Exercise 19.6.
The message is read in threes from a fixed start; one or two extra bases shift every codon downstream and the rest of the protein is gibberish, usually ending at a premature stop. Three extra bases add one codon and restore the frame: the protein has one extra residue and a few wrong ones between the insertions, and often still works.
Exercise 19.7 ★★
A mutation changes the codon CAG to UAG in the middle of a gene of codons; another changes CAG to CAA; a third changes it to CGG. Give the effect of each on the protein.
Solution
Solution of Exercise 19.7.
CAG (Gln) to UAG (stop): the chain ends at residue 150, a truncated, almost certainly inactive protein (nonsense mutation). CAG to CAA: still Gln, silent. CAG to CGG: Arg for Gln, a substitution (missense) whose effect depends on the position.
Exercise 19.8 ★★
Explain why the fidelity of translation rests on the aminoacyl-tRNA synthetases rather than on the ribosome, and describe an experiment that showed it (a cysteine attached to its tRNA is chemically converted to alanine; where does the alanine end up?).
Solution
Solution of Exercise 19.8.
The ribosome checks only the codon–anticodon pairing; it cannot see the amino acid. If cysteine on its tRNA is converted chemically to alanine, the ribosome inserts alanine wherever the message says cysteine: the tRNA, not the amino acid, is read. Hence the synthetase, which pairs each amino acid with the right tRNA, is the true translator.
Exercise 19.9 ★★
Compute the energy cost, in high-energy bonds, of a protein of residues, and the fraction of it spent on the ribosome.
Solution
Solution of Exercise 19.9.
bonds; the ribosome’s two GTP per residue are half of it, the synthetases’ two the other half.
Exercise 19.10 ★★★
In bacteria ribosomes begin translating a message while it is still being transcribed; in eukaryotes they cannot. Explain why (two reasons), and what this difference makes possible in eukaryotes.
Solution
Solution of Exercise 19.10.
In eukaryotes the transcript is made in the nucleus and the ribosomes are in the cytosol, separated by the envelope; and the transcript is not a messenger until it has been spliced and capped. The separation makes possible the processing itself — alternative splicing, the control of export, and a check that the message is complete before it is read.
Exercise 19.11 ★★★
A gene of five exons can be spliced to include or skip exon 3 ( nucleotides) and exon 4 ( nucleotides). List the possible messengers and say which ones keep the reading frame of exon 5. What does this show about alternative splicing?
Solution
Solution of Exercise 19.11.
Four messengers: with both exons (190 nucleotides added), with 3 only (90), with 4 only (100), with neither (0). A frame is kept if the added length is a multiple of 3: 0 and 90 keep it; 100 and 190 shift it, so the messengers with exon 4 alone or both exons read exon 5 out of frame and truncate the protein. Alternative splicing must respect the frame, and exons are often multiples of three for that reason.
Exercise 19.12 ★★★
“The code is a frozen accident: arbitrary, but too costly to change.” Discuss in a paragraph: what in the code looks arbitrary, what looks optimised (the third position, similar codons for similar amino acids), and why a mutation changing a synthetase’s specificity is almost always lethal.
Solution
Solution of Exercise 19.12.
Arbitrary: nothing in chemistry says that UUU must mean Phe; the assignments are conventions held by the synthetases. Optimised: the third position is the most degenerate, so the errors and mutations that fall there are often silent, and codons that differ by one base tend to encode similar amino acids, so a substitution is often mild — the code minimises the damage of error. Frozen: a synthetase that changed its specificity would alter every protein of the cell at once, thousands of them, in the same instant; almost no cell survives that, so the code cannot drift and has been fixed since the common ancestor of all living things.
19.7 Problem: From a Gene to a Protein
Problem 19.1
Weekend problem — a thirty-kilobase gene followed into a five-hundred-residue protein: transcription timed, splicing weighed, ribosomes counted, energy and errors reckoned, ending on the cost of one protein in ATP
A human gene spans with eight exons totalling nucleotides of mature messenger, of which code (500 residues plus the stop). Polymerase II transcribes at nucleotides per second; ribosomes elongate at residues per second and space themselves one per nucleotides; the messenger’s half-life is . Costs: two high-energy bonds per nucleotide transcribed, four per residue translated; take one high-energy bond as one ATP. Error rates: per nucleotide transcribed, per codon translated.
Part I — Transcription and splicing.
- How long does one polymerase take to transcribe the gene?
- What fraction of the primary transcript is removed by splicing?
- How many nucleotides are transcribed for every nucleotide of mature messenger?
- Compute the ATP spent transcribing one primary transcript.
- If a polymerase starts every , how many polymerases are on the gene at once, and how many messengers does the gene produce per hour?
- Compute the number of introns and their mean length.
- The spliceosome removes each intron in about a minute, in parallel. Does splicing or transcription set the time from gene to messenger?
- Compute the physical length of the gene and of the mature messenger ( per nucleotide).
Part II — Translation.
- How long does one ribosome take to translate the protein?
- How many ribosomes read one messenger at once?
- How many protein molecules does one messenger yield per hour?
- Over its half-life, how many does it yield in all (a message with half-life yields, on average, hours of full production)?
- Compute the ATP spent translating one protein.
- Compute the ATP spent on transcription per protein, sharing the transcript’s cost among the proteins of question 12.
- Compute the total cost of one protein and the fraction due to translation.
Part III — Errors.
- Compute the probability that a given messenger carries at least one transcription error in its coding nucleotides.
- Compute the probability that a given protein molecule carries at least one translation error.
- About a quarter of nucleotide changes are silent and a third of amino-acid substitutions are harmless. What fraction of the protein molecules made are defective?
- A defective messenger yields defective proteins for two hours; a defective protein is one molecule. Explain why the cell can afford a translation error rate ten times the transcription rate, and both far above the replication rate.
Part IV — The cell’s budget. The cell makes copies of this protein per hour, and in all residues of protein per day.
- How many messengers of this gene must be present at once (each yielding the number of question 11)?
- How many ribosomes are occupied by this protein?
- Compute the daily ATP the cell spends on all its protein synthesis (four per residue) and, at , the power in watts for a cell of .
- Compare with the cell’s total power if it consumes ATP per second.
- A drug blocks the spliceosome. Predict its effect on this protein and on a bacterial protein.
- State the result: the ATP cost of one molecule of the protein, split between translation and transcription, and the time from the start of transcription to the first finished protein.
Solution
Solution of Problem 19.1.
1. , about . 2. . 3. . 4. ATP. 5. polymerases on the gene; 360 messengers per hour. 6. Seven introns, mean nucleotides. 7. Transcription (); the introns are spliced as they are made, and the last one adds only a minute. 8. Gene ; messenger . 9. . 10. . 11. One chain finishes every : 180 per hour. 12. of full production: about 520 proteins. 13. ATP. 14. ATP per protein. 15. About 2100 ATP, of it in translation. 16. . 17. . 18. Transcription: of messengers defective, hence of proteins; translation: ; about of the molecules. 19. A protein error is confined to one molecule, which is soon degraded; a messenger error is copied into hundreds of proteins; a replication error is inherited by every descendant for ever. The cost of an error, and hence the accuracy worth paying for, rises at each step back. 20. messengers present. 21. ribosomes. 22. ATP per day , per day, i.e. — per cubic micrometre. 23. ATP per second is per day: protein synthesis is a hundredth of it in this slowly renewing cell (in a dividing bacterium it is half). 24. No mature messenger: the primary transcripts accumulate in the nucleus and the protein disappears as its messengers decay, within hours. The bacterial protein, whose gene has no introns, is unaffected. 25. About 2100 ATP per molecule — 2000 for translation, a hundred for its share of the transcript; first protein after of transcription, a minute of processing and export, and of translation: about twenty minutes.