High School Biology · Grades 10–12
14From Gene to Protein
In 1961 a biochemist fed a cell-free extract of bacteria with an artificial RNA made of nothing but the letter U, repeated. The extract made a protein made of nothing but one amino acid, phenylalanine, repeated. With that single experiment the first word of the genetic code was read: UUU means phenylalanine. Within five years all sixty-four words were known, and they turned out to be the same words in a bacterium, a wheat plant and a human. This chapter follows the path from a gene’s sequence of nucleotides to a protein’s sequence of amino acids — the path every trait of Chapter 13 runs along.
14.1 Genes make proteins
Proposition 14.1 (One gene, one protein)
A gene is expressed when the cell uses its sequence to make the protein it encodes. Most traits depend on proteins — enzymes, structural fibres, carriers, receptors — and a mutation changes a trait by changing the protein its gene produces.
Evidence. Beadle and Tatum (1941) irradiated spores of a bread mould and collected mutants that could no longer grow on minimal medium unless one specific substance, such as the amino acid arginine, was added. Each mutant lacked one enzyme of the chain of reactions that makes that substance; each defect was inherited as a single gene; and mutants blocked at different steps of the same chain carried mutations in different genes. One gene, one enzyme — and, as later work generalised, one gene, one polypeptide chain. ∎
Definition 14.2 (Protein, amino acid sequence)
A protein is a chain of amino acids, from a few dozen to several thousand long, drawn from a set of twenty kinds. The sequence — the order of the amino acids — is what the gene specifies; once made, the chain folds into a definite three-dimensional shape determined by that sequence, and the shape determines the function. Change one amino acid and the fold, and the function, may change.
14.2 Transcription: the gene copied into RNA
Definition 14.3 (Messenger RNA and transcription)
RNA (ribonucleic acid) is a single-stranded chain of nucleotides like those of DNA, except that its sugar is ribose and the base uracil (U) replaces thymine (T). Transcription is the copying of a gene’s sequence into a molecule of messenger RNA (mRNA): the enzyme RNA polymerase opens the double helix along the gene, reads one strand (the template strand) and assembles the complementary RNA, U facing A, A facing T, G facing C and C facing G. The mRNA has the sequence of the gene’s other strand, with U for T; it leaves the nucleus for the cytoplasm.
Example 14.4 (A transcript)
Template strand 3’-TAC GGA CTT ATC-5’ gives the mRNA 5’-AUG CCU GAA UAG-3’. The gene’s other strand reads 5’-ATG CCT GAA TAG-3’: the mRNA is that strand’s sequence with U for T, which is why gene sequences are conventionally written as that strand, the coding strand.
14.3 The genetic code
Definition 14.5 (Codon and the genetic code)
The mRNA is read in consecutive groups of three nucleotides, the codons, from a defined starting point. The genetic code is the correspondence between the 64 possible codons and the 20 amino acids: 61 codons specify an amino acid, three (UAA, UAG, UGA) are stop signals ending the chain, and AUG, which specifies methionine, also serves as the start codon. The code is degenerate — most amino acids have several codons — and universal: the same table serves every known organism, with rare minor variants.
| second letter | |||||
| first | U | C | A | G | third |
| U | UUU Phe | UCU Ser | UAU Tyr | UGU Cys | U |
| UUC Phe | UCC Ser | UAC Tyr | UGC Cys | C | |
| UUA Leu | UCA Ser | UAA stop | UGA stop | A | |
| UUG Leu | UCG Ser | UAG stop | UGG Trp | G | |
| C | CUU Leu | CCU Pro | CAU His | CGU Arg | U |
| CUC Leu | CCC Pro | CAC His | CGC Arg | C | |
| CUA Leu | CCA Pro | CAA Gln | CGA Arg | A | |
| CUG Leu | CCG Pro | CAG Gln | CGG Arg | G | |
| A | AUU Ile | ACU Thr | AAU Asn | AGU Ser | U |
| AUC Ile | ACC Thr | AAC Asn | AGC Ser | C | |
| AUA Ile | ACA Thr | AAA Lys | AGA Arg | A | |
| AUG Met | ACG Thr | AAG Lys | AGG Arg | G | |
| G | GUU Val | GCU Ala | GAU Asp | GGU Gly | U |
| GUC Val | GCC Ala | GAC Asp | GGC Gly | C | |
| GUA Val | GCA Ala | GAA Glu | GGA Gly | A | |
| GUG Val | GCG Ala | GAG Glu | GGG Gly | G | |
Proposition 14.6 (How the code was read)
Every codon’s meaning was established experimentally.
Evidence. Nirenberg and Matthaei (1961) added synthetic RNA of a single repeated nucleotide to a cell-free extract containing ribosomes, amino acids and the enzymes of protein synthesis: poly-U yielded a chain of phenylalanines, poly-C a chain of prolines, poly-A of lysines. Synthetic RNAs of repeating pairs and triplets, and later the binding of single triplets to ribosomes, assigned the remaining codons by 1966. The three-letter length had been inferred from the arithmetic (two letters give only 16 combinations, too few for 20 amino acids; three give 64) and confirmed by mutations: inserting one or two nucleotides into a gene destroys its protein, inserting three restores a nearly normal one. ∎
Method 14.7 (Translating a sequence)
- Write the coding strand (or transcribe the template strand): replace T by U to get the mRNA.
- Find the first AUG: it is the start, and the reading frame is fixed from it.
- Cut the mRNA into consecutive triplets from the AUG and read each in the table.
- Stop at the first stop codon; the amino acids listed, in order, are the protein’s sequence.
- For a mutant, repeat and compare: same protein (silent), one amino acid changed (missense), premature stop (nonsense), everything changed after the mutation (frameshift).
Example 14.8 (Four mutations, four outcomes)
Coding strand ATG CCT GAA TTC TAG: mRNA AUG CCU GAA UUC UAG, protein Met–Pro–Glu–Phe.
ATG CCC GAA TTC TAG: CCC is still Pro — silent.ATG CCT GCA TTC TAG: GCA is Ala — Met–Pro–Ala–Phe, missense.ATG CCT TAA TTC TAG: UAA is stop — Met–Pro, nonsense, a truncated protein.ATG CCAT GAA TTC TAG: an inserted A shifts the frame:AUG CCA UGA… Met–Pro–stop — frameshift.
14.4 Translation: the message read
Proposition 14.9 (Translation)
Translation takes place on the ribosomes, in the cytoplasm. A ribosome binds the mRNA at its start codon and moves along it codon by codon. Each codon is matched by a transfer RNA (tRNA), a small RNA carrying, at one end, the three-nucleotide anticodon complementary to the codon and, at the other, the corresponding amino acid. The ribosome joins each amino acid brought in to the growing chain and releases the empty tRNA; at a stop codon it releases the finished chain, which folds. Several ribosomes read one mRNA in succession, and one mRNA yields hundreds of copies of the protein before it is degraded.
Proof. Admitted at this level. ∎
Example 14.10 (Speed and yield)
A ribosome adds 5 to 20 amino acids per second: a protein of 300 amino acids takes under a minute. Ribosomes follow each other on an mRNA about 80 nucleotides apart, so a message of 900 nucleotides carries a dozen at once; over its life of a few hours a single mRNA produces a thousand protein molecules. A cell of E. coli makes some protein molecules per second; a red blood cell’s precursor, some haemoglobin molecules in its lifetime.
14.5 One gene, several messages
Proposition 14.11 (Maturation of the message in eukaryotes)
In eukaryotic cells the sequence of a gene is interrupted by introns, segments that are transcribed but then cut out of the RNA before it leaves the nucleus; the segments kept and joined end to end, the exons, form the mature mRNA. Many genes can be cut in more than one way, keeping different sets of exons (alternative splicing), so that one gene yields several distinct mRNAs and several related proteins: the human genes encode well over proteins.
Proof. Admitted at this level. ∎
Remark 14.12 (The flow of information)
DNA is transcribed into RNA, RNA is translated into protein, and protein makes the trait: the information flows one way. A changed protein never rewrites the gene, which is why a lifetime of training or sunburn is not inherited, and why only mutations of the DNA (Chapter 13) change what the next generation receives. The same three steps run in every cell of every organism, in the same code: the universality of Chapter 3, now down to the last word.
14.6 Exercises
Exercise 14.1 ★
Exercise 14.2 ★
Transcribe the template strand 3’-TAC CGA AAA ATT-5’ into mRNA.
Solution
Solution of Exercise 14.2.
5’-AUG GCU UUU UAA-3’.
Exercise 14.3 ★
Translate the mRNA AUG GGC AAA UGU UAA with the code table.
Solution
Solution of Exercise 14.3.
Met–Gly–Lys–Cys, then stop.
Exercise 14.4 ★
What are the roles of the ribosome and of the transfer RNAs in translation?
Solution
Solution of Exercise 14.4.
The ribosome reads the mRNA codon by codon and joins the amino acids into a chain; each transfer RNA matches one codon with its anticodon and brings the corresponding amino acid.
Exercise 14.5 ★
Why must the code use at least three nucleotides per amino acid?
Solution
Solution of Exercise 14.5.
With 4 letters, words of two give combinations, fewer than 20 amino acids; words of three give 64, enough.
Exercise 14.6 ★★
A protein has 412 amino acids. What is the minimum length of its mRNA’s coding part, stop codon included, and of the corresponding gene in base pairs?
Solution
Solution of Exercise 14.6.
nucleotides; a gene of at least 1239 base pairs (more in a eukaryote, with introns).
Exercise 14.7 ★★
The coding strand ATG AAA GGC TGG TAA is mutated to ATG AAA GGC TGA TAA. Give both proteins and classify the mutation.
Exercise 14.8 ★★
Same starting sequence, mutated to ATG AAG GGC TGG TAA and to ATG AAA GGG CTG GTA A. Classify each and give the proteins.
Solution
Solution of Exercise 14.8.
AAG is still Lys: silent, Met–Lys–Gly–Trp unchanged. The second is an insertion of G shifting the frame: AUG AAA GGG CUG GUA A… Met–Lys–Gly–Leu–Val…, a frameshift with no stop in the fragment.
Exercise 14.9 ★★
Explain why a substitution in the third position of a codon is often silent, using the table.
Solution
Solution of Exercise 14.9.
In most blocks of the table the four codons sharing the first two letters specify the same amino acid (Pro, Thr, Ala, Gly, Val, Ser, Leu…): changing the third letter then changes nothing. The degeneracy of the code sits mainly in the third position.
Exercise 14.10 ★★
Poly-UC RNA (UCUCUCUC…) gives a protein alternating serine and leucine. Show that this is consistent with a three-letter code and use it to assign two codons.
Solution
Solution of Exercise 14.10.
Read in threes, UCU CUC UCU CUC… alternates two codons, UCU and CUC, so the protein alternates two amino acids — as observed. With a two-letter code the same repeat would give a single codon and a single amino acid. Since poly-U gives Phe and the table’s UCU is Ser and CUC Leu, the experiment assigns UCU to Ser and CUC to Leu (or the reverse, settled by other experiments).
Exercise 14.11 ★★
A human gene of base pairs produces a protein of 500 amino acids. Explain the discrepancy.
Exercise 14.12 ★★★
A drug blocks bacterial ribosomes but not human ones. Explain why it can be an antibiotic, and what its existence implies about ribosomes across the two groups.
Solution
Solution of Exercise 14.12.
Blocking translation kills the bacterium while leaving the patient’s cells working. The drug’s selectivity implies that bacterial and human ribosomes, though they do the same job with the same code, differ in structure enough for a molecule to bind one and not the other — a difference accumulated since their separation.
Exercise 14.13 ★★★
Explain how the jellyfish gene of Chapter 3 could be read by a mouse, in the vocabulary of this chapter, and what would happen if the mouse used a different code.
Solution
Solution of Exercise 14.13.
The mouse’s RNA polymerase transcribed the jellyfish gene, and its ribosomes and tRNAs translated the mRNA with the same codon table, so the same amino acid sequence, hence the same fluorescent protein, was made. With a different code the codons would be read as other amino acids and the protein would be a meaningless, non-fluorescent chain.
Exercise 14.14 ★★★
A mutation changes the tRNA whose anticodon reads UAG so that it carries an amino acid instead of stopping. Predict its effect on the cell’s proteins in general, and on a mutant gene carrying a premature UAG.
Solution
Solution of Exercise 14.14.
UAG would no longer stop translation: every protein whose gene ends in UAG would be extended past its normal end, often losing function — a widespread harm. But a gene truncated by a premature UAG would now be read through, restoring a nearly full-length protein: the second mutation suppresses the first.
Exercise 14.15 ★★★
Discuss why a frameshift near the start of a gene is almost always worse than one near its end, and why a missense mutation can be anything from harmless to fatal.
Solution
Solution of Exercise 14.15.
A frameshift scrambles everything downstream and usually meets a stop within a few codons: near the start, essentially no correct protein is made; near the end, most of the protein is intact and may still fold and work. A missense changes one amino acid: harmless if that position tolerates the change, fatal if it is at the active site or breaks the fold — the effect depends on which amino acid, where.
14.7 Problem: A Gene Read Four Ways
Problem 14.1
Weekend problem — one short gene transcribed, translated, mutated three times and timed: from the DNA to the amino acids, and the sixty seconds a protein takes
The coding strand of a short gene reads:
ATG GCT TGG AAA CCC GAA TAC GGT CAT TTC TGA
Part I — Reading the gene.
- Write the template strand.
- Write the mRNA transcribed from the gene.
- Translate it, codon by codon, and give the protein’s sequence of amino acids.
- How many amino acids does the protein contain, and how many nucleotides of the mRNA are used to specify them, stop included?
- Which codon started the reading, and what fixes the reading frame for all the others?
Part II — Three mutants.
- Mutant 1:
ATG GCT TGG AAA CCG GAA TAC GGT CAT TTC TGA. Give the protein and classify the mutation. - Mutant 2:
ATG GCT TGG AAA CCC GAA TAC GGT CAT TTC TGAwith the ninth codonCATreplaced byCGT. Give the protein and classify. - Mutant 3:
ATG GCT TGA AAA CCC GAA TAC GGT CAT TTC TGA. Give the protein and classify. Which single substitution produced it? - Mutant 4: a T is inserted after the sixth nucleotide:
ATG GCT TTG GAA ACC CGA ATA CGG TCA TTT CTG A. Give the protein and classify. - Rank the four mutants from least to most disruptive for the protein and justify.
Part III — Why the code is what it is.
- How many codons specify leucine? Serine? Tryptophan? Which amino acid is most likely to be changed by a random substitution in its codon, and which least?
- Compute the fraction of all substitutions at the third position of a leucine codon
CUxthat are silent. - Explain why a code of two letters could not work, and why a code of four is not needed.
- Poly-U gives poly-phenylalanine and poly-A poly-lysine. Which entries of the table do these two experiments establish?
- The code is the same in the bacterium and in a human. State the consequence for transgenesis and the consequence for kinship.
Part IV — Timing the factory. A ribosome adds 10 amino acids per second and ribosomes follow each other 80 nucleotides apart along an mRNA; each mRNA lives 2 hours.
- How long does one ribosome take to translate this gene’s protein? And a protein of 600 amino acids?
- How many ribosomes can read an mRNA of 1800 nucleotides at once?
- If a new ribosome starts every 8 seconds, how many copies of the protein does one mRNA yield in its lifetime?
- A bacterium needs copies of an enzyme in 10 minutes. How many mRNAs of the gene, transcribed at once, are needed?
- State the result: the protein encoded by the gene, the mutation among the four that abolished it, and the time one ribosome needs to build it.
Solution
Solution of Problem 14.1.
1. 3’-TAC CGA ACC TTT GGG CTT ATG CCA GTA AAG ACT-5’.
2. AUG GCU UGG AAA CCC GAA UAC GGU CAU UUC UGA.
3. Met–Ala–Trp–Lys–Pro–Glu–Tyr–Gly–His–Phe, stop.
4. 10 amino acids; 33 nucleotides (11 codons including the stop).
5. AUG; the position of that first AUG fixes the frame, and every following codon is read in step from it.
6. CCG is still Pro: silent, protein unchanged.
7. CGU is Arg in place of His: missense, Met–Ala–Trp–Lys–Pro–Glu–Tyr–Gly–Arg–Phe.
8. UGA is stop: Met–Ala, nonsense. A single substitution G to A in the third codon (TGG to TGA).
9. mRNA AUG GCU UUG GAA ACC CGA AUA CGG UCA UUU CUG A: Met–Ala–Leu–Glu–Thr–Arg–Ile–Arg–Ser–Phe–Leu…, no stop in the fragment: frameshift, every amino acid after the second is changed.
10. Mutant 1 (silent, no change) mutant 2 (one amino acid changed, function possibly kept) mutant 4 (frameshift, all but two amino acids wrong) mutant 3 (nonsense, a two-amino-acid fragment: the protein is gone).
11. Leu 6, Ser 6, Trp 1. Trp, with a single codon, is changed by any substitution; Leu or Ser, with six, is most often unchanged.
12. CUU, CUC, CUA, CUG are all Leu: every one of the 3 possible substitutions at the third position is silent — 100%.
13. Two letters give 16 codons, too few for 20 amino acids plus a stop; three give 64, already more than enough, so four (256) would be waste.
14. UUU Phe and AAA Lys.
15. A gene transferred between species is read identically, so transgenesis works; and a shared code, inherited unchanged, is evidence that all organisms descend from a common ancestor that already used it.
16. 10 amino acids: . 600 amino acids: .
17. ribosomes at once.
18. copies.
19. One mRNA gives copies in 10 minutes; about 27 mRNAs are needed.
20. The protein Met–Ala–Trp–Lys–Pro–Glu–Tyr–Gly–His–Phe; mutant 3, the single G-to-A substitution creating a stop at the third codon, abolished it; one ribosome builds the ten amino acids in about one second.