Biology · Book 2 · Grades 10–12

High School Biology

High School Biology · Grades 10–12

14From Gene to Protein

In 1961 a biochemist fed a cell-free extract of bacteria with an artificial RNA made of nothing but the letter U, repeated. The extract made a protein made of nothing but one amino acid, phenylalanine, repeated. With that single experiment the first word of the genetic code was read: UUU means phenylalanine. Within five years all sixty-four words were known, and they turned out to be the same words in a bacterium, a wheat plant and a human. This chapter follows the path from a gene’s sequence of nucleotides to a protein’s sequence of amino acids — the path every trait of Chapter 13 runs along.

14.1 Genes make proteins

Proposition 14.1 (One gene, one protein)

A gene is expressed when the cell uses its sequence to make the protein it encodes. Most traits depend on proteinsenzymes, structural fibres, carriers, receptors — and a mutation changes a trait by changing the protein its gene produces.

Evidence. Beadle and Tatum (1941) irradiated spores of a bread mould and collected mutants that could no longer grow on minimal medium unless one specific substance, such as the amino acid arginine, was added. Each mutant lacked one enzyme of the chain of reactions that makes that substance; each defect was inherited as a single gene; and mutants blocked at different steps of the same chain carried mutations in different genes. One gene, one enzyme — and, as later work generalised, one gene, one polypeptide chain.

Definition 14.2 (Protein, amino acid sequence)

A protein is a chain of amino acids, from a few dozen to several thousand long, drawn from a set of twenty kinds. The sequence — the order of the amino acids — is what the gene specifies; once made, the chain folds into a definite three-dimensional shape determined by that sequence, and the shape determines the function. Change one amino acid and the fold, and the function, may change.

14.2 Transcription: the gene copied into RNA

Definition 14.3 (Messenger RNA and transcription)

RNA (ribonucleic acid) is a single-stranded chain of nucleotides like those of DNA, except that its sugar is ribose and the base uracil (U) replaces thymine (T). Transcription is the copying of a gene’s sequence into a molecule of messenger RNA (mRNA): the enzyme RNA polymerase opens the double helix along the gene, reads one strand (the template strand) and assembles the complementary RNA, U facing A, A facing T, G facing C and C facing G. The mRNA has the sequence of the gene’s other strand, with U for T; it leaves the nucleus for the cytoplasm.

Transcription. RNA polymerase opens the helix, reads the template strand and builds a messenger RNA complementary to it (U in place of T). Behind the enzyme the helix closes again; the RNA peels off.
Transcription. RNA polymerase opens the helix, reads the template strand and builds a messenger RNA complementary to it (U in place of T). Behind the enzyme the helix closes again; the RNA peels off.

Example 14.4 (A transcript)

Template strand 3’-TAC GGA CTT ATC-5’ gives the mRNA 5’-AUG CCU GAA UAG-3’. The gene’s other strand reads 5’-ATG CCT GAA TAG-3’: the mRNA is that strand’s sequence with U for T, which is why gene sequences are conventionally written as that strand, the coding strand.

14.3 The genetic code

Definition 14.5 (Codon and the genetic code)

The mRNA is read in consecutive groups of three nucleotides, the codons, from a defined starting point. The genetic code is the correspondence between the 64 possible codons and the 20 amino acids: 61 codons specify an amino acid, three (UAA, UAG, UGA) are stop signals ending the chain, and AUG, which specifies methionine, also serves as the start codon. The code is degenerate — most amino acids have several codons — and universal: the same table serves every known organism, with rare minor variants.

second letter
firstUCAGthird
UUUU PheUCU SerUAU TyrUGU CysU
UUC PheUCC SerUAC TyrUGC CysC
UUA LeuUCA SerUAA stopUGA stopA
UUG LeuUCG SerUAG stopUGG TrpG
CCUU LeuCCU ProCAU HisCGU ArgU
CUC LeuCCC ProCAC HisCGC ArgC
CUA LeuCCA ProCAA GlnCGA ArgA
CUG LeuCCG ProCAG GlnCGG ArgG
AAUU IleACU ThrAAU AsnAGU SerU
AUC IleACC ThrAAC AsnAGC SerC
AUA IleACA ThrAAA LysAGA ArgA
AUG MetACG ThrAAG LysAGG ArgG
GGUU ValGCU AlaGAU AspGGU GlyU
GUC ValGCC AlaGAC AspGGC GlyC
GUA ValGCA AlaGAA GluGGA GlyA
GUG ValGCG AlaGAG GluGGG GlyG
The genetic code, read on the mRNA. Each codon is found by its first letter (row), second letter (column) and third letter (line within the block). AUG is both methionine and the start signal; three codons are stops.

Proposition 14.6 (How the code was read)

Every codon’s meaning was established experimentally.

Evidence. Nirenberg and Matthaei (1961) added synthetic RNA of a single repeated nucleotide to a cell-free extract containing ribosomes, amino acids and the enzymes of protein synthesis: poly-U yielded a chain of phenylalanines, poly-C a chain of prolines, poly-A of lysines. Synthetic RNAs of repeating pairs and triplets, and later the binding of single triplets to ribosomes, assigned the remaining codons by 1966. The three-letter length had been inferred from the arithmetic (two letters give only 16 combinations, too few for 20 amino acids; three give 64) and confirmed by mutations: inserting one or two nucleotides into a gene destroys its protein, inserting three restores a nearly normal one.

Method 14.7 (Translating a sequence)

  1. Write the coding strand (or transcribe the template strand): replace T by U to get the mRNA.
  2. Find the first AUG: it is the start, and the reading frame is fixed from it.
  3. Cut the mRNA into consecutive triplets from the AUG and read each in the table.
  4. Stop at the first stop codon; the amino acids listed, in order, are the protein’s sequence.
  5. For a mutant, repeat and compare: same protein (silent), one amino acid changed (missense), premature stop (nonsense), everything changed after the mutation (frameshift).

Example 14.8 (Four mutations, four outcomes)

Coding strand ATG CCT GAA TTC TAG: mRNA AUG CCU GAA UUC UAG, protein Met–Pro–Glu–Phe.

  • ATG CCC GAA TTC TAG: CCC is still Pro — silent.
  • ATG CCT GCA TTC TAG: GCA is Ala — Met–Pro–Ala–Phe, missense.
  • ATG CCT TAA TTC TAG: UAA is stop — Met–Pro, nonsense, a truncated protein.
  • ATG CCAT GAA TTC TAG: an inserted A shifts the frame: AUG CCA UGA… Met–Pro–stop — frameshift.

14.4 Translation: the message read

Proposition 14.9 (Translation)

Translation takes place on the ribosomes, in the cytoplasm. A ribosome binds the mRNA at its start codon and moves along it codon by codon. Each codon is matched by a transfer RNA (tRNA), a small RNA carrying, at one end, the three-nucleotide anticodon complementary to the codon and, at the other, the corresponding amino acid. The ribosome joins each amino acid brought in to the growing chain and releases the empty tRNA; at a stop codon it releases the finished chain, which folds. Several ribosomes read one mRNA in succession, and one mRNA yields hundreds of copies of the protein before it is degraded.

Proof. Admitted at this level.

Translation. The ribosome holds two codons; a tRNA whose anticodon matches each codon brings its amino acid, the chain is transferred onto the newcomer, and the ribosome steps one codon along. At a stop codon the chain is released.
Translation. The ribosome holds two codons; a tRNA whose anticodon matches each codon brings its amino acid, the chain is transferred onto the newcomer, and the ribosome steps one codon along. At a stop codon the chain is released.

Example 14.10 (Speed and yield)

A ribosome adds 5 to 20 amino acids per second: a protein of 300 amino acids takes under a minute. Ribosomes follow each other on an mRNA about 80 nucleotides apart, so a message of 900 nucleotides carries a dozen at once; over its life of a few hours a single mRNA produces a thousand protein molecules. A cell of E. coli makes some 2000020\,000 protein molecules per second; a red blood cell’s precursor, some 5×1085 \times 10^{8} haemoglobin molecules in its lifetime.

14.5 One gene, several messages

Proposition 14.11 (Maturation of the message in eukaryotes)

In eukaryotic cells the sequence of a gene is interrupted by introns, segments that are transcribed but then cut out of the RNA before it leaves the nucleus; the segments kept and joined end to end, the exons, form the mature mRNA. Many genes can be cut in more than one way, keeping different sets of exons (alternative splicing), so that one gene yields several distinct mRNAs and several related proteins: the 2000020\,000 human genes encode well over 100000100\,000 proteins.

Proof. Admitted at this level.

A eukaryotic gene is transcribed whole, then the introns are removed. Keeping all the exons, or leaving one out, gives two mRNAs and two proteins from the same gene.
A eukaryotic gene is transcribed whole, then the introns are removed. Keeping all the exons, or leaving one out, gives two mRNAs and two proteins from the same gene.

Remark 14.12 (The flow of information)

DNA is transcribed into RNA, RNA is translated into protein, and protein makes the trait: the information flows one way. A changed protein never rewrites the gene, which is why a lifetime of training or sunburn is not inherited, and why only mutations of the DNA (Chapter 13) change what the next generation receives. The same three steps run in every cell of every organism, in the same code: the universality of Chapter 3, now down to the last word.

14.6 Exercises

Exercise 14.1

Give three differences between DNA and RNA.

Solution

Solution of Exercise 14.1.

RNA is single-stranded, its sugar is ribose, and it uses uracil in place of thymine (it is also short-lived and leaves the nucleus).

Exercise 14.2

Transcribe the template strand 3’-TAC CGA AAA ATT-5’ into mRNA.

Solution

Solution of Exercise 14.2.

5’-AUG GCU UUU UAA-3’.

Exercise 14.3

Translate the mRNA AUG GGC AAA UGU UAA with the code table.

Solution

Solution of Exercise 14.3.

Met–Gly–Lys–Cys, then stop.

Exercise 14.4

What are the roles of the ribosome and of the transfer RNAs in translation?

Solution

Solution of Exercise 14.4.

The ribosome reads the mRNA codon by codon and joins the amino acids into a chain; each transfer RNA matches one codon with its anticodon and brings the corresponding amino acid.

Exercise 14.5

Why must the code use at least three nucleotides per amino acid?

Solution

Solution of Exercise 14.5.

With 4 letters, words of two give 42=164^2 = 16 combinations, fewer than 20 amino acids; words of three give 64, enough.

Exercise 14.6 ★★

A protein has 412 amino acids. What is the minimum length of its mRNA’s coding part, stop codon included, and of the corresponding gene in base pairs?

Solution

Solution of Exercise 14.6.

412×3+3=1239412 \times 3 + 3 = 1239 nucleotides; a gene of at least 1239 base pairs (more in a eukaryote, with introns).

Exercise 14.7 ★★

The coding strand ATG AAA GGC TGG TAA is mutated to ATG AAA GGC TGA TAA. Give both proteins and classify the mutation.

Solution

Solution of Exercise 14.7.

Original: Met–Lys–Gly–Trp. Mutant: UGA is a stop, so Met–Lys–Gly: a nonsense mutation, truncating the protein.

Exercise 14.8 ★★

Same starting sequence, mutated to ATG AAG GGC TGG TAA and to ATG AAA GGG CTG GTA A. Classify each and give the proteins.

Solution

Solution of Exercise 14.8.

AAG is still Lys: silent, Met–Lys–Gly–Trp unchanged. The second is an insertion of G shifting the frame: AUG AAA GGG CUG GUA A… Met–Lys–Gly–Leu–Val…, a frameshift with no stop in the fragment.

Exercise 14.9 ★★

Explain why a substitution in the third position of a codon is often silent, using the table.

Solution

Solution of Exercise 14.9.

In most blocks of the table the four codons sharing the first two letters specify the same amino acid (Pro, Thr, Ala, Gly, Val, Ser, Leu…): changing the third letter then changes nothing. The degeneracy of the code sits mainly in the third position.

Exercise 14.10 ★★

Poly-UC RNA (UCUCUCUC…) gives a protein alternating serine and leucine. Show that this is consistent with a three-letter code and use it to assign two codons.

Solution

Solution of Exercise 14.10.

Read in threes, UCU CUC UCU CUC… alternates two codons, UCU and CUC, so the protein alternates two amino acids — as observed. With a two-letter code the same repeat would give a single codon and a single amino acid. Since poly-U gives Phe and the table’s UCU is Ser and CUC Leu, the experiment assigns UCU to Ser and CUC to Leu (or the reverse, settled by other experiments).

Exercise 14.11 ★★

A human gene of 3000030\,000 base pairs produces a protein of 500 amino acids. Explain the discrepancy.

Solution

Solution of Exercise 14.11.

Only 1503 base pairs are needed for the coding sequence; the rest are introns, transcribed and then cut out of the pre-mRNA, plus regulatory sequences at the gene’s ends.

Exercise 14.12 ★★★

A drug blocks bacterial ribosomes but not human ones. Explain why it can be an antibiotic, and what its existence implies about ribosomes across the two groups.

Solution

Solution of Exercise 14.12.

Blocking translation kills the bacterium while leaving the patient’s cells working. The drug’s selectivity implies that bacterial and human ribosomes, though they do the same job with the same code, differ in structure enough for a molecule to bind one and not the other — a difference accumulated since their separation.

Exercise 14.13 ★★★

Explain how the jellyfish gene of Chapter 3 could be read by a mouse, in the vocabulary of this chapter, and what would happen if the mouse used a different code.

Solution

Solution of Exercise 14.13.

The mouse’s RNA polymerase transcribed the jellyfish gene, and its ribosomes and tRNAs translated the mRNA with the same codon table, so the same amino acid sequence, hence the same fluorescent protein, was made. With a different code the codons would be read as other amino acids and the protein would be a meaningless, non-fluorescent chain.

Exercise 14.14 ★★★

A mutation changes the tRNA whose anticodon reads UAG so that it carries an amino acid instead of stopping. Predict its effect on the cell’s proteins in general, and on a mutant gene carrying a premature UAG.

Solution

Solution of Exercise 14.14.

UAG would no longer stop translation: every protein whose gene ends in UAG would be extended past its normal end, often losing function — a widespread harm. But a gene truncated by a premature UAG would now be read through, restoring a nearly full-length protein: the second mutation suppresses the first.

Exercise 14.15 ★★★

Discuss why a frameshift near the start of a gene is almost always worse than one near its end, and why a missense mutation can be anything from harmless to fatal.

Solution

Solution of Exercise 14.15.

A frameshift scrambles everything downstream and usually meets a stop within a few codons: near the start, essentially no correct protein is made; near the end, most of the protein is intact and may still fold and work. A missense changes one amino acid: harmless if that position tolerates the change, fatal if it is at the active site or breaks the fold — the effect depends on which amino acid, where.

14.7 Problem: A Gene Read Four Ways

Problem 14.1

Weekend problem — one short gene transcribed, translated, mutated three times and timed: from the DNA to the amino acids, and the sixty seconds a protein takes

The coding strand of a short gene reads:

ATG GCT TGG AAA CCC GAA TAC GGT CAT TTC TGA

Part I — Reading the gene.

  1. Write the template strand.
  2. Write the mRNA transcribed from the gene.
  3. Translate it, codon by codon, and give the protein’s sequence of amino acids.
  4. How many amino acids does the protein contain, and how many nucleotides of the mRNA are used to specify them, stop included?
  5. Which codon started the reading, and what fixes the reading frame for all the others?

Part II — Three mutants.

  1. Mutant 1: ATG GCT TGG AAA CCG GAA TAC GGT CAT TTC TGA. Give the protein and classify the mutation.
  2. Mutant 2: ATG GCT TGG AAA CCC GAA TAC GGT CAT TTC TGA with the ninth codon CAT replaced by CGT. Give the protein and classify.
  3. Mutant 3: ATG GCT TGA AAA CCC GAA TAC GGT CAT TTC TGA. Give the protein and classify. Which single substitution produced it?
  4. Mutant 4: a T is inserted after the sixth nucleotide: ATG GCT TTG GAA ACC CGA ATA CGG TCA TTT CTG A. Give the protein and classify.
  5. Rank the four mutants from least to most disruptive for the protein and justify.

Part III — Why the code is what it is.

  1. How many codons specify leucine? Serine? Tryptophan? Which amino acid is most likely to be changed by a random substitution in its codon, and which least?
  2. Compute the fraction of all substitutions at the third position of a leucine codon CUx that are silent.
  3. Explain why a code of two letters could not work, and why a code of four is not needed.
  4. Poly-U gives poly-phenylalanine and poly-A poly-lysine. Which entries of the table do these two experiments establish?
  5. The code is the same in the bacterium and in a human. State the consequence for transgenesis and the consequence for kinship.

Part IV — Timing the factory. A ribosome adds 10 amino acids per second and ribosomes follow each other 80 nucleotides apart along an mRNA; each mRNA lives 2 hours.

  1. How long does one ribosome take to translate this gene’s protein? And a protein of 600 amino acids?
  2. How many ribosomes can read an mRNA of 1800 nucleotides at once?
  3. If a new ribosome starts every 8 seconds, how many copies of the protein does one mRNA yield in its lifetime?
  4. A bacterium needs 20002000 copies of an enzyme in 10 minutes. How many mRNAs of the gene, transcribed at once, are needed?
  5. State the result: the protein encoded by the gene, the mutation among the four that abolished it, and the time one ribosome needs to build it.
Solution

Solution of Problem 14.1.

1. 3’-TAC CGA ACC TTT GGG CTT ATG CCA GTA AAG ACT-5’.

2. AUG GCU UGG AAA CCC GAA UAC GGU CAU UUC UGA.

3. Met–Ala–Trp–Lys–Pro–Glu–Tyr–Gly–His–Phe, stop.

4. 10 amino acids; 33 nucleotides (11 codons including the stop).

5. AUG; the position of that first AUG fixes the frame, and every following codon is read in step from it.

6. CCG is still Pro: silent, protein unchanged.

7. CGU is Arg in place of His: missense, Met–Ala–Trp–Lys–Pro–Glu–Tyr–Gly–Arg–Phe.

8. UGA is stop: Met–Ala, nonsense. A single substitution G to A in the third codon (TGG to TGA).

9. mRNA AUG GCU UUG GAA ACC CGA AUA CGG UCA UUU CUG A: Met–Ala–Leu–Glu–Thr–Arg–Ile–Arg–Ser–Phe–Leu…, no stop in the fragment: frameshift, every amino acid after the second is changed.

10. Mutant 1 (silent, no change) << mutant 2 (one amino acid changed, function possibly kept) << mutant 4 (frameshift, all but two amino acids wrong) \approx mutant 3 (nonsense, a two-amino-acid fragment: the protein is gone).

11. Leu 6, Ser 6, Trp 1. Trp, with a single codon, is changed by any substitution; Leu or Ser, with six, is most often unchanged.

12. CUU, CUC, CUA, CUG are all Leu: every one of the 3 possible substitutions at the third position is silent — 100%.

13. Two letters give 16 codons, too few for 20 amino acids plus a stop; three give 64, already more than enough, so four (256) would be waste.

14. UUU == Phe and AAA == Lys.

15. A gene transferred between species is read identically, so transgenesis works; and a shared code, inherited unchanged, is evidence that all organisms descend from a common ancestor that already used it.

16. 10 amino acids: 1s1\,\mathrm{s}. 600 amino acids: 60s60\,\mathrm{s}.

17. 1800/80221800/80 \approx 22 ribosomes at once.

18. 7200/8=9007200/8 = 900 copies.

19. One mRNA gives 600/8=75600/8 = 75 copies in 10 minutes; about 27 mRNAs are needed.

20. The protein Met–Ala–Trp–Lys–Pro–Glu–Tyr–Gly–His–Phe; mutant 3, the single G-to-A substitution creating a stop at the third codon, abolished it; one ribosome builds the ten amino acids in about one second.

Terms defined in this chapter

See all 479 terms in the glossary