Biology · Glossary

What is Massively parallel sequencing?

Definition 4.2 University Biology — Year 3 · Chapter 4 — Genomics and Sequencing

Second-generation instruments read millions to billions of fragments at once. In sequencing by synthesis the fragments, with adapters ligated to their ends, are bound to a glass flow cell and amplified in place into clusters of identical molecules; the clusters are then extended one base per cycle with fluorescent, reversibly blocked nucleotides, imaged, unblocked, and extended again, so that each cycle adds one base to every cluster’s read. Reads are 100 to 300bp100\text{ to }300\,\mathrm{bp}, usually from both ends of a fragment (paired ends), with an error rate of about 10310^{-3} per base, and one run yields up to 101210^{12} bases. Third-generation long-read instruments read single molecules without amplification: by watching one polymerase incorporate fluorescent nucleotides in real time, or by threading the DNA through a protein nanopore and recording the ionic current, which each sequence of bases modulates in its own way. Reads of 10 to 100kb10\text{ to }100\,\mathrm{kb} and more span the repeats that short reads cannot, at a higher raw error rate that consensus reduces.

Left: a flow cell, the glass slide on which billions of DNA clusters are grown and read one base per cycle. Right: a pocket-sized nanopore sequencer reading single molecules as changes in an ionic current. Left: a flow cell, the glass slide on which billions of DNA clusters are grown and read one base per cycle. Right: a pocket-sized nanopore sequencer reading single molecules as changes in an ionic current.
Left: a flow cell, the glass slide on which billions of DNA clusters are grown and read one base per cycle. Right: a pocket-sized nanopore sequencer reading single molecules as changes in an ionic current.

Examples

Example 4.4 (How much is enough)

At c=5c = 5 the unsequenced fraction is e5=0.7%e^{-5} = 0.7\,\% — for a 3.2Gb3.2\,\mathrm{Gb} genome, twenty million bases in some tens of thousands of gaps. At c=10c = 10 it is 4.5×1054.5\times 10^{-5}, 150kb150\,\mathrm{kb} in total. Human genomes are routinely sequenced at c=30c = 30 not because of coverage gaps (e301013e^{-30} \approx 10^{-13}) but because each base must be read several times on each of the two chromosomes to call a heterozygous variant with confidence against an error rate of 10310^{-3} per read. Bacterial genomes are sequenced at c=50c = 50100100 for the same reason and because it is cheap. The formula also shows what coverage cannot fix: a repeat longer than a read is a place where the overlap graph branches, and no amount of short reads resolves it. That is what long reads are for.

Read in context →