Second-generation instruments read millions to billions of fragments at once. In sequencing by synthesis the fragments, with adapters ligated to their ends, are bound to a glass flow cell and amplified in place into clusters of identical molecules; the clusters are then extended one base per cycle with fluorescent, reversibly blocked nucleotides, imaged, unblocked, and extended again, so that each cycle adds one base to every cluster’s read. Reads are , usually from both ends of a fragment (paired ends), with an error rate of about per base, and one run yields up to bases. Third-generation long-read instruments read single molecules without amplification: by watching one polymerase incorporate fluorescent nucleotides in real time, or by threading the DNA through a protein nanopore and recording the ionic current, which each sequence of bases modulates in its own way. Reads of and more span the repeats that short reads cannot, at a higher raw error rate that consensus reduces.
Examples
Example 4.4 (How much is enough)
At the unsequenced fraction is — for a genome, twenty million bases in some tens of thousands of gaps. At it is , in total. Human genomes are routinely sequenced at not because of coverage gaps () but because each base must be read several times on each of the two chromosomes to call a heterozygous variant with confidence against an error rate of per read. Bacterial genomes are sequenced at – for the same reason and because it is cheap. The formula also shows what coverage cannot fix: a repeat longer than a read is a place where the overlap graph branches, and no amount of short reads resolves it. That is what long reads are for.