Genomics Technologies
beginnerSequencing Generations Comparison
Three Generations of Sequencing
Sanger Sequencing (1st Generation)
Developed by Frederick Sanger in 1977, who won the Nobel Prize for it. Uses chain-terminating dideoxynucleotides (ddNTPs) - analogues of normal dNTPs that lack the 3'-OH group needed for chain extension. When a ddNTP is incorporated, elongation stops, producing a ladder of fragments of different lengths. Capillary electrophoresis separates by size; fluorescent dye colours identify each base.
Despite being 50 years old, Sanger sequencing remains the gold standard for single-variant confirmation in clinical labs. Any pathogenic variant identified by WES or NGS that will influence clinical management must be confirmed by Sanger before reporting.
- Read length: 600โ1000 bp (longest of any sequencing method)
- Accuracy: >99.9% raw - the most accurate sequencing available for a single read
- Throughput: 1โ96 reads per run (very low - not scalable to whole genomes)
- Cost: ~$5โ10 per reaction (cheap per reaction, expensive per genome)
- Turnaround: 1โ2 days (primer design + PCR + sequencing)
- Clinical use: confirm WES/NGS findings before reporting; targeted family testing; confirm pathogenic variants in known disease genes
- Limitation: cannot detect CNVs, structural variants, or variants below ~20% allele frequency
# Typical Sanger confirmation workflow # 1. Design primers flanking the variant (Primer3, NCBI Primer-BLAST) # Target: 200-400 bp amplicon; variant in middle third # 2. PCR amplification from patient DNA # 3. Cleanup (ExoSAP-IT or gel purification) # 4. Cycle sequencing with BigDye Terminator v3.1 # 5. Capillary electrophoresis on ABI 3730xl # 6. Interpret chromatogram with FinchTV or Sequencher # Example: Confirming BRCA1 c.5266dupC # Primer F: 5'-CTTACCTGTTTTATGCATTTT-3' # Primer R: 5'-TTTCATAAGAAAATTTTGAGCAGTT-3' # Expected: heterozygous insertion peak visible in chromatogram
Next Generation Sequencing (2nd Generation)
Next Generation Sequencing (NGS) revolutionised genomics from 2007 onwards by enabling massively parallel sequencing - millions of DNA fragments sequenced simultaneously in a single run. The key innovation was "sequencing by synthesis" combined with bridge amplification on a flow cell surface, generating millions of clusters each reading the same fragment.
The cost of sequencing a human genome dropped from $3 billion (2003) to under $200 (2023), following a trajectory faster than Moore's Law. This democratised genomics, making clinical WES affordable at โน15,000โ30,000 in India today.
Platform Comparison
- Illumina (NovaSeq, NextSeq, MiSeq) - 150โ300 bp reads, >99.9% accuracy, most widely used, ~$1/Gb
- MGI/BGI (DNBSEQ) - similar to Illumina, lower cost, used in large population studies
- Ion Torrent (Thermo Fisher) - semiconductor sequencing, fast runs, higher error in homopolymers
- Complete Genomics - proprietary combinatorial probe-anchor synthesis, population-scale
Long-Read Sequencing (3rd Generation)
- PacBio HiFi (SMRT) - 15โ25 kb reads, >99.9% HiFi accuracy, methylation detection, $10โ15/Gb
- Oxford Nanopore (MinION, PromethION) - real-time sequencing, up to Mb-length reads, ~85โ98% raw accuracy
- Advantages over short reads: resolves complex SVs, repeats, phasing, full-length transcripts
- Disadvantage: higher error rate in raw reads (Nanopore), higher cost per Gb (PacBio)
Library Preparation Key Concepts
- Fragmentation: sonication or enzymatic to target insert size (150โ300 bp for WES)
- End-repair & A-tailing: creates blunt ends then adds 3' adenine overhang
- Adapter ligation: adds platform-specific sequences with unique molecular identifiers (UMIs)
- PCR amplification: 8โ12 cycles; too many cycles creates duplicates
- Capture (WES): biotinylated probes hybridise to exonic regions, pull down with streptavidin beads
- Size selection: removes adapter dimers; typically 200โ400 bp final library size