Fundamental Biology
beginnerCentral Dogma Visual
Flow of Genetic Information
DNA Structure
DNA (deoxyribonucleic acid) is a double-stranded helix made of nucleotide monomers. Each nucleotide has three parts: a deoxyribose sugar (the backbone), a phosphate group (linking sugars), and one of four nitrogenous bases (A, T, C, G). The two strands run antiparallel (one 5'→3', the other 3'→5') and pair via specific hydrogen bonds - this base complementarity is the foundation of all molecular biology.
- Adenine (A) ↔ Thymine (T) - 2 hydrogen bonds (weaker pairing; AT-rich regions melt first)
- Cytosine (C) ↔ Guanine (G) - 3 hydrogen bonds (stronger; GC-rich genomes are more thermostable)
- Chargaff's rule: in any dsDNA, %A = %T and %C = %G - key quality check for DNA purity
- The sugar-phosphate backbone runs 5'→3'; all polymerases synthesise new DNA in this direction
- DNA coils around octamers of histone proteins (H2A, H2B, H3, H4) to form nucleosomes (~147 bp per nucleosome)
- Nucleosomes compact further into 30 nm fibres → loops → chromosome domains → chromosomes
- Histone modifications (H3K27ac = active enhancer; H3K4me3 = active promoter) regulate gene expression
Central Dogma
The central dogma describes the directional flow of genetic information in all living cells. Understanding each step is essential for interpreting variant consequences - a frameshift in an exon disrupts the coding sequence; a variant at a splice site disrupts pre-mRNA processing; a variant in the 3' UTR may destroy a miRNA binding site.
- Transcription: RNA Pol II reads the template strand (3'→5') and synthesises pre-mRNA (5'→3')
- 5' capping: 7-methylguanosine cap added co-transcriptionally - protects mRNA from exonucleases, required for ribosome binding
- Splicing: spliceosome removes introns at GT...AG splice sites (GU-AG rule in RNA); ~95% of human genes are multi-exonic
- Alternative splicing: one gene can produce multiple protein isoforms - BRCA1 has 23 splice variants; dystrophin has 78 exons
- Poly-A tail: ~200 adenines added after poly-A signal (AATAAA); essential for nuclear export and mRNA stability
- Translation: ribosome reads mRNA codons (3-base triplets) 5'→3'; tRNAs deliver amino acids matching each codon
- Start codon: AUG (methionine); stop codons: UAA, UAG, UGA - nonsense variants create premature stop codons
- Post-translational modification (PTM): phosphorylation, ubiquitination, glycosylation - variants can disrupt PTM sites
Gene Structure in Detail
A typical human gene spans 10–100 kb of genomic DNA but produces only 1–3 kb of protein-coding mRNA. Understanding every element of gene structure is critical for interpreting variant consequences: a variant 2 bp from an exon-intron boundary is a potential splice-site variant even if it's technically intronic.
- Promoter (−2000 to TSS): TATA box at −30 bp binds TFIID; CpG islands at TSS in ~70% of human genes
- Transcription start site (TSS): where RNA Pol II begins - variants here can reduce transcription initiation
- 5' UTR: from TSS to start codon; upstream ORFs (uORFs) in 5' UTR can dramatically reduce translation efficiency
- Exon: protein-coding sequence; on average 300 bp; human BRCA1 exon 11 is unusually large at 3.4 kb
- Intron: spliced out; average 3,400 bp; intronic variants at ±1/±2 positions from exon are classified as splice-site (PVS1)
- 3' UTR: from stop codon to poly-A tail; miRNA binding sites here regulate mRNA degradation (e.g., VEGFA 3' UTR)
- Enhancers: can be up to 1 Mb away but loop to contact promoters via CTCF-cohesin; ENCODE has mapped >1 million
- Pseudogenes: non-functional copies of genes; PMS2 has multiple pseudogenes that confound variant calling
Types of Genetic Variants - With Clinical Examples
Variants differ in their size, mechanism of pathogenicity, and detection method. Choosing the right sequencing approach depends on which variant type you're looking for.
- SNP (Single Nucleotide Polymorphism): >1% frequency; mostly benign; e.g., rs334 (HBB Val6Glu) causes sickle cell - rare exception where a common SNP is pathogenic
- SNV (Single Nucleotide Variant): any frequency; clinically relevant rare variants; e.g., BRCA1 c.5266dupC is actually a single-base insertion but historically called "185delAG" in old nomenclature
- Missense: amino acid change; e.g., TP53 p.Arg248Trp is a gain-of-function hotspot in many cancers; classified by SIFT, PolyPhen, CADD
- Nonsense: creates premature stop codon; e.g., CFTR c.1624G>T p.Glu542* causes cystic fibrosis; typically Pathogenic via PVS1
- Frameshift indel: shifts reading frame; e.g., BRCA1 c.5266dupC p.Gln1756Profs*74 - founder mutation in Ashkenazi Jewish population; truncates protein
- In-frame indel: maintains reading frame; may preserve partial function; e.g., DMD in-frame deletions cause Becker MD (milder than frameshift = Duchenne)
- CNV deletion/duplication: e.g., SMN1 deletion (both copies) causes spinal muscular atrophy; PMP22 duplication causes Charcot-Marie-Tooth 1A
- Trinucleotide repeat: CAG expansion >36 in HTT causes Huntington disease; FMR1 CGG expansion >200 causes Fragile X (anticipation)
- Chromosomal translocation: BCR::ABL1 fusion (Philadelphia chromosome) in CML; creates constitutively active tyrosine kinase
Real-World Example: Why Variant Type Matters for WES
A genetic counsellor refers a family with Duchenne muscular dystrophy (DMD). Cascade testing is requested for at-risk male relatives. The affected proband's DMD variant is a large exon 45-52 deletion - a CNV spanning ~170 kb.
- WES misses large CNVs: this deletion spans multiple exons but WES shows normal coverage in some exons due to read depth averaging
- Correct test: multiplex ligation-dependent probe amplification (MLPA) or chromosomal microarray to detect the deletion
- WES result: appears normal - a false negative for the family's mutation
- Lesson: WES detects SNVs and small indels well, but large deletions/duplications require CNV-specific tools (CNVkit, GATK gCNV) or MLPA
- ATGC Flow includes CNV detection; always check coverage plots for DMD/PARK2/PMP22/SMN1 in WES reports