All chapters
Learning Roadmap
beginnerLearning Journey Overview
Four Stages to Genomics Mastery
🌱Stage 1 (0-3 mo)Biology + Linux + File formats + Python/R basics + Statistics | Milestone: count reads in FASTQ
📚Stage 2 (3-9 mo)GATK pipeline + ANNOVAR + ACMG classification + RNA-seq | Milestone: NA12878 WES >99% sensitivity
🔬Stage 3 (9-18 mo)scRNA-seq + SVs + Long reads + Nextflow + Cloud | Milestone: DSL2 pipeline with Docker
🏆Stage 4 (18+ mo)ML + Population genetics + Clinical genomics + Multi-omics | Milestone: Publication or open-source contribution
Stage 1: Foundations (0–3 months)
Before touching any sequencing data, build a solid foundation in biology, computing, and statistics. Many people skip this and struggle later because they can run a pipeline but cannot interpret the output.
- Biology: DNA structure, central dogma, gene structure (exons/introns/UTRs), types of variants (SNV/indel/CNV/SV)
- Linux: navigation, file operations, grep/awk/sed, shell scripting, managing background jobs
- File formats: FASTQ (quality scores, Phred encoding), FASTA, SAM/BAM (FLAGS, CIGAR), VCF (genotype fields)
- Python or R basics: variables, loops, functions, data frames, reading/writing tabular files
- Statistics: mean, variance, normal distribution, p-value, multiple testing correction (Bonferroni, FDR/BH)
- Milestone check: can you write a bash script that counts reads in 10 FASTQ files and prints the totals?
- Free resources: Coursera "Genomic Data Science" (Johns Hopkins); "Bioinformatics Algorithms" (Compeau & Pevzner); EMBL-EBI training courses
Stage 2: Core Pipelines (3–9 months)
Run real pipelines on real data. Download a public WES dataset from the NCBI Sequence Read Archive (SRA) and process it end-to-end. NA12878 (HG001) is the perfect training sample - it has a validated truth set for benchmarking your results.
- FASTQ QC: FastQC, MultiQC, fastp - interpret every module; understand what adapter contamination looks like
- Alignment: BWA-MEM2 → samtools sort/index → understand CIGAR strings, read groups, mapping quality
- GATK best practices: MarkDuplicates → BQSR → HaplotypeCaller (GVCF mode) → GenotypeGVCFs → VQSR
- Variant annotation: ANNOVAR or VEP - add gnomAD AF, ClinVar significance, CADD score, gene consequence
- ACMG classification: manually classify 5 variants using the 2015 guidelines; compare to ClinVar classification
- RNA-seq starter: STAR alignment → featureCounts → DESeq2; produce a volcano plot
- Milestone: run NA12878 WES → compare your VCF to GIAB truth set using hap.py; target >99% SNV sensitivity
Stage 3: Specialisation (9–18 months)
- Single-cell: Seurat or Scanpy full pipeline on a GEO dataset; cell type annotation; trajectory analysis with Monocle3
- Structural variants: Manta + DELLY ensemble calling; CNVkit for copy number from WES; SV interpretation
- Long reads: Minimap2 for PacBio/Nanopore; HiFi variant calling with DeepVariant; phasing with WhatsHap
- Workflow automation: write a 3-step Nextflow DSL2 pipeline with Docker containers; use -resume for checkpointing
- Cloud computing: run a pipeline on AWS Batch or Google Cloud Life Sciences; understand spot instance cost savings
- Choose a focus: clinical genomics (ACMG, report writing), cancer genomics (somatic calling, mutational signatures), or research (population genetics, multi-omics)
- Milestone: submit a methods section describing your pipeline to a lab meeting or bioRxiv preprint
Stage 4: Advanced / Expert (18+ months)
- Machine learning: train a random forest on ClinVar P/B variants; learn PyTorch for deep learning; understand AlphaMissense architecture
- Population genetics: GWAS quality control pipeline; LD pruning; PCA for ancestry; fine-mapping with FINEMAP
- Clinical genomics mastery: phenotype-driven analysis with Exomiser; ACMG secondary findings reporting; variant re-interpretation
- Multi-omics: integrate WES + RNA-seq + ATAC-seq; eQTL analysis; MOFA+ for factor analysis
- Contribute open source: add a module to nf-core, submit a Bioconductor package, or fix a bug in a popular tool
- Certifications: EMQN Next Generation Sequencing scheme for clinical labs; CAP/CLIA competency testing
- Career paths: clinical molecular geneticist, research bioinformatician, computational genomics scientist, genomics product manager
Real-World Learning Path: From MSc to First Job
Priya, an MSc biotechnology graduate from Hyderabad with no prior coding experience, became a clinical bioinformatician at a major diagnostic lab in 18 months:
- Month 1–2: Completed "Introduction to Genomic Technologies" on Coursera; learned Linux basics with "The Linux Command Line" book
- Month 3–4: Downloaded NA12878 SRA data; ran FastQC; struggled with BWA index for 3 days (lesson: RTFM)
- Month 5–7: Completed full GATK WES pipeline; achieved 98.9% SNV sensitivity vs GIAB truth set
- Month 8–10: Annotated variants with ANNOVAR; manually classified 50 ClinVar variants; built a Python ACMG scorer
- Month 11–14: Joined a local genomics lab as a research assistant; analysed 20 clinical WES cases under supervision
- Month 15–18: Built a Nextflow pipeline for the lab; presented at a local genomics conference; received job offer
- Key insight: "I spent 80% of my time on 20% of the tasks. Debugging alignment and understanding VCF format took longer than running the tools themselves."