All chapters

Learning Roadmap

beginner

Learning Journey Overview

Four Stages to Genomics Mastery

🌱Stage 1 (0-3 mo)Biology + Linux + File formats + Python/R basics + Statistics | Milestone: count reads in FASTQ
📚Stage 2 (3-9 mo)GATK pipeline + ANNOVAR + ACMG classification + RNA-seq | Milestone: NA12878 WES >99% sensitivity
🔬Stage 3 (9-18 mo)scRNA-seq + SVs + Long reads + Nextflow + Cloud | Milestone: DSL2 pipeline with Docker
🏆Stage 4 (18+ mo)ML + Population genetics + Clinical genomics + Multi-omics | Milestone: Publication or open-source contribution

Stage 1: Foundations (0–3 months)

Before touching any sequencing data, build a solid foundation in biology, computing, and statistics. Many people skip this and struggle later because they can run a pipeline but cannot interpret the output.

  • Biology: DNA structure, central dogma, gene structure (exons/introns/UTRs), types of variants (SNV/indel/CNV/SV)
  • Linux: navigation, file operations, grep/awk/sed, shell scripting, managing background jobs
  • File formats: FASTQ (quality scores, Phred encoding), FASTA, SAM/BAM (FLAGS, CIGAR), VCF (genotype fields)
  • Python or R basics: variables, loops, functions, data frames, reading/writing tabular files
  • Statistics: mean, variance, normal distribution, p-value, multiple testing correction (Bonferroni, FDR/BH)
  • Milestone check: can you write a bash script that counts reads in 10 FASTQ files and prints the totals?
  • Free resources: Coursera "Genomic Data Science" (Johns Hopkins); "Bioinformatics Algorithms" (Compeau & Pevzner); EMBL-EBI training courses

Stage 2: Core Pipelines (3–9 months)

Run real pipelines on real data. Download a public WES dataset from the NCBI Sequence Read Archive (SRA) and process it end-to-end. NA12878 (HG001) is the perfect training sample - it has a validated truth set for benchmarking your results.

  • FASTQ QC: FastQC, MultiQC, fastp - interpret every module; understand what adapter contamination looks like
  • Alignment: BWA-MEM2 → samtools sort/index → understand CIGAR strings, read groups, mapping quality
  • GATK best practices: MarkDuplicates → BQSR → HaplotypeCaller (GVCF mode) → GenotypeGVCFs → VQSR
  • Variant annotation: ANNOVAR or VEP - add gnomAD AF, ClinVar significance, CADD score, gene consequence
  • ACMG classification: manually classify 5 variants using the 2015 guidelines; compare to ClinVar classification
  • RNA-seq starter: STAR alignment → featureCounts → DESeq2; produce a volcano plot
  • Milestone: run NA12878 WES → compare your VCF to GIAB truth set using hap.py; target >99% SNV sensitivity

Stage 3: Specialisation (9–18 months)

  • Single-cell: Seurat or Scanpy full pipeline on a GEO dataset; cell type annotation; trajectory analysis with Monocle3
  • Structural variants: Manta + DELLY ensemble calling; CNVkit for copy number from WES; SV interpretation
  • Long reads: Minimap2 for PacBio/Nanopore; HiFi variant calling with DeepVariant; phasing with WhatsHap
  • Workflow automation: write a 3-step Nextflow DSL2 pipeline with Docker containers; use -resume for checkpointing
  • Cloud computing: run a pipeline on AWS Batch or Google Cloud Life Sciences; understand spot instance cost savings
  • Choose a focus: clinical genomics (ACMG, report writing), cancer genomics (somatic calling, mutational signatures), or research (population genetics, multi-omics)
  • Milestone: submit a methods section describing your pipeline to a lab meeting or bioRxiv preprint

Stage 4: Advanced / Expert (18+ months)

  • Machine learning: train a random forest on ClinVar P/B variants; learn PyTorch for deep learning; understand AlphaMissense architecture
  • Population genetics: GWAS quality control pipeline; LD pruning; PCA for ancestry; fine-mapping with FINEMAP
  • Clinical genomics mastery: phenotype-driven analysis with Exomiser; ACMG secondary findings reporting; variant re-interpretation
  • Multi-omics: integrate WES + RNA-seq + ATAC-seq; eQTL analysis; MOFA+ for factor analysis
  • Contribute open source: add a module to nf-core, submit a Bioconductor package, or fix a bug in a popular tool
  • Certifications: EMQN Next Generation Sequencing scheme for clinical labs; CAP/CLIA competency testing
  • Career paths: clinical molecular geneticist, research bioinformatician, computational genomics scientist, genomics product manager

Real-World Learning Path: From MSc to First Job

Priya, an MSc biotechnology graduate from Hyderabad with no prior coding experience, became a clinical bioinformatician at a major diagnostic lab in 18 months:

  • Month 1–2: Completed "Introduction to Genomic Technologies" on Coursera; learned Linux basics with "The Linux Command Line" book
  • Month 3–4: Downloaded NA12878 SRA data; ran FastQC; struggled with BWA index for 3 days (lesson: RTFM)
  • Month 5–7: Completed full GATK WES pipeline; achieved 98.9% SNV sensitivity vs GIAB truth set
  • Month 8–10: Annotated variants with ANNOVAR; manually classified 50 ClinVar variants; built a Python ACMG scorer
  • Month 11–14: Joined a local genomics lab as a research assistant; analysed 20 clinical WES cases under supervision
  • Month 15–18: Built a Nextflow pipeline for the lab; presented at a local genomics conference; received job offer
  • Key insight: "I spent 80% of my time on 20% of the tasks. Debugging alignment and understanding VCF format took longer than running the tools themselves."