All chapters

Hands-on Projects

beginner

Project Difficulty Ladder

From Beginner to Portfolio-Level Projects

1️⃣
FASTQ QC Pipeline (Beginner)

SRA download → FastQC → fastp → MultiQC; skills: Linux, basic QC interpretation

2️⃣
VCF Filter & Parse (Beginner)

ClinVar VCF → cyvcf2 → filter P/LP → TSV output; skills: Python, VCF format

3️⃣
WES Gold Standard (Intermediate)

NA12878 WES → full GATK pipeline → hap.py benchmark; skills: full pipeline, metrics

4️⃣
TCGA RNA-seq DEG (Intermediate)

GDC download → DESeq2 → volcano plot → pathway; skills: R, statistics, biology

5️⃣
Nextflow DSL2 Pipeline (Advanced)

Modular NF pipeline with Docker containers + HPC submission; skills: workflow, DevOps

6️⃣
Multi-omics Integration (Portfolio)

TCGA WES+RNA+methylation → MOFA+ → cancer subtypes; skills: all disciplines, publishable

Beginner Projects with Step-by-Step Guides

Each project is designed to build a specific skill. Complete them in order - each one assumes skills from the previous.

  • Project 1 - FASTQ QC Pipeline: Download SRR10045924 from NCBI SRA using "prefetch SRR10045924 && fastq-dump --split-files --gzip SRR10045924". Run FastQC, then fastp with --detect_adapter_for_pe, then MultiQC. Write a 1-page report: what QC issues did you find? Would you proceed with analysis?
  • Project 2 - VCF Filter Script: Download clinvar.vcf.gz from NCBI FTP. Use cyvcf2 to extract all Pathogenic/Likely Pathogenic variants in BRCA1 and BRCA2. Output a TSV with: CHROM, POS, REF, ALT, CLNSIG, gene, consequence. Count how many P/LP variants each gene has.
  • Project 3 - Coverage Visualisation: Download NA12878 chr17 BAM from 1000 Genomes project. Run "samtools depth -a -r chr17:43044294-43125364 NA12878.bam" to get per-base BRCA1 coverage. Plot with matplotlib: x-axis = genomic position, y-axis = depth, add exon boundaries as vertical lines.
  • Project 4 - FASTQ Stats from Scratch: Without using FastQC, write a Python script that parses a FASTQ.gz file and reports: total reads, total bases, mean read length, % bases ≥ Q30, GC content (%). Compare your results to FastQC output to verify.
  • Project 5 - gnomAD API Lookup: Take a VCF with 20 variants. For each, query the gnomAD GraphQL API (gnomad.broadinstitute.org/api) to retrieve gnomAD v4.1 allele frequency in NFE and global. Add frequencies as new columns. Flag variants with AF > 0.01 as likely benign.

Intermediate Projects with Expected Outcomes

  • Project 6 - WES Gold Standard Benchmark: Download NA12878 WES data (SRA: SRR098401). Run full GATK4 best-practices pipeline. Compare your output VCF to GIAB HG001 v4.2.1 truth set using hap.py. Target: SNV sensitivity >99%, precision >98%, indel sensitivity >95%. Document every parameter and version number.
  • Project 7 - Python ACMG Classifier: Build a CLI tool that reads an ANNOVAR-annotated VCF, applies BA1 (gnomAD_AF>0.05), PM2 (gnomAD_AF<0.0001), PVS1 (frameshift/nonsense in disease gene), PP3 (CADD>20 + SIFT<0.05), and BP4 (CADD<10 + PolyPhen<0.1), outputs a classification for each variant. Test on 50 ClinVar variants; measure concordance.
  • Project 8 - TCGA RNA-seq DEG Analysis: Download TCGA-BRCA RNA-seq raw counts (GDC portal). Compare ER+ vs triple-negative breast cancer. Run DESeq2. Expected: hundreds of DEGs; top genes should include ERBB2, ESR1, MKI67. Run clusterProfiler GO enrichment. Make a publication-quality volcano plot with ggplot2.
  • Project 9 - Mutational Signatures: Download TCGA LUAD somatic VCF calls from GDC. Run SigProfilerExtractor on the mutation matrix. Identify APOBEC signature (SBS2/SBS13 = C>T and C>G at TCW context) and smoking signature (SBS4 = C>A at CC context). Correlate signature exposure with patient smoking status from TCGA clinical data.
  • Project 10 - Population Ancestry PCA: Download 1000 Genomes Phase 3 VCF for chr22 (fast to work with). QC filter: MAF>0.05, call rate>95%. LD prune with PLINK (--indep-pairwise 50 10 0.1). PCA with PLINK --pca 10. Plot PC1 vs PC2 coloured by superpopulation (AFR, AMR, EAS, EUR, SAS). Can you distinguish Indian (SAS) from other populations?

Advanced Projects - Portfolio Level

  • Project 11 - scRNA-seq PBMC Atlas: Download GSE130148 (10X Chromium PBMC from GEO). Run Cell Ranger count, then Seurat: QC (nFeature 200–6000, MT<20%), SCTransform normalisation, PCA, Harmony batch correction, UMAP, clustering at resolution 0.5. Annotate clusters with CellTypist. Expected cell types: CD4 T, CD8 T, NK, B, Monocyte, DC, Platelet. Write a 2-page methods section describing every parameter choice.
  • Project 12 - Production Nextflow Pipeline: Build a Nextflow DSL2 pipeline: (1) fastpQC process, (2) bwaMem process, (3) samtools flagstat process, (4) GATK HaplotypeCaller process. Each process should use a Docker container (biocontainers). Add a nextflow.config with resource requests. Test on local and submit to SLURM. Add retry logic for failed processes.
  • Project 13 - SV Caller Benchmark: Take GIAB HG002 WGS BAM. Call SVs with Manta, DELLY, and PBSV (or Sniffles for Nanopore). Merge calls with SURVIVOR. Evaluate against GIAB HG002 SV truth set v0.6 using truvari (--passonly --sizemin 50). Report: precision, recall, F1 per SV type (DEL, INS, DUP, INV). Which caller performs best for deletions vs insertions?
  • Project 14 - ML Pathogenicity Model: Download ClinVar VCF. Extract 2000 Pathogenic and 2000 Benign variants. Annotate with CADD, REVEL, gnomAD AF, SIFT, PolyPhen, SpliceAI, phyloP. Train a random forest classifier (scikit-learn). Use 5-fold cross-validation. Plot ROC curve (AUC target >0.95). Compare to REVEL alone. What features are most important (feature importance plot)?
  • Project 15 - Multi-omics TCGA Integration: Download TCGA-COAD WES somatic mutations, RNA-seq normalised counts, and DNA methylation 450K array data for 200 samples. Run MOFA+ (Multi-Omics Factor Analysis) to identify latent factors. Visualise factors in UMAP. Do any factors correlate with microsatellite instability status (from TCGA clinical data)? This is a publishable analysis.