All chapters

Best Practices

beginner

Best Practices at a Glance

Six Pillars of Good Genomics Practice

🔄ReproducibilityGit version control; pinned software versions; Nextflow -resume; config files not hardcoded paths
✅Data QualityFastQC before processing; VerifyBamID2 contamination; coverage checks; Ts/Tv monitoring
📋Variant ReportingSanger confirmation; HGVS nomenclature; ACMG criteria documented; VUS annual review
🔐Data SecurityEncrypted storage; HIPAA/GDPR compliance; de-identification; data access agreements
📚DocumentationJupyter/RMarkdown notebooks; parameter logs; data provenance tracking; lab notebook
👥Code ReviewPeer review all pipeline changes; clinical consequences of bugs; test with known samples

Reproducibility & Documentation

  • Version control everything: Git commit all scripts, configs, and analysis notebooks - not just code
  • Pin software versions: conda env export > environment.yml; docker build with specific tags (bwa:0.7.17)
  • Workflow management: Nextflow/Snakemake with checksums on inputs; -resume/--rerun-incomplete for failed jobs
  • Document parameters: save all CLI flags used; use config files, not hardcoded paths
  • Lab notebook: Jupyter/RMarkdown for analysis with narrative, code, and figures together
  • Data provenance: track which pipeline version processed which samples; store in LIMS or database

Data Quality Standards

  • Raw QC: always run FastQC before ANY processing; fail samples with Q30 <80% or unusual GC distribution
  • Coverage checks: verify mean depth AND uniformity before calling variants; don't assume coverage from file size
  • Contamination: run VerifyBamID2 on every clinical sample; FREEMIX >0.03 = investigate/fail
  • Ts/Tv monitoring: SNP Ts/Tv <2.5 for WES = likely quality issue; check filter settings
  • Sex concordance: infer genetic sex from X heterozygosity; mismatch = sample swap
  • Relatedness checks: in multi-sample studies, calculate kinship coefficients (king, PLINK --genome)

Clinical Variant Reporting

  • Always confirm reportable variants by Sanger sequencing or orthogonal method
  • Never report a variant without applying all relevant ACMG criteria
  • Use HGVS nomenclature correctly: NM_007294.4(BRCA1):c.5266dupC p.(Gln1756Profs*74)
  • Disclose limitations: variants in pseudogenes, repeat regions, or low coverage may be missed
  • Document evidence: save literature references and database screenshots for each classified variant
  • Re-interpretation: review VUS variants annually as new evidence emerges; proactive recontact policy
  • Never phone a pathogenic variant result - always written report reviewed by clinical geneticist

Computing & Security

  • Never store patient genomic data on personal devices or public cloud without encryption + consent
  • GDPR/HIPAA compliance: de-identify data before sharing; use data access agreements for public datasets
  • Backup: 3-2-1 rule - 3 copies, 2 different media, 1 offsite; test restores quarterly
  • Resource management: don't run GATK with default -Xmx4g on a 50 GB genome; profile memory usage
  • Cluster etiquette: request appropriate resources; kill stuck jobs; don't monopolise shared nodes
  • Code review: bioinformatics bugs have clinical consequences - peer-review all pipeline changes