All chapters
Variant Annotation
intermediateAnnotation Database Landscape
Key Annotation Databases by Category
📊Population FreqgnomAD v4.1, 1000G, TOPMed, UK Biobank
🏥ClinicalClinVar, OMIM, HGMD, ClinGen, DECIPHER
🎯Pathogenicity ScoresCADD, REVEL, SIFT, PolyPhen, AlphaMissense
🧬Gene FunctionrefGene, GENCODE, Ensembl, UniProt
✂️SplicingSpliceAI, dbscSNV, MaxEntScan
🔬CancerCOSMIC, CIViC, OncoKB, TCGA
What Annotation Adds
- Gene/transcript: which gene, exon number, cDNA/protein change (e.g., BRCA1:c.5266dupC:p.Gln1756Profs*74)
- Functional consequence: synonymous, missense, frameshift, stop-gained, splice-site
- Population frequency: gnomAD allele frequency in NFE, AFR, EAS, SAS, AMR populations
- Clinical significance: ClinVar classification (Pathogenic/VUS/Benign), OMIM disease
- Pathogenicity scores: SIFT (<0.05=deleterious), PolyPhen-2 (>0.85=probably damaging), CADD (>20=top 1%), REVEL (>0.5=likely pathogenic)
- Conservation: PhyloP, PhastCons (high score = evolutionarily conserved = likely important)
- Splicing prediction: SpliceAI delta score >0.5 = likely splicing impact
ANNOVAR
code
# Direct VCF input (recommended) table_annovar.pl sample.vcf humandb/ \ --vcfinput \ --buildver hg38 \ --out sample_annotated \ --remove \ --protocol refGene,dbnsfp47a,clinvar_20240611,gnomad211_exome,avsnp150,cosmic70 \ --operation gx,f,f,f,f,f \ --nastring . \ --thread 8 \ --xreffile gene_fullxref.txt # Protocol operations: # g = gene-based (refGene, ensGene, knownGene) # f = filter-based (frequency/pathogenicity databases) # r = region-based (regulatory, conserved regions)
VEP (Ensembl)
code
# Offline mode (fastest) vep \ --input_file sample.vcf \ --output_file annotated.vcf \ --vcf \ --offline \ --assembly GRCh38 \ --cache \ --everything \ --fork 8 \ --stats_html vep_stats.html # --everything enables: # --sift b --polyphen b --ccds --uniprot --hgvs # --symbol --numbers --domains --regulatory # --canonical --protein --biotype --af --af_1kg # --af_gnomad --max_af --pubmed --var_synonyms
Annotation Databases (hg38)
- refGene / GENCODE / Ensembl - gene models; use GENCODE v45 for comprehensive transcripts
- dbnsfp47a - 30+ pathogenicity scores in one file (SIFT, PolyPhen, CADD, REVEL, AlphaMissense)
- gnomad211_exome - 125,748 exomes; allele frequencies in 8 populations
- clinvar_20240611 - 2.6M variants with clinical interpretations
- cosmic70 - 30M somatic mutations from 1.5M tumour samples
- spliceai_filtered - pre-computed SpliceAI delta scores for all SNVs
- intervar - automated ACMG/AMP classification for missense variants