All chapters
Understanding VCF Files
intermediateVCF Field Structure
Anatomy of a VCF Data Line
📍CHROMchr1 ... chrX, chrY, chrM
🔢POS1-based position on chromosome
🏷IDrsID from dbSNP or . if novel
🔄REF/ALTReference and alternate alleles
⭐QUALPhred confidence variant exists
🔍FILTERPASS or filter name (LowQual)
📊INFODP, AF, gnomAD_AF, CADD...
🧬FORMAT/GT0/1 het, 1/1 hom-alt, 0/0 ref
VCF Structure
code
## Meta-information lines ##fileformat=VCFv4.2 ##reference=file:///data/hg38.fa ##contig=<ID=chr1,length=248956422> ##INFO=<ID=DP,Number=1,Type=Integer,Description="Total read depth"> ##INFO=<ID=AF,Number=A,Type=Float,Description="Allele frequency"> ##FILTER=<ID=LowQual,Description="Low quality"> ##FORMAT=<ID=GT,Number=1,Type=String,Description="Genotype"> ##FORMAT=<ID=AD,Number=R,Type=Integer,Description="Allelic depths"> ##FORMAT=<ID=GQ,Number=1,Type=Integer,Description="Genotype quality"> ## Data lines #CHROM POS ID REF ALT QUAL FILTER INFO FORMAT SAMPLE chr1 925952 rs2296290 G A 1256 PASS DP=87;AF=0.52;QD=14.4 GT:AD:GQ 0/1:42,45:99 chr17 43045629 rs28897696 G A . . . GT 0/1
Genotype Fields Explained
- GT 0/0 - homozygous reference (both alleles = REF)
- GT 0/1 - heterozygous (one REF, one ALT) - expected in germline Mendelian
- GT 1/1 - homozygous alternate (both alleles = ALT)
- GT 0/2 - heterozygous for second ALT allele at multi-allelic site
- GT ./. - missing genotype (insufficient coverage or filtered)
- GT 0|1 - phased: pipe separates maternal|paternal allele
- GQ (Genotype Quality): Phred-scaled confidence; GQ<20 = unreliable genotype
- AD (Allelic Depth): reads supporting REF,ALT (e.g., 42,45 for balanced het)
- PL (Phred-scaled Likelihoods): 0/0,0/1,1/1 genotype likelihoods; lowest = called GT
VCF Manipulation with bcftools
code
# Count variants by type bcftools stats sample.vcf.gz | grep "^SN" # Filter PASS variants only bcftools view -f PASS sample.vcf.gz -o pass.vcf.gz -O z # Extract specific region bcftools view sample.vcf.gz chr17:43000000-44000000 # Merge multiple VCFs bcftools merge -m all *.vcf.gz -o merged.vcf.gz -O z # Normalise indels (left-align and split multi-allelics) bcftools norm -m-any -f hg38.fa sample.vcf.gz -o norm.vcf.gz # Extract sample-level fields to TSV bcftools query -f '[%CHROM\t%POS\t%REF\t%ALT\t%GT\t%GQ\n]' sample.vcf.gz