All chapters
Interview Questions
beginnerInterview Topic Map
Bioinformatics Interview Topic Areas
🎓Interview Topics
🔍Pipeline QCTs/Tv, FREEMIX, coverage metrics
⚙️GATK/ToolsBQSR, GVCF, VQSR, MarkDup
📋VCF FieldsQUAL vs GQ vs VQSLOD
📜ACMG RulesAll 28 criteria + examples
📊Statisticsp-values, FDR, PCA, LD
🧩ScenariosDebug failing pipelines
🎯Variant InterpClassify novel variants
Technical Questions with Model Answers
These are the most common technical questions in bioinformatics and clinical genomics interviews, with model answers. Study the reasoning, not just the answer.
- Q: What is the difference between WGS and WES? → WGS sequences all 3.2 Gb at 30× (~$300) and detects all variant types including structural variants and non-coding variants. WES sequences only ~50 Mb exome at 100× (~$150) and detects ~95% of Mendelian disease variants. Choose WES for cost-effective rare disease diagnosis; WGS for unsolved cases, suspected non-coding disease, or structural variant investigation.
- Q: What does Phred Q30 mean? → Q = -10 × log₁₀(P_error). Q30 = 1 error per 1000 bases = 99.9% accuracy. Clinical labs require ≥80% of bases at Q30 or above. Q20 = 1% error; Q40 = 0.01% error (HiFi PacBio standard).
- Q: What is PCR duplicate marking and why do we mark rather than remove? → PCR duplicates arise from amplification of the same DNA fragment. They inflate coverage and create false variant evidence. GATK MarkDuplicates flags them (SAM FLAG 1024) but keeps them so total reads and metrics reflect the original library. GATK tools automatically ignore marked duplicates during variant calling.
- Q: Explain BQSR in plain language. → BQSR (Base Quality Score Recalibration) corrects systematic errors in the base quality scores that the sequencer assigns. Sequencers miscalibrate quality for specific positions, cycle numbers, and base contexts. BQSR uses machine learning to detect these patterns by looking at known variant sites (dbSNP, Mills) and correcting the systematic over/under-confidence. Result: ~5–10% improvement in variant calling accuracy.
- Q: What is GVCF mode in HaplotypeCaller and why do we use it? → GVCF (Genomic VCF) outputs one record per genomic position, including non-variant blocks. This is required for joint genotyping across multiple samples - you must know that position X was not a variant in sample A (as opposed to simply missing). Joint calling improves variant detection in low-coverage regions and reduces false positives through better statistical models.
- Q: What is the Ts/Tv ratio and what does it tell you? → Transitions (Ts) are purine-purine or pyrimidine-pyrimidine changes (A↔G, C↔T). Transversions (Tv) are purine-pyrimidine changes (A/G↔C/T). In the human genome, Ts are 2× more likely than Tv due to CpG hypermutation and chemistry. Expected: ~2.0–2.1 for WGS, ~3.0–3.3 for WES (exonic regions have more CpG sites). Ts/Tv <2.0 in WES = too many false transversion calls = poor quality filtering.
- Q: What is the difference between QUAL, GQ, and VQSLOD? → QUAL: Phred-scaled confidence that a variant exists at this position (site-level). GQ: Phred-scaled confidence that the genotype call is correct for this specific sample. VQSLOD: VQSR log-odds score - negative means likely artifact, positive means likely true variant. Use GQ≥20 and VQSLOD>0 as basic filters.
Variant Interpretation Questions with Answers
- Q: Walk through classifying a novel missense variant. → (1) Check gnomAD AF: if >5% → BA1 = Benign. If absent → PM2. (2) Check ClinVar/HGMD for same or adjacent variants. (3) Assess variant type: is it PVS1-eligible? Missense → no. (4) Check functional domain: is it in a mutational hotspot? → PM1. (5) Run in silico: SIFT, PolyPhen, CADD, REVEL → PP3/BP4. (6) Check if different pathogenic missense at same codon → PM5. (7) Sum criteria → assign 1 of 5 tiers.
- Q: A variant is at AF 0.003 in NFE but 0.0001 in EAS. Does this matter? → Yes. Use the highest population frequency when applying BA1/BS1. At 0.3% in NFE, this is above 0.1% (typical BS1 threshold for many Mendelian diseases) but well below 1%. Apply BS1 at reduced strength or note it in the report. Population-specific founder variants (e.g., BRCA1 c.5266dupC is ~1% in Ashkenazi Jewish but very rare globally) require gene-specific thresholds from ClinGen.
- Q: What is compound heterozygosity? → Two pathogenic variants in trans (on different chromosomes) of the same autosomal recessive gene. Example: CFTR c.1521_1523delCTT p.(Phe508del) from father + CFTR c.1624G>T p.(Glu542*) from mother → both copies non-functional → cystic fibrosis. Confirmation requires parental testing or phasing by long-read sequencing.
- Q: How do SpliceAI scores inform classification? → SpliceAI predicts change in splice donor/acceptor probability for any variant within 50 bp of a splice site. Delta score ≥0.5 = strong evidence of splicing disruption → applies PVS1 modifier or PP3 at strong level. Delta score 0.2–0.5 = moderate; <0.1 = likely no splicing effect (BP7 supporting for synonymous).
- Q: What additional evidence reclassifies a VUS? → Functional assays (PS3/BS3 Strong), segregation data from additional affected family members (PP1 at varying strength based on LOD score), de novo confirmation (PS2), RNA splicing assays showing aberrant splicing, population data from new databases, or clinical classification by ClinGen expert panels.
Statistics & Concepts with Answers
- Q: Why is p<0.05 not sufficient for GWAS? → With 1 million SNPs tested simultaneously, 50,000 false positives are expected at p<0.05 by chance. GWAS requires p<5×10⁻⁸ (genome-wide significance), which corresponds to Bonferroni correction for ~1 million independent tests. Most GWAS hits replicate; most p=0.04 hits are false positives.
- Q: Explain batch effects and how to correct them. → Batch effects = technical variation introduced when samples are processed in different runs, labs, or on different instruments. Detected by PCA - samples cluster by batch, not biology. Correction: ComBat (R/sva package) for known batches; PEER factors for unknown technical variation; Harmony for single-cell data. Always check PCA before and after correction.
- Q: What is PPV and why does it matter for rare variant interpretation? → PPV (Positive Predictive Value) = TP/(TP+FP) = probability that a positive test result is truly positive. For a disease with 1/10,000 prevalence, even a test with 99% sensitivity and 99% specificity has PPV = ~1% (100 true positives vs 9,999 false positives in 1 million people). This is why population frequency filters (PM2/BA1) are so important - rare variants in a rare disease gene have much higher PPV than common variants.
- Q: Explain LD and why it matters for GWAS. → Linkage disequilibrium is the non-random association of alleles at two loci - variants inherited together more often than expected by chance. In GWAS, a significant SNP usually represents a haplotype in LD with the true causal variant, not the causal variant itself. Fine-mapping (FINEMAP, SuSiE) identifies the most likely causal variant. Always check r² values when interpreting GWAS hits.
Scenario-Based Questions
These open-ended scenarios test your problem-solving and are common in senior bioinformatics interviews.
- Scenario 1: "Your WES pipeline suddenly shows Ts/Tv ratios of 1.2 for all new samples. What do you do?" → Check if reference genome changed (chr vs no-chr prefix mismatch); check if BQSR was applied; check if VQSR models were retrained; examine which transitions are missing vs expected; this could indicate a BWA index issue or GATK parameter change.
- Scenario 2: "A child's WES shows no candidate variants, but clinical features strongly suggest a genetic disorder. What next?" → First check coverage gaps - are relevant genes well covered? Then consider: (1) WGS to find intronic or regulatory variants; (2) RNA-seq to detect splicing or expression abnormalities; (3) Trio analysis if not already done; (4) Mitochondrial genome analysis; (5) Repeat expansion testing (Fragile X, myotonic dystrophy); (6) Chromosomal microarray for CNVs missed by WES.
- Scenario 3: "You find a BRCA2 variant that is Pathogenic by ACMG criteria in a patient referred for cardiomyopathy. What do you do?" → This is a secondary/incidental finding. Under ACMG SF v3.2, BRCA2 is on the recommended secondary findings list. Report it separately from the primary indication finding, with appropriate counselling and consent documentation. Refer the patient to a hereditary cancer genetics clinic.
- Scenario 4: "Your variant annotation pipeline is adding gnomAD AF of 0 to variants that are actually present in gnomAD. What is the likely cause?" → Chromosome naming mismatch (chr1 vs 1), genome build mismatch (hg19 vs hg38), left-normalisation not applied before lookup, or multi-allelic sites not split. Run bcftools norm -m-any first, then re-annotate.