All chapters
Linux for Genomics
beginnerLinux Command Categories
Essential Linux for Genomics
📁Navigationpwd, ls, cd, find
📋File Opscp, mv, rm, ln, mkdir
🔍Text Searchgrep, awk, sed, cut
📦Compressiongzip, bgzip, tabix, zcat
⚡Jobs & HPCnohup, screen, tmux, kill
🧬Genomics Filessamtools, bcftools, bedtools
Essential Commands
code
# Navigation pwd # print working directory ls -lh # list files with human-readable sizes ls -lhS # sort by size (largest first) cd /path/to/dir # change directory cd ~ # go to home directory cd - # go to previous directory # File operations mkdir -p data/fastq # create nested directories cp -r src/ dest/ # copy directory recursively mv old.fastq new.fastq # move/rename rm -rf dir/ # remove directory (careful!) ln -s /path/file link # create symbolic link # Viewing files cat file.txt # print entire file head -n 20 file.txt # first 20 lines tail -n 20 file.txt # last 20 lines tail -f logfile.txt # follow log in real-time less file.txt # interactive viewer (q to quit) wc -l file.txt # count lines
Searching & Text Processing
code
# grep - search patterns
grep "PASS" sample.vcf # lines containing PASS
grep -v "^#" sample.vcf # exclude header lines
grep -c "PASS" sample.vcf # count matching lines
grep -i "brca" genes.txt # case-insensitive
grep -E "chr[0-9]+\s" file.bed # extended regex
# awk - column processing
awk '{print $1, $2}' file.tsv # print columns 1 and 2
awk '$6 > 30' variants.tsv # filter by quality
awk 'NR>1' file.tsv # skip header row
awk -F'\t' '{print $3}' file.tsv # tab-delimited
# sed - stream editor
sed 's/chr//g' file.vcf # remove "chr" prefix
sed -n '10,20p' file.txt # print lines 10–20
sed '/^#/d' file.vcf # delete header lines
# sort & uniq
sort -k1,1 -k2,2n file.bed # sort by chr then position
sort file.txt | uniq -c # count unique occurrences
cut -f1,2,3 file.tsv # extract columns 1–3Working with Genomics Files
code
# FASTQ operations
zcat sample.fastq.gz | head -8 # view compressed FASTQ
zcat sample.fastq.gz | wc -l | awk '{print $1/4}' # count reads
zgrep -c "^@" sample.fastq.gz # count reads (faster)
# Disk & jobs
du -sh data/ # directory size
df -h # disk space
top / htop # system resources
nohup command & # run in background
jobs # list background jobs
kill %1 # kill job 1
screen -S session1 # named screen session
tmux new -s pipeline # tmux session
# File compression
gzip file.fastq # compress
bgzip file.vcf # bgzip for tabix indexing
tabix -p vcf file.vcf.gz # index VCF
tabix -p bed file.bed.gz # index BEDShell Scripting Basics
code
#!/bin/bash
# Script to run FastQC on all FASTQ files
SAMPLE_DIR="/data/fastq"
OUT_DIR="/data/qc"
mkdir -p $OUT_DIR
for R1 in $SAMPLE_DIR/*_R1.fastq.gz; do
SAMPLE=$(basename $R1 _R1.fastq.gz)
R2="${R1/_R1/_R2}"
echo "Processing: $SAMPLE"
fastqc "$R1" "$R2" -o "$OUT_DIR" -t 4
done
echo "All QC done. Running MultiQC..."
multiqc $OUT_DIR -o $OUT_DIRReal-World Linux Workflows in Genomics
These one-liners solve common daily problems in a genomics lab. Knowing them saves hours of manual work.
code
# Count reads in a FASTQ file (fast)
echo "Reads: $(( $(zcat sample_R1.fastq.gz | wc -l) / 4 ))"
# Check if a specific gene region has coverage
samtools depth -a -r chr17:43044294-43125364 sample.bam | awk '{sum+=$3; n++} END{print "BRCA1 mean depth:", sum/n}'
# Extract all PASS variants from VCF
bcftools view -f PASS variants.vcf.gz -O z -o pass_only.vcf.gz
# Count variants per chromosome
bcftools stats pass_only.vcf.gz | grep "^SN" | grep "per chromosome"
# Find samples with very low coverage (batch QC)
for bam in *.bam; do
MEAN=$(samtools depth -a $bam | awk '{sum+=$3; n++} END{print sum/n}')
echo "$bam: ${MEAN}x"
done | sort -t: -k2 -n
# Check Ts/Tv ratio - should be 2.0-2.1 for WGS, 3.0-3.3 for WES
bcftools stats sample.vcf.gz | grep "^TSTV"
# Rename samples in bulk (add prefix)
for f in *.fastq.gz; do mv "$f" "COHORT_$f"; done
# Convert chromosome names: chr1 → 1 (for tools expecting no "chr" prefix)
sed 's/^chr//' sample.vcf > sample_nochr.vcf