All chapters

Linux for Genomics

beginner

Linux Command Categories

Essential Linux for Genomics

📁Navigationpwd, ls, cd, find
📋File Opscp, mv, rm, ln, mkdir
🔍Text Searchgrep, awk, sed, cut
📦Compressiongzip, bgzip, tabix, zcat
⚡Jobs & HPCnohup, screen, tmux, kill
🧬Genomics Filessamtools, bcftools, bedtools

Essential Commands

code
# Navigation
pwd                     # print working directory
ls -lh                   # list files with human-readable sizes
ls -lhS                  # sort by size (largest first)
cd /path/to/dir          # change directory
cd ~                     # go to home directory
cd -                     # go to previous directory

# File operations
mkdir -p data/fastq      # create nested directories
cp -r src/ dest/         # copy directory recursively
mv old.fastq new.fastq   # move/rename
rm -rf dir/              # remove directory (careful!)
ln -s /path/file link    # create symbolic link

# Viewing files
cat file.txt             # print entire file
head -n 20 file.txt      # first 20 lines
tail -n 20 file.txt      # last 20 lines
tail -f logfile.txt      # follow log in real-time
less file.txt            # interactive viewer (q to quit)
wc -l file.txt           # count lines

Searching & Text Processing

code
# grep - search patterns
grep "PASS" sample.vcf              # lines containing PASS
grep -v "^#" sample.vcf             # exclude header lines
grep -c "PASS" sample.vcf           # count matching lines
grep -i "brca" genes.txt            # case-insensitive
grep -E "chr[0-9]+\s" file.bed     # extended regex

# awk - column processing
awk '{print $1, $2}' file.tsv       # print columns 1 and 2
awk '$6 > 30' variants.tsv          # filter by quality
awk 'NR>1' file.tsv                 # skip header row
awk -F'\t' '{print $3}' file.tsv   # tab-delimited

# sed - stream editor
sed 's/chr//g' file.vcf             # remove "chr" prefix
sed -n '10,20p' file.txt            # print lines 10–20
sed '/^#/d' file.vcf               # delete header lines

# sort & uniq
sort -k1,1 -k2,2n file.bed          # sort by chr then position
sort file.txt | uniq -c             # count unique occurrences
cut -f1,2,3 file.tsv               # extract columns 1–3

Working with Genomics Files

code
# FASTQ operations
zcat sample.fastq.gz | head -8      # view compressed FASTQ
zcat sample.fastq.gz | wc -l | awk '{print $1/4}'  # count reads
zgrep -c "^@" sample.fastq.gz       # count reads (faster)

# Disk & jobs
du -sh data/                        # directory size
df -h                               # disk space
top / htop                          # system resources
nohup command &                     # run in background
jobs                                # list background jobs
kill %1                             # kill job 1
screen -S session1                  # named screen session
tmux new -s pipeline                # tmux session

# File compression
gzip file.fastq                     # compress
bgzip file.vcf                      # bgzip for tabix indexing
tabix -p vcf file.vcf.gz            # index VCF
tabix -p bed file.bed.gz            # index BED

Shell Scripting Basics

code
#!/bin/bash
# Script to run FastQC on all FASTQ files

SAMPLE_DIR="/data/fastq"
OUT_DIR="/data/qc"

mkdir -p $OUT_DIR

for R1 in $SAMPLE_DIR/*_R1.fastq.gz; do
  SAMPLE=$(basename $R1 _R1.fastq.gz)
  R2="${R1/_R1/_R2}"
  echo "Processing: $SAMPLE"
  fastqc "$R1" "$R2" -o "$OUT_DIR" -t 4
done

echo "All QC done. Running MultiQC..."
multiqc $OUT_DIR -o $OUT_DIR

Real-World Linux Workflows in Genomics

These one-liners solve common daily problems in a genomics lab. Knowing them saves hours of manual work.

code
# Count reads in a FASTQ file (fast)
echo "Reads: $(( $(zcat sample_R1.fastq.gz | wc -l) / 4 ))"

# Check if a specific gene region has coverage
samtools depth -a -r chr17:43044294-43125364 sample.bam | awk '{sum+=$3; n++} END{print "BRCA1 mean depth:", sum/n}'

# Extract all PASS variants from VCF
bcftools view -f PASS variants.vcf.gz -O z -o pass_only.vcf.gz

# Count variants per chromosome
bcftools stats pass_only.vcf.gz | grep "^SN" | grep "per chromosome"

# Find samples with very low coverage (batch QC)
for bam in *.bam; do
  MEAN=$(samtools depth -a $bam | awk '{sum+=$3; n++} END{print sum/n}')
  echo "$bam: ${MEAN}x"
done | sort -t: -k2 -n

# Check Ts/Tv ratio - should be 2.0-2.1 for WGS, 3.0-3.3 for WES
bcftools stats sample.vcf.gz | grep "^TSTV"

# Rename samples in bulk (add prefix)
for f in *.fastq.gz; do mv "$f" "COHORT_$f"; done

# Convert chromosome names: chr1 → 1 (for tools expecting no "chr" prefix)
sed 's/^chr//' sample.vcf > sample_nochr.vcf