All chapters

Common File Formats

beginner

Genomics File Format Map

File Formats Through the Pipeline

📄FASTQ
📍SAM/BAM
🔬VCF/BCF
📏BED
🗺GTF/GFF
📊TSV/CSV
FASTQ: Raw reads + quality scores
SAM/BAM: Aligned reads + CIGAR
VCF/BCF: Variant calls + genotypes
BED: Genomic coordinates
GTF/GFF: Gene annotations
TSV/CSV: Annotated variants table

FASTQ Format

FASTQ stores raw sequencing reads with per-base quality scores. Each read = exactly 4 lines.

code
@SRR001666.1 HWI-EAS27_10_8 length=36     # header (read ID + metadata)
GATTTGGGGTTCAAAGCAGTATCGATCAAATAGTAAATCC  # DNA sequence
+                                           # separator (optionally repeats header)
!''*((((***+))%%%++)(%%%%).1***-+*'''')"`18  # Phred+33 quality scores
  • Phred score: Q = -10 × log₁₀(P_error)
  • Q20 = 99% accuracy (1 error per 100 bases)
  • Q30 = 99.9% accuracy (1 error per 1000 bases) - clinical minimum
  • Q40 = 99.99% accuracy - HiFi PacBio standard
  • ASCII encoding: Phred+33 (Illumina 1.8+), Phred+64 (old Illumina)
  • Paired-end files: _R1 (forward read) and _R2 (reverse read)

SAM / BAM / CRAM

SAM (Sequence Alignment/Map) stores aligned reads. BAM is binary-compressed SAM (~5× smaller). CRAM is reference-based compression (~2–3× smaller than BAM).

code
# SAM header
@HD  VN:1.6  SO:coordinate
@SQ  SN:chr1  LN:248956422
@RG  ID:sample1  SM:patient001  PL:ILLUMINA

# SAM alignment (11 mandatory fields)
#QNAME    FLAG  RNAME  POS     MAPQ  CIGAR  RNEXT  PNEXT  TLEN  SEQ          QUAL
read1     99    chr1   925952  60    101M   =      926100 249   ACGT...      IIII...
  • FLAG: bitwise value (99=paired, properly paired, mate on reverse strand)
  • MAPQ: mapping quality (60=uniquely mapped, 0=multimapper)
  • CIGAR: alignment description (101M=101 matches, 5S=5 soft-clipped, 2I=2 insertion)
  • samtools view -c -F 4 file.bam - count mapped reads
  • samtools depth -a file.bam - per-base coverage
  • samtools stats file.bam - full alignment statistics

VCF Format

code
##fileformat=VCFv4.2
##reference=hg38
##INFO=<ID=DP,Number=1,Type=Integer,Description="Read Depth">
##FORMAT=<ID=GT,Number=1,Type=String,Description="Genotype">
#CHROM  POS      ID          REF  ALT  QUAL  FILTER  INFO         FORMAT   SAMPLE
chr1    925952   rs1234567   G    A    100   PASS    DP=45;AF=0.5 GT:AD    0/1:22,23
chr17   43045629 rs28897696  G    A    .     .        .            GT       0/1
  • GT: 0/0=homozygous ref, 0/1=heterozygous, 1/1=homozygous alt, ./. =missing
  • AD: allelic depth (ref reads, alt reads)
  • DP: total read depth at position
  • GQ: genotype quality (Phred scale)
  • AF: allele frequency in INFO field (somatic) vs AF in gnomAD annotation
  • Multi-allelic sites: ALT has multiple alleles (A,T); genotype 0/2 = ref/second-alt

BED, GTF and Other Formats

code
# BED format (0-based, half-open intervals)
chr1  1000  2000  gene_name  0  +

# GTF format (1-based, closed intervals)
chr1  HAVANA  gene  11869  14409  .  +  .  gene_id "ENSG00000223972";
chr1  HAVANA  exon  11869  12227  .  +  .  gene_id "ENSG00000223972"; exon_number "1";

# FASTA format
>chr1  GRCh38
ACGTACGTACGTACGTACGT...
  • BED: 0-based coordinates; col3-col2 = feature length; widely used for intervals
  • GTF/GFF3: gene annotation; GENCODE and RefSeq provide hg38 GTF files
  • bigWig/bigBed: indexed binary track formats for genome browsers
  • MAF (Mutation Annotation Format): TCGA standard for somatic mutations
  • PLINK (.bed/.bim/.fam): GWAS data format for population genetics