All chapters
Common File Formats
beginnerGenomics File Format Map
File Formats Through the Pipeline
📄FASTQ
📍SAM/BAM
🔬VCF/BCF
📏BED
🗺GTF/GFF
📊TSV/CSV
FASTQ: Raw reads + quality scores
SAM/BAM: Aligned reads + CIGAR
VCF/BCF: Variant calls + genotypes
BED: Genomic coordinates
GTF/GFF: Gene annotations
TSV/CSV: Annotated variants table
FASTQ Format
FASTQ stores raw sequencing reads with per-base quality scores. Each read = exactly 4 lines.
code
@SRR001666.1 HWI-EAS27_10_8 length=36 # header (read ID + metadata) GATTTGGGGTTCAAAGCAGTATCGATCAAATAGTAAATCC # DNA sequence + # separator (optionally repeats header) !''*((((***+))%%%++)(%%%%).1***-+*'''')"`18 # Phred+33 quality scores
- Phred score: Q = -10 × log₁₀(P_error)
- Q20 = 99% accuracy (1 error per 100 bases)
- Q30 = 99.9% accuracy (1 error per 1000 bases) - clinical minimum
- Q40 = 99.99% accuracy - HiFi PacBio standard
- ASCII encoding: Phred+33 (Illumina 1.8+), Phred+64 (old Illumina)
- Paired-end files: _R1 (forward read) and _R2 (reverse read)
SAM / BAM / CRAM
SAM (Sequence Alignment/Map) stores aligned reads. BAM is binary-compressed SAM (~5× smaller). CRAM is reference-based compression (~2–3× smaller than BAM).
code
# SAM header @HD VN:1.6 SO:coordinate @SQ SN:chr1 LN:248956422 @RG ID:sample1 SM:patient001 PL:ILLUMINA # SAM alignment (11 mandatory fields) #QNAME FLAG RNAME POS MAPQ CIGAR RNEXT PNEXT TLEN SEQ QUAL read1 99 chr1 925952 60 101M = 926100 249 ACGT... IIII...
- FLAG: bitwise value (99=paired, properly paired, mate on reverse strand)
- MAPQ: mapping quality (60=uniquely mapped, 0=multimapper)
- CIGAR: alignment description (101M=101 matches, 5S=5 soft-clipped, 2I=2 insertion)
- samtools view -c -F 4 file.bam - count mapped reads
- samtools depth -a file.bam - per-base coverage
- samtools stats file.bam - full alignment statistics
VCF Format
code
##fileformat=VCFv4.2 ##reference=hg38 ##INFO=<ID=DP,Number=1,Type=Integer,Description="Read Depth"> ##FORMAT=<ID=GT,Number=1,Type=String,Description="Genotype"> #CHROM POS ID REF ALT QUAL FILTER INFO FORMAT SAMPLE chr1 925952 rs1234567 G A 100 PASS DP=45;AF=0.5 GT:AD 0/1:22,23 chr17 43045629 rs28897696 G A . . . GT 0/1
- GT: 0/0=homozygous ref, 0/1=heterozygous, 1/1=homozygous alt, ./. =missing
- AD: allelic depth (ref reads, alt reads)
- DP: total read depth at position
- GQ: genotype quality (Phred scale)
- AF: allele frequency in INFO field (somatic) vs AF in gnomAD annotation
- Multi-allelic sites: ALT has multiple alleles (A,T); genotype 0/2 = ref/second-alt
BED, GTF and Other Formats
code
# BED format (0-based, half-open intervals) chr1 1000 2000 gene_name 0 + # GTF format (1-based, closed intervals) chr1 HAVANA gene 11869 14409 . + . gene_id "ENSG00000223972"; chr1 HAVANA exon 11869 12227 . + . gene_id "ENSG00000223972"; exon_number "1"; # FASTA format >chr1 GRCh38 ACGTACGTACGTACGTACGT...
- BED: 0-based coordinates; col3-col2 = feature length; widely used for intervals
- GTF/GFF3: gene annotation; GENCODE and RefSeq provide hg38 GTF files
- bigWig/bigBed: indexed binary track formats for genome browsers
- MAF (Mutation Annotation Format): TCGA standard for somatic mutations
- PLINK (.bed/.bim/.fam): GWAS data format for population genetics