All chapters

Single-cell Genomics

advanced

Single-Cell vs Bulk RNA-seq

Bulk vs Single-Cell vs Spatial

🧫Bulk RNA-seqAverage of all cells | Cheap ($100-300) | Established methods | Misses cell heterogeneity
🔬scRNA-seq500-10K cells | $1000-5000 | UMAP clusters | Cell type resolution | Tumour microenvironment
🗺SpatialGene expression + location | Visium/Xenium | Tissue context | Highest cost ($3000+)

Why Single-Cell?

Bulk RNA-seq averages gene expression across thousands of cells, masking cell-type-specific signals. scRNA-seq resolves cell-level heterogeneity - critical for understanding tumour microenvironments, developmental trajectories, and rare cell populations.

  • 10X Genomics Chromium: droplet-based; 500–10,000 cells/run; 3' gene expression standard
  • Smart-seq2: plate-based; full-length transcripts; lower throughput, higher sensitivity
  • CITE-seq: simultaneous RNA + surface protein (ADT) quantification
  • Spatial transcriptomics (Visium, Xenium): gene expression with tissue location
  • scATAC-seq: chromatin accessibility at single-cell level
  • Multiome: simultaneous scRNA + scATAC from same cell

Standard scRNA-seq Pipeline (Seurat)

code
library(Seurat)

# Load Cell Ranger output
data <- Read10X("cellranger_output/filtered_feature_bc_matrix/")
obj <- CreateSeuratObject(counts = data, min.cells = 3, min.features = 200)

# Quality control
obj[["percent.mt"]] <- PercentageFeatureSet(obj, pattern = "^MT-")
obj <- subset(obj,
  subset = nFeature_RNA > 200 &
           nFeature_RNA < 6000 &
           percent.mt < 20)

# Normalisation
obj <- NormalizeData(obj)  # log1p(count/total * 10000)

# Feature selection
obj <- FindVariableFeatures(obj, nfeatures = 3000)

# Scale and PCA
obj <- ScaleData(obj, vars.to.regress = "percent.mt")
obj <- RunPCA(obj, npcs = 50)

# Clustering
obj <- FindNeighbors(obj, dims = 1:30)
obj <- FindClusters(obj, resolution = 0.5)

# Visualisation
obj <- RunUMAP(obj, dims = 1:30)
DimPlot(obj, label = TRUE)

# Marker genes
markers <- FindAllMarkers(obj, only.pos = TRUE, min.pct = 0.25)

Cell Type Annotation

  • Manual: check known markers per cluster (CD3E=T cells, CD19=B cells, CD14=monocytes, EPCAM=epithelial)
  • SingleR: automated annotation against reference datasets (Human Cell Atlas, ENCODE)
  • CellTypist: logistic regression classifier; 300+ cell types; clinical grade accuracy
  • scType: uses curated marker gene database; works offline
  • Azimuth (Seurat): reference-based label transfer using multimodal reference atlases
  • Doublet detection: DoubletFinder or Scrublet before annotation (doublets = 2 cells in 1 droplet)

Advanced Analyses

  • Trajectory/pseudotime: Monocle3, scVelo (RNA velocity) - order cells in developmental time
  • Cell-cell communication: CellChat, NicheNet - infer ligand-receptor interactions
  • Batch correction: Harmony (fast), scVI (deep learning), Seurat integration - essential for multi-sample
  • Differential abundance: edgeR-based testing for cell-type frequency changes between conditions
  • Gene regulatory networks: SCENIC - transcription factor regulons from co-expression + motifs
  • CNV inference (cancer): inferCNV, CopyKAT - detect tumour cells from scRNA-seq