All chapters
Single-cell Genomics
advancedSingle-Cell vs Bulk RNA-seq
Bulk vs Single-Cell vs Spatial
🧫Bulk RNA-seqAverage of all cells | Cheap ($100-300) | Established methods | Misses cell heterogeneity
🔬scRNA-seq500-10K cells | $1000-5000 | UMAP clusters | Cell type resolution | Tumour microenvironment
🗺SpatialGene expression + location | Visium/Xenium | Tissue context | Highest cost ($3000+)
Why Single-Cell?
Bulk RNA-seq averages gene expression across thousands of cells, masking cell-type-specific signals. scRNA-seq resolves cell-level heterogeneity - critical for understanding tumour microenvironments, developmental trajectories, and rare cell populations.
- 10X Genomics Chromium: droplet-based; 500–10,000 cells/run; 3' gene expression standard
- Smart-seq2: plate-based; full-length transcripts; lower throughput, higher sensitivity
- CITE-seq: simultaneous RNA + surface protein (ADT) quantification
- Spatial transcriptomics (Visium, Xenium): gene expression with tissue location
- scATAC-seq: chromatin accessibility at single-cell level
- Multiome: simultaneous scRNA + scATAC from same cell
Standard scRNA-seq Pipeline (Seurat)
code
library(Seurat)
# Load Cell Ranger output
data <- Read10X("cellranger_output/filtered_feature_bc_matrix/")
obj <- CreateSeuratObject(counts = data, min.cells = 3, min.features = 200)
# Quality control
obj[["percent.mt"]] <- PercentageFeatureSet(obj, pattern = "^MT-")
obj <- subset(obj,
subset = nFeature_RNA > 200 &
nFeature_RNA < 6000 &
percent.mt < 20)
# Normalisation
obj <- NormalizeData(obj) # log1p(count/total * 10000)
# Feature selection
obj <- FindVariableFeatures(obj, nfeatures = 3000)
# Scale and PCA
obj <- ScaleData(obj, vars.to.regress = "percent.mt")
obj <- RunPCA(obj, npcs = 50)
# Clustering
obj <- FindNeighbors(obj, dims = 1:30)
obj <- FindClusters(obj, resolution = 0.5)
# Visualisation
obj <- RunUMAP(obj, dims = 1:30)
DimPlot(obj, label = TRUE)
# Marker genes
markers <- FindAllMarkers(obj, only.pos = TRUE, min.pct = 0.25)Cell Type Annotation
- Manual: check known markers per cluster (CD3E=T cells, CD19=B cells, CD14=monocytes, EPCAM=epithelial)
- SingleR: automated annotation against reference datasets (Human Cell Atlas, ENCODE)
- CellTypist: logistic regression classifier; 300+ cell types; clinical grade accuracy
- scType: uses curated marker gene database; works offline
- Azimuth (Seurat): reference-based label transfer using multimodal reference atlases
- Doublet detection: DoubletFinder or Scrublet before annotation (doublets = 2 cells in 1 droplet)
Advanced Analyses
- Trajectory/pseudotime: Monocle3, scVelo (RNA velocity) - order cells in developmental time
- Cell-cell communication: CellChat, NicheNet - infer ligand-receptor interactions
- Batch correction: Harmony (fast), scVI (deep learning), Seurat integration - essential for multi-sample
- Differential abundance: edgeR-based testing for cell-type frequency changes between conditions
- Gene regulatory networks: SCENIC - transcription factor regulons from co-expression + motifs
- CNV inference (cancer): inferCNV, CopyKAT - detect tumour cells from scRNA-seq