Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Seurat Under the Hood

Motivation

Seurat is a standard bioinformatic tool used to explore and analyze single-cell RNA sequencing (scRNA-seq) data. Its “Guided Clustering Tutorial” serves as a foundational resource for newcomers learning standard single-cell workflows.

While the tutorial provides a solid scaffold by demonstrating key commands and high-level concepts, it lacks deep explanations of the underlying statistical and algorithmic methods. This creates a gap for readers who want to understand what happens “under the hood”.

This blog serves as an extension of the Seurat vignette, detailing the processes running behind these straightforward commands and explaining the reasoning behind each step.

Content

The content in this blog strictly follows typical functional steps in the Seurat (v5.5.0) vignette, breaking down the main command and underlying mechanisms of each step. The method for knitting content of this blog detailed in Methodology.

Overview

First, the raw count matrix (such as from 10X Genomics) undergoes pre-processing to reduce technical noise and unwanted variance. This traditional pipeline includes three key steps: Normalization, Feature selection, and Scaling data. Alternatively, SCTransform can be performed as an all-in-one approach.

Next, purified matrix is performed Principal Component Analysis (PCA) to capture major axes of biological variation across cells, which facilitates cell clustering. From all of that, Uniform Manifold Approximation and Projection (UMAP) is employed to fine-tuning clusters.

Finally, marker genes are identified each cells proving bases for cell-type determination.