Blogs
Beyond Expression Values: Using Single-Cell Foundation Models to Unlock Richer Biological Signals
Single-cell sequencing allows researchers to study disease, development, and treatment response by measuring gene activity across thousands of individual cells rather than averaging it across bulk tissue. Most standard analysis workflows still use each gene's expression value as the primary representation for downstream analysis. Downstream tools such as differential gene expression (DGE), clustering, and cell-type annotation operate primarily on those expression values.
The challenge is that genes do not function in isolation. Every gene operates within biological pathways, interacts with dozens of others, and its biological role changes depending on cellular and disease context. Reducing it to a single expression value leaves much of that context outside the analysis.
Recent research suggests an alternative approach, and Strand is actively incorporating the same principle into its single-cell analysis capabilities.
The Limitation of Classical Single-Cell Analysis
A gene expression matrix cannot distinguish a two-fold rise in a transcription factor that rewires an entire regulatory program from a two-fold rise in a housekeeping gene unless additional context is added, because numerically they are identical.
This is due to the fact that traditional workflows, including differential gene expression (DGE), biomarker identification, clustering, and cell-type annotation, rely on gene expression matrices as the primary input. Each gene contributes one numerical value per cell. This approach is well established and interpretable, but it has a limitation: it cannot represent that a gene belongs to a specific pathway, interacts with particular partners, or carries functional context that changes its relevance from one disease to another.
This limitation can restrict performance in applications such as patient stratification, biomarker discovery, and machine learning, which often benefit from richer feature representations.
Context-Enriched Gene Representations
A study published in Patterns (Liu et al., 2026) presents one approach to addressing this limitation through scELMo, a framework that combines large language models with single-cell analysis.
scELMo performs cell clustering, batch-effect correction, and cell-type annotation without training a new model, and its fine-tuning framework extends to in-silico treatment and perturbation analysis, all with a lighter structure and lower resource requirements than a pre-trained single-cell foundation model. The mechanism here is simple, where an LLM writes a functional description of each gene and that description becomes a numerical embedding, which is combined with the measured expression.
The same framework was also applied to computationally screen candidate therapeutic targets by simulating gene knockouts and evaluating whether diseased cell states shifted toward healthy profiles before laboratory validation.
The takeaway is that you don't need to build a new foundation model to get foundation-model-level insight. Existing foundation models, used directly or after fine-tuning for specific datasets, can provide richer representations when genes are treated as components of interconnected biological systems rather than isolated measurements.
Foundation Model Representations for Single-Cell Analysis
The same principle underlies the single-cell foundation model (scFM) capability Strand is developing for clients. Strand evaluates whether an existing single-cell foundation model, used directly or after project-specific fine-tuning, can generate a richer representation of every cell, rather than defaulting immediately to established analyses such as DGE, biomarker discovery, or cell-type annotation.
Workflow
Omics measurements are processed through a single-cell foundation model to generate an embedding that represents each cell and, collectively, each patient. Instead of relying solely on the gene-by-cell expression matrix, the workflow generates a higher-dimensional representation in which each value reflects patterns learned from large-scale biological datasets, including gene identity, pathway information, and gene-gene relationships.
Downstream Analysis
- Patient Stratification: Embeddings capture coordinated biological relationships beyond individual expression changes, often producing more clearly separated patient groups.
- Machine Learning Readiness: Dense, information-rich embedding vectors reduce the need for extensive feature engineering while supporting predictive modeling.
- Biomarker Discovery with Context: Candidate biomarkers are identified within representations that retain pathway and interaction context, supporting downstream biological interpretation.
A Complementary Layer for Existing Workflows
Foundation-model-derived embeddings are intended to extend, not replace, established single-cell analysis methods. Differential gene expression, clustering, and cell-type annotation remain essential components of single-cell studies because they provide interpretable and well-validated biological insights.
The additional representation layer captures biological relationships that conventional expression matrices were not designed to represent, making it particularly useful for downstream applications such as patient stratification, predictive modeling, and translational target discovery where richer feature representations can improve analysis.
As single-cell foundation models continue to mature, applying existing models or adapting them through targeted fine-tuning will become a practical way of capturing additional biological information from datasets that research groups already generate, allowing downstream analyses to begin with representations that capture more than expression values alone.
To explore how foundation-model-derived embeddings can be applied to your single-cell data, connect with Dr. Swaraj Basu.
Reference: Liu, T., Chen, T., Zheng, W., Luo, X., Chen, Y., & Zhao, H. (2026). Embeddings from language models are good learners for single-cell data analysis. Patterns, 7, 101431.
Let's Connect
Let's Connect
download the case study.