用基因标记集指导细胞表征优化,提升单细胞模型泛化能力
Prototype Guided Post-pretraining for Single-Cell Representation Learning

- 引入标记基因集作为结构先验,指导预训练后表征优化
- 在多个任务中提升下游性能,最高增益达15%
- 适合需要强泛化能力的单细胞数据建模场景
从基因表达数据中进行单细胞表征学习(SCRL)为揭示细胞功能背后的复杂调控逻辑提供了途径。受自然语言建模中大语言模型的启发,近期已提出若干将基因视为词元、细胞视为句子的单细胞预训练模型。然而,这些模型受限于细胞类型分布的长尾特性,在基因表达数据的协变量偏移下难以泛化。尽管微调常被用于缓解此问题,但性能仍存在上限。为此,我们提出CellRefine,一种位于单细胞基础模型预训练与微调之间的后预训练方法。CellRefine采用多目标策略,利用标记基因集作为结构先验,引导细胞潜在嵌入流形的优化。在多个计算生物学任务中,实验结果表明,CellRefine持续提升下游性能,最高增益达15%。
原文摘要 · Abstract (English)
Single-cell representation learning (SCRL) from gene expression data offers a way to uncover the complex regulatory logic underlying cellular function. Inspired by large language models in natural language modeling, several single-cell pretrained models have recently been proposed that treat genes as tokens and cells as sentences. However, these models are fundamentally limited by the long-tailed nature of cell-type distributions and struggle to generalize under covariate shifts in gene expression data. While fine-tuning is often used to mitigate these issues, we observe that performance remains bounded. To address this challenge, we introduce CellRefine, a post-pretraining method that operates between the pretraining and fine-tuning stages of a single-cell foundation model. CellRefine uses a multi-faceted objective that incorporates marker-gene sets as structural priors to guide post-pretraining and refine the latent embedding manifold of cells. Across multiple computational biology tasks, empirical results show that CellRefine consistently improves downstream performance, yielding gains up to 15%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。