用文本知识增强单细胞基因表达数据表示,提升分析效果。
Language-Enhanced Representation Learning for Single-Cell Transcriptomics
- 融合基因表达与文本描述,双模态学习细胞特征。
- 在细胞注释和聚类任务中优于现有方法,泛化能力更强。
- 适合生物信息学研究者及单细胞数据分析人员。
单细胞RNA测序(scRNA-seq)为细胞异质性提供了精细洞察。近期研究利用单细胞大语言模型(scLLMs)实现有效的表示学习,但这些模型仅依赖转录组数据,忽略了文本描述中的互补生物学知识。为此,我们提出scMMGPT,一种新型多模态框架,用于增强单细胞转录组的语言感知表示学习。与现有方法不同,scMMGPT采用稳健的细胞表示提取机制,保留定量基因表达数据,并引入创新的两阶段预训练策略,结合判别精度与生成灵活性。大量实验表明,scMMGPT在细胞注释、聚类等关键下游任务中显著优于单模态和多模态基线模型,并在分布外场景中表现出更优的泛化能力。
原文摘要 · Abstract (English)
Single-cell RNA sequencing (scRNA-seq) offers detailed insights into cellular heterogeneity. Recent advancements leverage single-cell large language models (scLLMs) for effective representation learning. These models focus exclusively on transcriptomic data, neglecting complementary biological knowledge from textual descriptions. To overcome this limitation, we propose scMMGPT, a novel multimodal framework designed for language-enhanced representation learning in single-cell transcriptomics. Unlike existing methods, scMMGPT employs robust cell representation extraction, preserving quantitative gene expression data, and introduces an innovative two-stage pre-training strategy combining discriminative precision with generative flexibility. Extensive experiments demonstrate that scMMGPT significantly outperforms unimodal and multimodal baselines across key downstream tasks, including cell annotation and clustering, and exhibits superior generalization in out-of-distribution scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。