用开放知识增强单细胞模型,抗噪能力更强。
Open World Knowledge Aided Single-Cell Foundation Model with Robust Cross-Modal Cell-Language Pre-training
- 用大模型检索生成细胞文本描述,引入外部知识
- 通过可靠性评估与对比学习,提升对噪声数据的鲁棒性
- 支持零样本分类和双向检索,适合多模态研究
单细胞多组学(尤其是RNA-seq)的进步揭示了细胞异质性和基因调控的深层机制。尽管基于预训练语言模型的单细胞基础模型已展现出潜力,但其仍受限于个体特征整合不足以及对多模态数据中噪声影响的忽视。为此,我们提出开放世界语言知识辅助的鲁棒单细胞基础模型OKR-CELL。该模型基于跨模态细胞-文本预训练框架,包含两项关键创新:(1) 利用基于大语言模型(LLM)的工作流结合检索增强生成(RAG),利用开放世界知识丰富细胞文本描述;(2) 设计跨模态鲁棒对齐(CRA)目标,融合样本可靠性评估、课程学习与耦合动量对比学习,增强模型对噪声数据的抵抗力。在3200万对细胞-文本数据上预训练后,OKR-CELL在6项评估任务中达到领先性能,涵盖细胞聚类、细胞类型注释、批次效应校正、少样本注释等标准基准,并在更广泛的多模态应用中表现优异,包括零样本细胞类型注释与双向细胞-文本检索。
原文摘要 · Abstract (English)
Recent advancements in single-cell multi-omics, particularly RNA-seq, have provided profound insights into cellular heterogeneity and gene regulation. While pre-trained language model (PLM) paradigm based single-cell foundation models have shown promise, they remain constrained by insufficient integration of in-depth individual profiles and neglecting the influence of noise within multi-modal data. To address both issues, we propose an Open-world Language Knowledge-Aided Robust Single-Cell Foundation Model (OKR-CELL). It is built based on a cross-modal Cell-Language pre-training framework, which comprises two key innovations: (1) leveraging Large Language Models (LLMs) based workflow with retrieval-augmented generation (RAG) enriches cell textual descriptions using open-world knowledge; (2) devising a Cross-modal Robust Alignment (CRA) objective that incorporates sample reliability assessment, curriculum learning, and coupled momentum contrastive learning to strengthen the model's resistance to noisy data. After pretraining on 32M cell-text pairs, OKR-CELL obtains cutting-edge results across 6 evaluation tasks. Beyond standard benchmarks such as cell clustering, cell-type annotation, batch-effect correction, and few-shot annotation, the model also demonstrates superior performance in broader multi-modal applications, including zero-shot cell-type annotation and bidirectional cell-text retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。