GeneMamba用状态空间模型高效处理单细胞数据,比Transformer快且更准确。
GeneMamba: An Efficient and Effective Foundation Model on Single Cell Data
- 基于双向Mamba架构,线性复杂度捕捉基因双向上下文。
- 在近三千万细胞上预训练,多任务表现优于主流Transformer模型。
- 适合生物信息、基因组学研究者用于大规模单细胞数据分析。
单细胞RNA测序(scRNA-seq)能高分辨率解析细胞异质性,但其高维、稀疏和批次效应带来重大计算挑战。基于Transformer的模型虽有进展,却受限于二次复杂度和长程依赖处理不佳。本文提出GeneMamba,一种基于状态空间建模的可扩展单细胞转录组基础模型。采用Bi-Mamba架构,以线性时间复杂度捕获双向基因上下文,显著优于Transformer基线。模型在近3000万细胞上预训练,并引入通路感知对比损失与基于秩的基因编码等生物先验目标。在多批次整合、细胞类型注释和基因-基因相关性等任务中表现优异,兼具强性能、可解释性和鲁棒性。GeneMamba为构建生物可信、可扩展的大规模单细胞分析工具提供了实用方案。
原文摘要 · Abstract (English)
Single-cell RNA sequencing (scRNA-seq) enables high-resolution analysis of cellular heterogeneity, but its complexity, which is marked by high dimensionality, sparsity, and batch effects, which poses major computational challenges. Transformer-based models have made significant advances in this domain but are often limited by their quadratic complexity and suboptimal handling of long-range dependencies. In this work, we introduce GeneMamba, a scalable and efficient foundation model for single-cell transcriptomics built on state space modeling. Leveraging the Bi-Mamba architecture, GeneMamba captures bidirectional gene context with linear-time complexity, offering substantial computational gains over transformer baselines. The model is pretrained on nearly 30 million cells and incorporates biologically informed objectives, including pathway-aware contrastive loss and rank-based gene encoding. We evaluate GeneMamba across diverse tasks, including multi-batch integration, cell type annotation, and gene-gene correlation, demonstrating strong performance, interpretability, and robustness. These results position GeneMamba as a practical and powerful alternative to transformer-based methods, advancing the development of biologically grounded, scalable tools for large-scale single-cell data analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。