arXiv:2506.20697q-bio.CBcs.LG2025-06被引 4

无需特征筛选,scMamba高效整合单细胞多组学数据

scMamba: A Scalable Foundation Model for Single-Cell Multi-Omics Integration Beyond Highly Variable Feature Selection

  • 将基因组区域视为词元,构建细胞语句的分块编码策略
  • 在多个数据集上优于现有方法,显著提升聚类与注释准确率
  • 适合大规模单细胞图谱构建与生物机制探索

单细胞多组学技术使同时解析个体细胞中多种组学层成为可能。整合这些多模态数据能深入揭示细胞身份、调控过程和疾病机制。然而,当前方法常依赖高变基因或峰的选择进行预处理,可能无意中丢失关键生物学信息。我们提出scMamba,一种无需先验特征选择即可整合单细胞多组学数据的基础模型,同时保留基因组位置信息。scMamba采用基于补丁的细胞标记化策略,将基因组区域视为词元(tokens),细胞视为句子。依托状态空间对偶性,从高维稀疏的单细胞多组学数据中提取丰富生物见解。此外,我们设计了一种新的对比学习方法,结合余弦相似度正则化,相比传统方法显著提升跨组学层对齐效果。在多个数据集上的系统评估表明,scMamba在保留生物变异、对齐组学层及提升聚类、细胞类型注释和轨迹推断等下游任务表现上均显著优于现有先进方法。研究结果表明,scMamba是大规模单细胞多组学整合的强大工具,具备处理大规模图谱并推动生物学发现的能力。

原文摘要 · Abstract (English)

The advent of single-cell multi-omics technologies has enabled the simultaneous profiling of diverse omics layers within individual cells. Integrating such multimodal data provides unprecedented insights into cellular identity, regulatory processes, and disease mechanisms. However, it remains challenging, as current methods often rely on selecting highly variable genes or peaks during preprocessing, which may inadvertently discard crucial biological information. Here, we present scMamba, a foundation model designed to integrate single-cell multi-omics data without the need for prior feature selection while preserving genomic positional information. scMamba introduces a patch-based cell tokenization strategy that treats genomics regions as words (tokens) and cells as sentences. Building upon the concept of state space duality, scMamba distills rich biological insights from high-dimensional, sparse single-cell multi-omics data. Additionally, our novel contrastive learning approach, enhanced with cosine similarity regularization, enables superior alignment across omics layers compared to traditional methods. Systematic benchmarking across multiple datasets demonstrates that scMamba significantly outperforms state-of-the-art methods in preserving biological variation, aligning omics layers, and enhancing key downstream tasks such as clustering, cell type annotation, and trajectory inference. Our findings position scMamba as a powerful tool for large-scale single-cell multi-omics integration, capable of handling large-scale atlases and advancing biological discovery.

单细胞多组学基础模型数据整合生物信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。