scMamba提升神经退行性疾病单核RNA测序分析质量
scMamba: A Pre-Trained Model for Single-Nucleus RNA Sequencing Analysis in Neurodegenerative Disorders
- 基于Mamba架构设计,融合基因嵌入与双向块处理数据
- 预训练后在细胞注释、双峰检测等任务中超越基准方法
- 无需变量基因筛选,适合多源脑组织数据整合分析
单核RNA测序(snRNA-seq)显著推进了对神经退行性疾病病因的理解。然而,死后脑组织样本质量低,加之疾病异质性导致的高变异性,使得多源snRNA-seq数据整合分析面临挑战。为此,我们提出scMamba,一种专为神经退行性疾病设计的预训练模型,旨在提升snRNA-seq分析的质量与实用性。受近期Mamba模型启发,scMamba引入线性适配层、基因嵌入和双向Mamba块,实现对snRNA-seq数据的高效处理并保留原始输入信息。值得注意的是,scMamba通过在snRNA-seq数据上进行预训练,学习到可泛化的细胞与基因特征,无需依赖降维或高变基因选择。我们在多种下游任务中验证了其优越性,包括细胞类型注释、双峰检测、数据填补及差异表达基因识别。
原文摘要 · Abstract (English)
Single-nucleus RNA sequencing (snRNA-seq) has significantly advanced our understanding of the disease etiology of neurodegenerative disorders. However, the low quality of specimens derived from postmortem brain tissues, combined with the high variability caused by disease heterogeneity, makes it challenging to integrate snRNA-seq data from multiple sources for precise analyses. To address these challenges, we present scMamba, a pre-trained model designed to improve the quality and utility of snRNA-seq analysis, with a particular focus on neurodegenerative diseases. Inspired by the recent Mamba model, scMamba introduces a novel architecture that incorporates a linear adapter layer, gene embeddings, and bidirectional Mamba blocks, enabling efficient processing of snRNA-seq data while preserving information from the raw input. Notably, scMamba learns generalizable features of cells and genes through pre-training on snRNA-seq data, without relying on dimension reduction or selection of highly variable genes. We demonstrate that scMamba outperforms benchmark methods in various downstream tasks, including cell type annotation, doublet detection, imputation, and the identification of differentially expressed genes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。