arXiv:2509.15429cs.LGphysics.bio-ph2025-09被引 1

用随机矩阵理论优化单细胞数据主成分分析,提升降维精度与稳定性。

Random Matrix Theory-guided sparse PCA for single-cell RNA-seq data

  • 基于随机矩阵理论设计自适应去噪算法,自动确定稀疏主成分的稀疏度。
  • 在7种技术、4种算法下均优于PCA和主流深度学习方法,细胞类型分类准确率更高。
  • 无需人工调参,适合处理异质性强、噪声大的单细胞测序数据。

单细胞RNA测序提供个体细胞的分子快照,但噪声显著。变异源于生物差异及扩增偏差、捕获效率低等技术因素,导致计算流程难以适配异质数据集或新兴技术。目前多数研究仍依赖主成分分析(PCA)进行降维,因其可解释性强且稳健,但高维下存在固有偏差。本文提出一种基于随机矩阵理论(RMT)的方法,指导现有稀疏PCA算法推断稀疏主成分。我们引入一种新型双白化算法,可自洽估计每个基因在单个细胞中的转录组噪声强度,无需假设特定噪声分布。这使得能采用RMT准则自动选择稀疏度,使稀疏PCA近乎无参数。该数学严谨的方法在保持PCA可解释性的同时,实现鲁棒、全自动的稀疏主成分推断。在七种单细胞RNA-seq技术及四种稀疏PCA算法中,该方法系统提升了主子空间重建效果,并在细胞类型分类任务中持续优于PCA、自编码器和扩散模型方法。

原文摘要 · Abstract (English)

Single-cell RNA-seq provides detailed molecular snapshots of individual cells but is notoriously noisy. Variability stems from biological differences and technical factors, such as amplification bias and limited RNA capture efficiency, making it challenging to adapt computational pipelines to heterogeneous datasets or evolving technologies. As a result, most studies still rely on principal component analysis (PCA) for dimensionality reduction, valued for its interpretability and robustness, in spite of its known bias in high dimensions. Here, we improve upon PCA with a Random Matrix Theory (RMT)-based approach that guides the inference of sparse principal components using existing sparse PCA algorithms. We first introduce a novel biwhitening algorithm which self-consistently estimates the magnitude of transcriptomic noise affecting each gene in individual cells, without assuming a specific noise distribution. This enables the use of an RMT-based criterion to automatically select the sparsity level, rendering sparse PCA nearly parameter-free. Our mathematically grounded approach retains the interpretability of PCA while enabling robust, hands-off inference of sparse principal components. Across seven single-cell RNA-seq technologies and four sparse PCA algorithms, we show that this method systematically improves the reconstruction of the principal subspace and consistently outperforms PCA-, autoencoder-, and diffusion-based methods in cell-type classification tasks.

单细胞测序主成分分析稀疏学习随机矩阵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。