arXiv:2603.24626q-bio.GNcs.LG2026-03

对比15种单细胞基因组填补方法,发现传统模型更稳定,效果不等于生物意义提升。

A Large-Scale Comparative Analysis of Imputation Methods for Single-Cell RNA Sequencing Data

  • 横跨30个数据集、10种实验协议,系统评估15种填补方法
  • 传统方法在基因表达恢复上优于深度学习类方法
  • 填补效果好≠下游分析更准确,需按任务选方法

单细胞RNA测序(scRNA-seq)可在单细胞水平进行基因表达分析,但受技术限制存在缺失事件(dropout),导致真实表达基因被误记为零值,扭曲表达分布并影响后续分析。已有众多填补方法被提出,涵盖传统统计模型与基于深度学习(DL)的方法。然而,现有基准测试仅覆盖有限的方法、数据集和下游分析,性能对比尚不清晰。本研究构建了涵盖7类方法的15种scRNA-seq填补方法的全面基准,评估对象包括30个来自10种实验协议的数据集,以及6项下游分析任务。结果表明:传统方法(如基于模型、平滑、低秩矩阵的方法)整体表现优于深度学习方法(包括基于扩散、GAN、GNN和自编码器的方法)。此外,数值上基因表达恢复效果好,并不必然带来下游分析中生物学可解释性的提升,包括细胞聚类、差异表达分析、标记基因识别、轨迹推断和细胞类型注释。同时,方法性能在不同数据集、实验协议和分析任务间差异显著,无单一方法始终最优。研究为针对特定分析目标选择合适的填补方法提供了实践指导,强调应根据任务特性评估填补性能。

原文摘要 · Abstract (English)

Background: Single-cell RNA sequencing (scRNA-seq) enables gene expression profiling at cellular resolution but is inherently affected by sparsity caused by dropout events, where expressed genes are recorded as zeros due to technical limitations. These artifacts distort gene expression distributions and compromise downstream analyses. Numerous imputation methods have been proposed to recover latent transcriptional signals. These methods range from traditional statistical models to deep learning (DL)-based methods. However, their comparative performance remains unclear, as existing benchmarks evaluate only a limited subset of methods, datasets, and downstream analyses. Results: We present a comprehensive benchmark of 15 scRNA-seq imputation methods spanning 7 methodological categories, including traditional and DL-based methods. Methods are evaluated across 30 datasets from 10 experimental protocols on 6 downstream analyses. Results show that traditional methods, such as model-based, smoothing-based, and low-rank matrix-based methods, generally outperform DL-based methods, including diffusion-based, GAN-based, GNN-based, and autoencoder-based methods. In addition, strong performance in numerical gene expression recovery does not necessarily translate into improved biological interpretability in downstream analyses, including cell clustering, differential expression analysis, marker gene analysis, trajectory analysis, and cell type annotation. Furthermore, method performance varies substantially across datasets, protocols, and downstream analyses, with no single method consistently outperforming others. Conclusions: Our findings provide practical guidance for selecting imputation methods tailored to specific analytical objectives and underscore the importance of task-specific evaluation when assessing imputation performance in scRNA-seq data analysis.

单细胞测序数据填补方法对比生物信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。