arXiv:2601.14653cs.LGq-bio.GN2026-01中稿 · ACM-BCB 2026被引 1

用最优传输方法高效填补单细胞数据中的大块缺失值

Efficient Imputation for Patch-based Missing Single-cell Data via Cluster-regularized Optimal Transport

  • 基于聚类正则化最优传输,捕捉数据内在结构
  • 在高缺失率下仍保持高精度,计算速度显著提升
  • 适合处理异质性高、维度大的缺失数据,生物医学研究适用

单细胞测序数据中的缺失值严重影响生物学洞察的提取。现有填补方法通常假设数据均匀且完整,难以应对大面积缺失的情况。本文提出CROT(聚类正则化最优传输)算法,基于最优传输框架,专为表格格式的片状缺失数据设计。该方法能有效捕获高缺失率下的数据内在结构,在保证高填补精度的同时大幅降低运行时间,展现出对大规模数据集的可扩展性和高效性。本工作为具有结构性缺失的异质性高维数据提供了稳健的填补方案,解决了生物与临床数据分析中的关键挑战。代码已开源:https://github.com/yuyuliu11037/CROT。

原文摘要 · Abstract (English)

Missing data in single-cell sequencing datasets poses significant challenges for extracting meaningful biological insights. However, existing imputation approaches, which often assume uniformity and data completeness, struggle to address cases with large patches of missing data. In this paper, we present CROT (Cluster-Regularized Optimal Transport), an optimal transport-based imputation algorithm designed to handle patch-based missing data in tabular formats. Our approach effectively captures the underlying data structure in the presence of significant missingness. Notably, it achieves superior imputation accuracy while significantly reducing runtime, demonstrating its scalability and efficiency for large-scale datasets. This work introduces a robust solution for imputation in heterogeneous, high-dimensional datasets with structured data absence, addressing critical challenges in both biological and clinical data analysis. Our code is available on GitHub, https://github.com/yuyuliu11037/CROT.

单细胞分析数据填补最优传输高维数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。