无需逐点标注,用类别标签学全局距离度量,提升单细胞数据对比精度。
Global Ground Metric Learning with Applications to scRNA data
- 基于类别标签学习共享空间中的全局距离度量,不依赖点级标注。
- 在多疾病单细胞数据上,优化了嵌入、聚类与分类性能。
- 适用于无共同支持分布的任意概率分布比较,解释性强。
最优传输为比较概率分布提供了稳健框架,其效果高度依赖底层基度量的选择。传统方法要么使用预定义度量(如欧氏距离),要么通过有标签数据监督学习特定任务的度量。但预定义度量难以捕捉特征内在结构和重要性差异,现有监督学习方法通常无法跨类别泛化,或仅适用于具有共享支撑集的分布。为此,我们提出一种新方法,用于在共享度量空间中学习任意分布的全局度量。该方法提供类似全局度量的点间距离,但训练仅需分布级别的类别标签。学习到的全局基度量可显著提升最优传输距离的准确性,进而改善嵌入、聚类和分类表现。我们在涵盖多种疾病的患者级scRNA-seq数据上验证了该方法的有效性与可解释性。
原文摘要 · Abstract (English)
Optimal transport provides a robust framework for comparing probability distributions. Its effectiveness is significantly influenced by the choice of the underlying ground metric. Traditionally, the ground metric has either been (i) predefined, e.g., as the Euclidean distance, or (ii) learned in a supervised way, by utilizing labeled data to learn a suitable ground metric for enhanced task-specific performance. Yet, predefined metrics typically cannot account for the inherent structure and varying importance of different features in the data, and existing supervised approaches to ground metric learning often do not generalize across multiple classes or are restricted to distributions with shared supports. To address these limitations, we propose a novel approach for learning metrics for arbitrary distributions over a shared metric space. Our method provides a distance between individual points like a global metric, but requires only class labels on a distribution-level for training. The learned global ground metric enables more accurate optimal transport distances, leading to improved performance in embedding, clustering and classification tasks. We demonstrate the effectiveness and interpretability of our approach using patient-level scRNA-seq data spanning multiple diseases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。