提出树-沃瑟斯坦距离,可从高维数据中学习隐含特征层次结构。
Tree-Wasserstein Distance for High Dimensional Data with a Latent Feature Hierarchy
- 通过扩散几何将特征嵌入多尺度双曲空间,构建树形结构
- 基于观测数据计算的TWD能准确恢复潜在特征层次结构
- 适用于文本、单细胞测序等高维数据,性能优于现有方法
在高维数据中寻找有意义的距离是重要的科学任务。为此,我们提出一种新的树-沃瑟斯坦距离(TWD),针对具有隐含特征层次结构的数据——即特征位于分层空间中,不同于通常将样本嵌入双曲空间的做法。同时,传统TWD主要用于加速沃瑟斯坦距离的计算,而我们利用其内在的树结构来学习隐含特征层次。核心思想是通过扩散几何将特征嵌入多尺度双曲空间,并建立双曲嵌入与树之间的类比关系,提出新的树解码方法。我们证明,基于数据观测计算的TWD能有效恢复由潜在特征层次定义的真值TWD,且计算高效可扩展。我们在词-文档和单细胞RNA测序数据集上展示了该方法的有效性,相较于现有TWD及基于预训练模型的方法具有明显优势。
原文摘要 · Abstract (English)
Finding meaningful distances between high-dimensional data samples is an important scientific task. To this end, we propose a new tree-Wasserstein distance (TWD) for high-dimensional data with two key aspects. First, our TWD is specifically designed for data with a latent feature hierarchy, i.e., the features lie in a hierarchical space, in contrast to the usual focus on embedding samples in hyperbolic space. Second, while the conventional use of TWD is to speed up the computation of the Wasserstein distance, we use its inherent tree as a means to learn the latent feature hierarchy. The key idea of our method is to embed the features into a multi-scale hyperbolic space using diffusion geometry and then present a new tree decoding method by establishing analogies between the hyperbolic embedding and trees. We show that our TWD computed based on data observations provably recovers the TWD defined with the latent feature hierarchy and that its computation is efficient and scalable. We showcase the usefulness of the proposed TWD in applications to word-document and single-cell RNA-sequencing datasets, demonstrating its advantages over existing TWDs and methods based on pre-trained models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。