同时学习样本与特征的层次结构,提升数据表示效果。
Joint Hierarchical Representation Learning of Samples and Features via Informed Tree-Wasserstein Distance
- 交替构建样本与特征的树结构,利用树-沃尔瑟斯坦距离协同优化。
- 在词-文档和单细胞数据上优于基线方法,提升稀疏近似与无监督距离学习性能。
- 可作为超曲面图卷积网络的预处理,助力链接预测与节点分类任务。
高维数据在样本和特征两个维度上常呈现层次结构,但现有方法多仅关注单一维度。本文提出一种基于树-沃尔瑟斯坦距离(TWD)的无监督联合层次表示学习方法,通过在两个数据模式间交替迭代:先为一模式构建树,再基于该树计算另一模式的TWD,进而用此距离构建第二模式的树。重复此过程可逐步优化两棵树及对应的TWD,捕捉数据的有意义层次表示。理论分析表明方法收敛。将该方法集成至超曲面图卷积网络中作为预处理,显著提升链接预测与节点分类性能;在词-文档与单细胞RNA测序数据集上,其在稀疏近似与无监督沃尔瑟斯坦距离学习任务中均优于基线方法。
原文摘要 · Abstract (English)
High-dimensional data often exhibit hierarchical structures in both modes: samples and features. Yet, most existing approaches for hierarchical representation learning consider only one mode at a time. In this work, we propose an unsupervised method for jointly learning hierarchical representations of samples and features via Tree-Wasserstein Distance (TWD). Our method alternates between the two data modes. It first constructs a tree for one mode, then computes a TWD for the other mode based on that tree, and finally uses the resulting TWD to build the second mode's tree. By repeatedly alternating through these steps, the method gradually refines both trees and the corresponding TWDs, capturing meaningful hierarchical representations of the data. We provide a theoretical analysis showing that our method converges. We show that our method can be integrated into hyperbolic graph convolutional networks as a pre-processing technique, improving performance in link prediction and node classification tasks. In addition, our method outperforms baselines in sparse approximation and unsupervised Wasserstein distance learning tasks on word-document and single-cell RNA-sequencing datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。