通过多视角对比学习提升异构图表示能力,无需标签也能超越部分有监督方法。
Incorporating Attributes and Multi-Scale Structures for Heterogeneous Graph Contrastive Learning
- 融合属性、高阶与低阶结构三视图进行对比学习
- 在四个真实数据集上超越主流无监督模型,部分超过有监督基线
- 提出属性增强正样本选择策略,缓解采样偏差问题
异构图(HGs)由多种节点和边类型构成,能更有效捕捉现实世界中复杂的关联结构。然而,在实际应用中,标注数据往往难以获取,限制了半监督方法的使用。自监督学习可通过自动从数据中学习有用特征,有效应对标注数据不足的问题。本文提出一种新的异构图对比学习框架(ASHGCL),包含三个不同视角:分别关注节点属性、高阶与低阶结构信息,以有效捕获节点的属性特征、高阶结构与低阶结构。此外,引入属性增强的正样本选择策略,结合结构与属性信息,有效缓解采样偏差问题。在四个真实数据集上的大量实验表明,ASHGCL显著优于现有无监督基线,并在某些情况下超越部分有监督基准。
原文摘要 · Abstract (English)
Heterogeneous graphs (HGs) are composed of multiple types of nodes and edges, making it more effective in capturing the complex relational structures inherent in the real world. However, in real-world scenarios, labeled data is often difficult to obtain, which limits the applicability of semi-supervised approaches. Self-supervised learning aims to enable models to automatically learn useful features from data, effectively addressing the challenge of limited labeling data. In this paper, we propose a novel contrastive learning framework for heterogeneous graphs (ASHGCL), which incorporates three distinct views, each focusing on node attributes, high-order and low-order structural information, respectively, to effectively capture attribute information, high-order structures, and low-order structures for node representation learning. Furthermore, we introduce an attribute-enhanced positive sample selection strategy that combines both structural information and attribute information, effectively addressing the issue of sampling bias. Extensive experiments on four real-world datasets show that ASHGCL outperforms state-of-the-art unsupervised baselines and even surpasses some supervised benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。