arXiv:2503.07852cs.LGcs.AI2025-03中稿 · the WSDM 2025 Oral被引 7

利用条件独立性设计更优的图自监督掩码策略,提升表征学习效果。

CIMAGE: Exploiting the Conditional Independence in Masked Graph Auto-encoders

  • 基于条件独立性在隐空间生成两个上下文,通过伪标签引导掩码。
  • 在多个图基准上节点分类和链接预测平均排名显著提升。
  • 适合研究图自监督学习与表示学习的学者参考。

近期基于图神经网络中掩码的自监督学习方法在捕获关系信息方面表现优异。然而,多数方法依赖特征或图空间的随机掩码,难以充分捕捉任务相关特征。我们提出,该局限源于掩码与未掩码部分间冗余过高、与下游任务相关性不足。条件独立性(CI)天然满足最小冗余与最大相关性要求,但通常需依赖下游标签。为此,我们提出CIMAGE,利用条件独立性指导隐空间中的掩码策略。CIMAGE通过无监督图聚类获取高置信度伪标签,进行条件独立感知的隐因子分解,生成两个不同上下文;预训练任务是仅凭第一个上下文重建被掩码的第二个上下文。理论分析表明,所学嵌入具有近似线性可分性,有利于下游任务准确预测。在多个图基准上的全面评估显示,CIMAGE在节点分类和链接预测任务中取得更高平均排名,揭示了条件独立性在增强图自监督学习中的潜力,为有效图表示学习提供了新见解。

原文摘要 · Abstract (English)

Recent Self-Supervised Learning (SSL) methods encapsulating relational information via masking in Graph Neural Networks (GNNs) have shown promising performance. However, most existing approaches rely on random masking strategies in either feature or graph space, which may fail to capture task-relevant information fully. We posit that this limitation stems from an inability to achieve minimum redundancy between masked and unmasked components while ensuring maximum relevance of both to potential downstream tasks. Conditional Independence (CI) inherently satisfies the minimum redundancy and maximum relevance criteria, but its application typically requires access to downstream labels. To address this challenge, we introduce CIMAGE, a novel approach that leverages Conditional Independence to guide an effective masking strategy within the latent space. CIMAGE utilizes CI-aware latent factor decomposition to generate two distinct contexts, leveraging high-confidence pseudo-labels derived from unsupervised graph clustering. In this framework, the pretext task involves reconstructing the masked second context solely from the information provided by the first context. Our theoretical analysis further supports the superiority of CIMAGE's novel CI-aware masking method by demonstrating that the learned embedding exhibits approximate linear separability, which enables accurate predictions for the downstream task. Comprehensive evaluations across diverse graph benchmarks illustrate the advantage of CIMAGE, with notably higher average rankings on node classification and link prediction tasks. Notably, our proposed model highlights the under-explored potential of CI in enhancing graph SSL methodologies and offers enriched insights for effective graph representation learning.

图神经网络自监督学习条件独立表示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。