arXiv:2510.02731cs.LG2025-10NeurIPS被引 2

提出新方法同时增强节点与边的表示,提升图聚类鲁棒性。

Hybrid-Collaborative Augmentation and Contrastive Sample Adaptive-Differential Awareness for Robust Attributed Graph Clustering

  • 同时对节点和边进行混合协同增强,构建更全面的相似性度量。
  • 通过自适应加权机制区分难易样本对,显著提升判别能力。
  • 在六个基准数据集上优于现有先进方法,适合复杂图结构聚类任务。

由于具备强大的自监督表征学习与聚类能力,对比属性图聚类(CAGC)取得了显著进展,主要依赖于有效的数据增强和对比目标设置。然而,多数CAGC方法仅利用边作为辅助信息获取节点级嵌入表示,并且仅关注节点级嵌入增强,忽略了边级嵌入增强以及节点与边级嵌入增强在不同粒度下的交互。此外,它们通常对所有对比样本对一视同仁,忽视了困难与简单正负样本对之间的显著差异,最终限制了其判别能力。为此,本文提出一种新型鲁棒属性图聚类(RAGC),融合混合协同增强(HCA)与对比样本自适应差分感知(CSADA)。首先,同时执行节点级与边级嵌入表示及增强,建立更全面的相似性度量标准以支持后续对比学习;反之,判别性相似性进一步有意识地引导边增强。其次,利用高置信度伪标签信息,设计了自适应识别所有对比样本对并采用创新加权函数差异化处理的策略。HCA与CSADA模块形成良性循环相互促进,从而增强表征学习的判别性。在六个基准数据集上的综合图聚类评估表明,所提RAGC方法优于多种先进CAGC方法。

原文摘要 · Abstract (English)

Due to its powerful capability of self-supervised representation learning and clustering, contrastive attributed graph clustering (CAGC) has achieved great success, which mainly depends on effective data augmentation and contrastive objective setting. However, most CAGC methods utilize edges as auxiliary information to obtain node-level embedding representation and only focus on node-level embedding augmentation. This approach overlooks edge-level embedding augmentation and the interactions between node-level and edge-level embedding augmentations across various granularity. Moreover, they often treat all contrastive sample pairs equally, neglecting the significant differences between hard and easy positive-negative sample pairs, which ultimately limits their discriminative capability. To tackle these issues, a novel robust attributed graph clustering (RAGC), incorporating hybrid-collaborative augmentation (HCA) and contrastive sample adaptive-differential awareness (CSADA), is proposed. First, node-level and edge-level embedding representations and augmentations are simultaneously executed to establish a more comprehensive similarity measurement criterion for subsequent contrastive learning. In turn, the discriminative similarity further consciously guides edge augmentation. Second, by leveraging pseudo-label information with high confidence, a CSADA strategy is elaborately designed, which adaptively identifies all contrastive sample pairs and differentially treats them by an innovative weight modulation function. The HCA and CSADA modules mutually reinforce each other in a beneficent cycle, thereby enhancing discriminability in representation learning. Comprehensive graph clustering evaluations over six benchmark datasets demonstrate the effectiveness of the proposed RAGC against several state-of-the-art CAGC methods.

图聚类对比学习嵌入增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。