针对多模态图数据噪声问题,提出双图滤波与跨模态对比学习新方法。
Cross-Contrastive Clustering for Multimodal Attributed Graphs with Dual Graph Filtering
- 设计双图滤波机制,分离特征级噪声并增强节点表示
- 通过跨模态、邻域、社区三重对比学习提升聚类性能
- 在8个基准数据集上显著优于现有方法,适合复杂图数据聚类
多模态属性图(MMAG)是一种用于表示多源异构实体间复杂关联的数据模型,广泛应用于社交社区发现、医疗数据分析等场景。然而,现有方法过度依赖多视图属性间的高相关性,忽视了预训练语言与视觉模型生成的属性中普遍存在的低视图内相关性与强特征噪声问题,导致聚类效果不佳。本文基于图信号处理理论分析,提出双图滤波(DGF)方案,创新性地引入特征级去噪模块,有效克服传统图滤波的局限。在此基础上,DGF采用三重跨对比学习策略,实现模态间、邻域间与社区间的实例级对比学习,以获得更鲁棒、更具区分性的节点表示。在8个基准MMAG数据集上的全面实验表明,DGF在聚类质量上持续且显著优于多种前沿基线方法。
原文摘要 · Abstract (English)
Multimodal Attributed Graphs (MMAGs) are an expressive data model for representing the complex interconnections among entities that associate attributes from multiple data modalities (text, images, etc.). Clustering over such data finds numerous practical applications in real scenarios, including social community detection, medical data analytics, etc. However, as revealed by our empirical studies, existing multi-view clustering solutions largely rely on the high correlation between attributes across various views and overlook the unique characteristics (e.g., low modality-wise correlation and intense feature-wise noise) of multimodal attributes output by large pre-trained language and vision models in MMAGs, leading to suboptimal clustering performance. Inspired by foregoing empirical observations and our theoretical analyses with graph signal processing, we propose the Dual Graph Filtering (DGF) scheme, which innovatively incorporates a feature-wise denoising component into node representation learning, thereby effectively overcoming the limitations of traditional graph filters adopted in the extant multi-view graph clustering approaches. On top of that, DGF includes a tri-cross contrastive training strategy that employs instance-level contrastive learning across modalities, neighborhoods, and communities for learning robust and discriminative node representations. Our comprehensive experiments on eight benchmark MMAG datasets exhibit that DGF is able to outperform a wide range of state-of-the-art baselines consistently and significantly in terms of clustering quality measured against ground-truth labels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。