arXiv:2511.16778cs.LG2025-11AAAI

提出GCL-OT框架,解决异质文本图中结构与文本对齐难题。

GCL-OT: Graph Contrastive Learning with Optimal Transport for Heterophilic Text-Attributed Graphs

  • 引入最优传输机制,分类型处理完全、部分和潜在同质性异质性。
  • 在9个基准上超越现有方法,显著提升异质图分类性能。
  • 适合处理复杂异质性文本图的科研人员与工业应用开发者。

近期,结构-文本对比学习在文本属性图上展现出良好性能,通过融合图神经网络与语言模型的优势。然而,现有方法通常依赖同质性假设进行相似性估计和硬优化目标,限制了其在异质图上的适用性。尽管已有方法通过结构调整或邻居聚合缓解异质性,但通常将文本嵌入视为静态目标,导致对齐效果不佳。本文识别出文本属性图中存在的多粒度异质性,包括完全异质、部分异质和潜在同质,使得结构-文本对齐面临混合、噪声和缺失语义关联的挑战。为此,我们提出GCL-OT,一种基于最优传输的新型图对比学习框架,针对不同异质类型设计相应机制:针对部分异质,采用RealSoftMax相似性估计器,强化关键邻域-词语交互并抑制背景噪声;针对完全异质,引入提示引导过滤器,在最优传输对齐中自适应剔除无关噪声;此外,结合OT引导的软监督,发现具有相似语义的潜在邻居,增强潜在同质性的学习。理论分析表明,GCL-OT可提升互信息界和贝叶斯误差保证。在九个基准数据集上的大量实验表明,GCL-OT优于当前最先进方法,验证了其有效性与鲁棒性。

原文摘要 · Abstract (English)

Recently, structure-text contrastive learning has shown promising performance on text-attributed graphs by leveraging the complementary strengths of graph neural networks and language models. However, existing methods typically rely on homophily assumptions in similarity estimation and hard optimization objectives, which limit their applicability to heterophilic graphs. Although existing methods can mitigate heterophily through structural adjustments or neighbor aggregation, they usually treat textual embeddings as static targets, leading to suboptimal alignment. In this work, we identify multi-granular heterophily in text-attributed graphs, including complete heterophily, partial heterophily, and latent homophily, which makes structure-text alignment particularly challenging due to mixed, noisy, and missing semantic correlations. To achieve flexible and bidirectional alignment, we propose GCL-OT, a novel graph contrastive learning framework with optimal transport, equipped with tailored mechanisms for each type of heterophily. Specifically, for partial heterophily, we design a RealSoftMax-based similarity estimator to emphasize key neighbor-word interactions while easing background noise. For complete heterophily, we introduce a prompt-based filter that adaptively excludes irrelevant noise during optimal transport alignment. Furthermore, we incorporate OT-guided soft supervision to uncover potential neighbors with similar semantics, enhancing the learning of latent homophily. Theoretical analysis shows that GCL-OT can improve the mutual information bound and Bayes error guarantees. Extensive experiments on nine benchmarks show that GCL-OT outperforms state-of-the-art methods, demonstrating its effectiveness and robustness.

图对比学习最优传输异质图文本属性图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。