研究节点特征异质性对图神经网络链接预测的影响,揭示了新优化方向。
On the Impact of Feature Heterophily on Link Prediction with Graph Neural Networks
- 提出同质/异质链接预测的正式定义与理论框架
- 实证发现异质性会降低传统GNN性能,需改进解码器和消息传递机制
- 适合研究图学习泛化性或设计鲁棒链接预测模型的研究者
异质性(即相连节点具有不同类别标签或不相似特征)被证实是图神经网络(GNN)面临的重要挑战。尽管在类别标签强异质性时进行节点分类的困难已有深入理解,但当类别标签不可用时,异质性对其他重要图学习任务的影响仍不清楚。本文聚焦链接预测任务,系统分析节点特征异质性对GNN性能的影响。理论上,首次提出同质与异质链接预测任务的正式定义,并构建理论框架,揭示两类任务所需的不同优化策略。进一步分析不同编码器与解码器在异质性水平变化下的适应能力,提出改进设计。在多种合成及真实世界数据集上的实验验证了理论洞察,强调在消息传递中采用可学习解码器及分离自环与邻居嵌入的编码器对提升链接预测性能的重要性。
原文摘要 · Abstract (English)
Heterophily, or the tendency of connected nodes in networks to have different class labels or dissimilar features, has been identified as challenging for many Graph Neural Network (GNN) models. While the challenges of applying GNNs for node classification when class labels display strong heterophily are well understood, it is unclear how heterophily affects GNN performance in other important graph learning tasks where class labels are not available. In this work, we focus on the link prediction task and systematically analyze the impact of heterophily in node features on GNN performance. Theoretically, we first introduce formal definitions of homophilic and heterophilic link prediction tasks, and present a theoretical framework that highlights the different optimizations needed for the respective tasks. We then analyze how different link prediction encoders and decoders adapt to varying levels of feature homophily and introduce designs for improved performance. Our empirical analysis on a variety of synthetic and real-world datasets confirms our theoretical insights and highlights the importance of adopting learnable decoders and GNN encoders with ego- and neighbor-embedding separation in message passing for link prediction tasks beyond homophily.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。