图神经网络在分布外场景下表现更优,图注意力模型比传统方法更具泛化能力。
Exploring Graph-Transformer Out-of-Distribution Generalization Abilities
- 对比图注意力与消息传递模型,评估其在分布外场景下的性能差异。
- 四组基准测试中,图注意力模型无需特殊算法即表现更优。
- 提出新分析方法,可可视化不同数据域的聚类结构和分类分离度。
图深度学习在社交网络、生物物理、交通网络和推荐系统等众多领域取得显著成功。然而,现有方法普遍依赖训练与测试数据同分布的假设,这在真实场景中难以满足。尽管图变换器(GT)在多个分布内(ID)基准上已超越传统消息传递神经网络(MPNN),但其在分布偏移下的泛化能力仍不明确。本文系统评估了GT与混合型骨干网络在分布外(OOD)设置下的表现,并与MPNN进行对比。为此,我们适配多种主流领域泛化(DG)算法用于GT,并在一套涵盖多种分布偏移的基准上测试其性能。结果表明,即使未使用专门的DG算法,GT及混合型GT-MPNN骨干网络在六项基准中的四项上也展现出更强的泛化能力。此外,我们提出一种新型后训练分析方法,通过比较完整分布内与分布外测试集的聚类结构,具体考察领域对齐与类别分离情况。该方法具有模型无关性,为理解GT与MPNN的泛化能力提供了超越准确率的新视角,适用于更广泛的领域泛化任务。研究揭示了图变换器在真实世界图学习中的潜力,并为未来分布外泛化研究指明新方向。
原文摘要 · Abstract (English)
Deep learning on graphs has shown remarkable success across numerous applications, including social networks, bio-physics, traffic networks, and recommendation systems. Regardless of their successes, current methods frequently depend on the assumption that training and testing data share the same distribution, a condition rarely met in real-world scenarios. While graph-transformer (GT) backbones have recently outperformed traditional message-passing neural networks (MPNNs) in multiple in-distribution (ID) benchmarks, their effectiveness under distribution shifts remains largely unexplored. In this work, we address the challenge of out-of-distribution (OOD) generalization for graph neural networks, with a special focus on the impact of backbone architecture. We systematically evaluate GT and hybrid backbones in OOD settings and compare them to MPNNs. To do so, we adapt several leading domain generalization (DG) algorithms to work with GTs and assess their performance on a benchmark designed to test a variety of distribution shifts. Our results reveal that GT and hybrid GT-MPNN backbones demonstrate stronger generalization ability compared to MPNNs, even without specialized DG algorithms (on four out of six benchmarks). Additionally, we propose a novel post-training analysis approach that compares the clustering structure of the entire ID and OOD test datasets, specifically examining domain alignment and class separation. Highlighting its model-agnostic design, the method yielded valuable insights into both GT and MPNN backbones and appears well suited for broader DG applications beyond graph learning, offering a deeper perspective on generalization abilities that goes beyond standard accuracy metrics. Together, our findings highlight the promise of graph-transformers for robust, real-world graph learning and set a new direction for future research in OOD generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。