arXiv:2603.17569stat.MLcs.LG2026-03被引 1

理论揭示图Transformer比图卷积更抗过平滑,保持社区结构

Gaussian Process Limit Reveals Structural Benefits of Graph Transformers

  • 通过高斯过程极限分析图Transformer的深层传播机制
  • 证明其能保留社区信息,深度层仍具区分性节点表示
  • 适合研究图神经网络理论或设计深层模型的读者

图Transformer在图结构数据学习中表现优异,但其理论机制尚不清晰。本文研究了无限宽度与无限头数下图Transformer(GAT、Graphormer、Specformer)的神经网络高斯过程极限,推导出各层的节点级与边级核函数。结果揭示了节点特征与图结构在注意力层中的传播规律。以具体案例证明,图Transformer能结构性地保持社区信息,即使在深层也维持可区分的节点表示,从而避免过平滑。在合成与真实世界图上提供了实证支持,表明融入先验信息与位置编码可提升深层图Transformer性能。

原文摘要 · Abstract (English)

Graph transformers are the state-of-the-art for learning from graph-structured data and are empirically known to avoid several pitfalls of message-passing architectures. However, there is limited theoretical analysis on why these models perform well in practice. In this work, we prove that attention-based architectures have structural benefits over graph convolutional networks in the context of node-level prediction tasks. Specifically, we study the neural network gaussian process limits of graph transformers (GAT, Graphormer, Specformer) with infinite width and infinite heads, and derive the node-level and edge-level kernels across the layers. Our results characterise how the node features and the graph structure propagate through the graph attention layers. As a specific example, we prove that graph transformers structurally preserve community information and maintain discriminative node representations even in deep layers, thereby preventing oversmoothing. We provide empirical evidence on synthetic and real-world graphs that validate our theoretical insights, such as integrating informative priors and positional encoding can improve performance of deep graph transformers.

图神经网络注意力机制理论分析过平滑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。