提出通用图注意力架构,揭示高效图神经网络设计规律。
Generalizable Insights for Graph Transformers in Theory and Practice
- 构建统一的广义距离变换器,整合近年图注意力关键改进
- 在超800万张图上验证,少样本迁移无需微调即表现优异
- 提炼出跨领域适用的图变换器设计与训练通用准则
图变换器(GTs)表现出强大的经验性能,但现有架构在注意力机制、位置嵌入(PEs)和表达能力方面差异显著。现有表达能力结果常依赖特定设计,且缺乏大规模数据的全面实证验证,导致理论与实践之间存在鸿沟。本文提出广义距离变换器(GDT),采用标准注意力机制,融合近年来图变换器的多项进展,并深入分析其在注意力与位置嵌入方面的表征能力。通过大量实验,我们识别出在多种应用、任务和模型规模下均表现稳定的架构设计选择,在无需微调的少样本迁移设置中表现强劲。评估涵盖超过八百万张图、约2.7亿个标记,覆盖图像目标检测、分子性质预测、代码摘要及分布外算法推理等多个领域。我们将理论与实践发现凝练为若干可泛化的图变换器设计、训练与推理洞见。
原文摘要 · Abstract (English)
Graph Transformers (GTs) have shown strong empirical performance, yet current architectures vary widely in their use of attention mechanisms, positional embeddings (PEs), and expressivity. Existing expressivity results are often tied to specific design choices and lack comprehensive empirical validation on large-scale data. This leaves a gap between theory and practice, preventing generalizable insights that exceed particular application domains. Here, we propose the Generalized-Distance Transformer (GDT), a GT architecture using standard attention that incorporates many advancements for GTs from recent years, and develop a fine-grained understanding of the GDT's representation power in terms of attention and PEs. Through extensive experiments, we identify design choices that consistently perform well across various applications, tasks, and model scales, demonstrating strong performance in a few-shot transfer setting without fine-tuning. Our evaluation covers over eight million graphs with roughly 270M tokens across diverse domains, including image-based object detection, molecular property prediction, code summarization, and out-of-distribution algorithmic reasoning. We distill our theoretical and practical findings into several generalizable insights about effective GT design, training, and inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。