arXiv:2602.08855cs.LG2026-02

提出新方法提升图神经网络在分布外数据上的稳定性。

Rethinking Graph Generalization through the Lens of Sharpness-Aware Minimization

  • 用局部鲁棒半径衡量损失曲面平坦性,揭示模型变尖导致误判。
  • 训练中鲁棒半径持续下降,暴露模型稳定性退化问题。
  • 设计能量驱动增强框架,有效提升图模型泛化能力。

图神经网络在各类图任务中表现优异,但对分布偏移敏感。本文关注一种常见却未被充分研究的现象——最小扰动翻转(MSF),即测试样本仅轻微偏离训练分布时即被错误分类。通过锐度感知最小化(SAM)视角,我们引入局部鲁棒半径来量化使预测翻转所需的最小扰动,建立局部稳定性与泛化性能的理论联系。实验发现,训练过程中鲁棒半径持续减小,表明损失曲面日益尖锐,引发MSF。为此,我们提出能量基公式,其与鲁棒半径单调相关,可作为可计算的平坦性目标。基于此,构建能量驱动生成增强框架(E2A),利用能量引导的潜在扰动生成伪分布外样本,提升模型泛化。大量实验证明,E2A在多个基准上持续优于现有先进方法。

原文摘要 · Abstract (English)

Graph Neural Networks (GNNs) have achieved remarkable success across various graph-based tasks but remain highly sensitive to distribution shifts. In this work, we focus on a prevalent yet under-explored phenomenon in graph generalization, Minimal Shift Flip (MSF),where test samples that slightly deviate from the training distribution are abruptly misclassified. To interpret this phenomenon, we revisit MSF through the lens of Sharpness-Aware Minimization (SAM), which characterizes the local stability and sharpness of the loss landscape while providing a theoretical foundation for modeling generalization error. To quantify loss sharpness, we introduce the concept of Local Robust Radius, measuring the smallest perturbation required to flip a prediction and establishing a theoretical link between local stability and generalization. Building on this perspective, we further observe a continual decrease in the robust radius during training, indicating weakened local stability and an increasingly sharp loss landscape that gives rise to MSF. To jointly solve the MSF phenomenon and the intractability of radius, we develop an energy-based formulation that is theoretically proven to be monotonically correlated with the robust radius, offering a tractable and principled objective for modeling flatness and stability. Building on these insights, we propose an energy-driven generative augmentation framework (E2A) that leverages energy-guided latent perturbations to generate pseudo-OOD samples and enhance model generalization. Extensive experiments across multiple benchmarks demonstrate that E2A consistently improves graph OOD generalization, outperforming state-of-the-art baselines.

图神经网络泛化能力稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。