arXiv:2502.10706cs.LGcs.AI2025-02TPAMI被引 6

提出新方法提升图数据分布外泛化能力,解决环境建模难与类别混淆问题。

Raising the Bar in Graph OOD Generalization: Invariant Learning Beyond Explicit Environment Modeling

  • 用超球面特征提取和多原型分类,避免显式建模环境
  • 在11个基准数据集上达到最优性能,显著超越现有方法
  • 适合研究图学习泛化与鲁棒性的人群参考

图数据的分布外(OOD)泛化已成为关键挑战,因真实世界图数据常呈现多样且动态变化的环境,传统模型难以跨域泛化。一种有前景的解决方案是图不变学习(GIL),旨在通过分离与标签相关联的不变子图与环境特异子图来学习不变表示。然而现有GIL方法面临两大难题:(1)难以捕捉和建模图数据中多样的环境;(2)语义悬崖问题——不同类别的不变子图难以区分,导致类别可分性差、误分类增多。为此,本文提出多原型超球面不变学习(MPHIL),引入两项核心创新:(1)超球面不变特征提取,实现鲁棒且高判别力的特征表示;(2)多原型超球面分类,以类别原型作为中间变量,无需显式环境建模,缓解语义悬崖。基于GIL理论框架,设计两个新目标函数:不变原型匹配损失,确保样本与正确类别原型对齐;原型分离损失,增强不同类别原型在超球面上的区分度。在11个图数据分布外泛化基准数据集上的大量实验表明,MPHIL在多种领域及分布偏移场景下均取得当前最佳表现。

原文摘要 · Abstract (English)

Out-of-distribution (OOD) generalization has emerged as a critical challenge in graph learning, as real-world graph data often exhibit diverse and shifting environments that traditional models fail to generalize across. A promising solution to address this issue is graph invariant learning (GIL), which aims to learn invariant representations by disentangling label-correlated invariant subgraphs from environment-specific subgraphs. However, existing GIL methods face two major challenges: (1) the difficulty of capturing and modeling diverse environments in graph data, and (2) the semantic cliff, where invariant subgraphs from different classes are difficult to distinguish, leading to poor class separability and increased misclassifications. To tackle these challenges, we propose a novel method termed Multi-Prototype Hyperspherical Invariant Learning (MPHIL), which introduces two key innovations: (1) hyperspherical invariant representation extraction, enabling robust and highly discriminative hyperspherical invariant feature extraction, and (2) multi-prototype hyperspherical classification, which employs class prototypes as intermediate variables to eliminate the need for explicit environment modeling in GIL and mitigate the semantic cliff issue. Derived from the theoretical framework of GIL, we introduce two novel objective functions: the invariant prototype matching loss to ensure samples are matched to the correct class prototypes, and the prototype separation loss to increase the distinction between prototypes of different classes in the hyperspherical space. Extensive experiments on 11 OOD generalization benchmark datasets demonstrate that MPHIL achieves state-of-the-art performance, significantly outperforming existing methods across graph data from various domains and with different distribution shifts.

图学习泛化能力不变学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。