arXiv:2410.20295cs.LG2024-10被引 3

提出因果解耦框架DeCaf,提升图神经网络在分布外数据上的泛化能力

DeCaf: A Causal Decoupling Framework for OOD Generalization on Node Classification

  • 基于结构因果模型重构图数据生成过程,精准定位分布偏移源头
  • 独立学习特征与结构的无偏映射,有效缓解多种分布偏移影响
  • 在真实与合成数据上验证,适合高安全要求场景的图学习任务

图神经网络(GNN)对分布偏移敏感,在关键领域存在脆弱性和安全隐患。现有方法通常依赖对数据生成过程的过度简化假设,难以反映图数据中分布偏移的真实动态。本文引入结构因果模型(SCMs)构建更现实的图数据生成模型,明确界定分布偏移的来源。基于此,提出因果解耦框架DeCaf,独立学习无偏的特征-标签和结构-标签映射。我们提供了详细的理论分析,证明该方法可有效缓解多种分布偏移的影响。在真实世界和合成数据集上的实验验证了DeCaf在提升GNN泛化性能方面的有效性。

原文摘要 · Abstract (English)

Graph Neural Networks (GNNs) are susceptible to distribution shifts, creating vulnerability and security issues in critical domains. There is a pressing need to enhance the generalizability of GNNs on out-of-distribution (OOD) test data. Existing methods that target learning an invariant (feature, structure)-label mapping often depend on oversimplified assumptions about the data generation process, which do not adequately reflect the actual dynamics of distribution shifts in graphs. In this paper, we introduce a more realistic graph data generation model using Structural Causal Models (SCMs), allowing us to redefine distribution shifts by pinpointing their origins within the generation process. Building on this, we propose a casual decoupling framework, DeCaf, that independently learns unbiased feature-label and structure-label mappings. We provide a detailed theoretical framework that shows how our approach can effectively mitigate the impact of various distribution shifts. We evaluate DeCaf across both real-world and synthetic datasets that demonstrate different patterns of shifts, confirming its efficacy in enhancing the generalizability of GNNs.

图神经网络因果学习分布外泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。