arXiv:2505.17293cs.LG2025-05NeurIPS被引 7

不依赖模型,用最优传输选数据提升图迁移学习效果

Graph Data Selection for Domain Adaptation: A Model-Free Approach

  • 基于最优传输理论选择源域中对目标域最有帮助的训练数据
  • 在多种分布偏移下显著提升分类性能,且只需更少训练数据
  • 可与现有模型方法配合使用,适合资源受限场景

图领域自适应(GDA)是图机器学习中的基础任务,现有方法如鲁棒图神经网络(GNN)和专用训练流程虽有效,但在严重分布偏移和计算资源受限时表现不佳。为此,我们提出一种新型无模型框架GRADATE(GRAph DATa sElector),从源域中选择最适合作为目标域分类任务的训练样本。GRADATE无需依赖任何GNN的预测或训练方案,利用最优传输理论捕捉并适应分布变化。该方法数据高效、可扩展,并能与现有模型中心的GDA方法互补。在多个真实世界图级数据集和多种协变量偏移类型上的实验表明,GRADATE优于现有选择方法,且能以更少训练数据显著增强现成的GDA方法性能。

原文摘要 · Abstract (English)

Graph domain adaptation (GDA) is a fundamental task in graph machine learning, with techniques like shift-robust graph neural networks (GNNs) and specialized training procedures to tackle the distribution shift problem. Although these model-centric approaches show promising results, they often struggle with severe shifts and constrained computational resources. To address these challenges, we propose a novel model-free framework, GRADATE (GRAph DATa sElector), that selects the best training data from the source domain for the classification task on the target domain. GRADATE picks training samples without relying on any GNN model's predictions or training recipes, leveraging optimal transport theory to capture and adapt to distribution changes. GRADATE is data-efficient, scalable and meanwhile complements existing model-centric GDA approaches. Through comprehensive empirical studies on several real-world graph-level datasets and multiple covariate shift types, we demonstrate that GRADATE outperforms existing selection methods and enhances off-the-shelf GDA methods with much fewer training data.

图神经网络领域自适应数据选择最优传输

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。