arXiv:2505.24178cs.LGcs.AI2025-05中稿 · AISTATS 2025被引 7

提出可区分不变与变化边的模型,提升时序图在分布外场景下的泛化能力

Invariant Link Selector for Spatial-Temporal Out-of-Distribution Problem

  • 基于信息瓶颈思想设计误差有界的不变边选择器
  • 在多个真实数据集上实现超越当前最优方法的推荐性能
  • 适用于学术引用、商品推荐等时序关系推理任务

在基础模型时代,分布外(OOD)问题——即训练环境与测试环境之间的数据差异——严重阻碍了AI的泛化能力。尤其对于违反独立同分布(IID)假设的关系型数据,如时序图,该问题更为严峻。为实现时序图上的鲁棒不变学习,本文探究哪些图结构成分对标签最具不变性与代表性。基于信息瓶颈(IB)方法,我们提出一种误差有界的不变边选择器,在训练过程中区分不变与变化成分,从而提升模型在不同测试场景下的泛化能力。此外,我们推导出一系列可推广的优化函数,并引入任务特定损失(如时序边预测),使预训练模型能解决实际应用任务,如论文引用推荐和商品推荐。实验表明,该方法在多个基准上达到当前最优(SOTA)性能。代码已公开于https://github.com/kthrn22/OOD-Linker。

原文摘要 · Abstract (English)

In the era of foundation models, Out-of- Distribution (OOD) problems, i.e., the data discrepancy between the training environments and testing environments, hinder AI generalization. Further, relational data like graphs disobeying the Independent and Identically Distributed (IID) condition makes the problem more challenging, especially much harder when it is associated with time. Motivated by this, to realize the robust invariant learning over temporal graphs, we want to investigate what components in temporal graphs are most invariant and representative with respect to labels. With the Information Bottleneck (IB) method, we propose an error-bounded Invariant Link Selector that can distinguish invariant components and variant components during the training process to make the deep learning model generalizable for different testing scenarios. Besides deriving a series of rigorous generalizable optimization functions, we also equip the training with task-specific loss functions, e.g., temporal link prediction, to make pretrained models solve real-world application tasks like citation recommendation and merchandise recommendation, as demonstrated in our experiments with state-of-the-art (SOTA) methods. Our code is available at https://github.com/kthrn22/OOD-Linker.

时序图分布外泛化不变学习推荐系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。