arXiv:2502.08105cs.LG2025-02综述被引 10

梳理图数据分布外检测的分类与挑战,助力模型应对真实场景分布偏移。

Out-of-Distribution Detection on Graphs: A Survey

  • 按增强、重建、信息传播、分类四类系统归纳图OOD检测方法。
  • 指出图数据分布偏移会严重降低模型性能,需主动识别异常图结构。
  • 适合关注图学习鲁棒性与安全性的研究者,尤其在工业部署中实用。

图机器学习发展迅速,但实际应用中训练与测试数据分布常不一致,导致模型性能下降。为应对这一问题,图分布外(GOOD)检测成为研究热点,旨在识别偏离训练分布的图数据,提升模型鲁棒性。本文首次严格定义GOOD检测,将现有方法系统分为四类:增强型、重建型、信息传播型和分类型,并分析其原理与机制。同时厘清了GOOD检测与图异常检测、离群点检测及分布外泛化等领域的区别。此外,探讨了实际应用场景与理论基础,揭示图数据特有的挑战。最后,总结主要难点并提出未来研究方向。相关资源可在 https://github.com/ca1man-2022/Awesome-GOOD-Detection 获取。

原文摘要 · Abstract (English)

Graph machine learning has witnessed rapid growth, driving advancements across diverse domains. However, the in-distribution assumption, where training and testing data share the same distribution, often breaks in real-world scenarios, leading to degraded model performance under distribution shifts. This challenge has catalyzed interest in graph out-of-distribution (GOOD) detection, which focuses on identifying graph data that deviates from the distribution seen during training, thereby enhancing model robustness. In this paper, we provide a rigorous definition of GOOD detection and systematically categorize existing methods into four types: enhancement-based, reconstruction-based, information propagation-based, and classification-based approaches. We analyze the principles and mechanisms of each approach and clarify the distinctions between GOOD detection and related fields, such as graph anomaly detection, outlier detection, and GOOD generalization. Beyond methodology, we discuss practical applications and theoretical foundations, highlighting the unique challenges posed by graph data. Finally, we discuss the primary challenges and propose future directions to advance this emerging field. The repository of this survey is available at https://github.com/ca1man-2022/Awesome-GOOD-Detection.

图学习分布外检测鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。