图预训练模型能高效检测异常,尤其在标注少时表现超越现有方法。
Graph Pre-Training Models Are Strong Anomaly Detectors
- 利用图预训练学习全局结构,不依赖端到端微调
- 在少量标注下超越最新端到端模型,对远距离异常检测更优
- 适合数据稀缺或需发现隐蔽异常的场景
图异常检测(GAD)是极具挑战性且实用的研究方向,图神经网络(GNN)近年展现出良好效果。现有GNN的效果主要源于端到端联合学习节点表征与分类器。而图预训练(如DGI、GraphMAE等两阶段范式)虽在利用无标签图数据提升下游任务方面显示潜力,其在GAD中的作用仍待深入探索。本文首次揭示:图预训练模型本身即具备强异常检测能力。具体而言,在监督信息有限时,预训练模型显著优于当前最优的端到端训练模型。进一步分析表明,预训练能有效增强对远离已知异常点、未充分表示的无标签异常的检测能力,这些异常超出已知异常2跳邻域范围,解释了其性能优势。此外,我们还拓展至图级异常检测的可行性分析。本研究呼吁重新评估预训练在GAD中的价值,为未来研究提供关键洞见。
原文摘要 · Abstract (English)
Graph Anomaly Detection (GAD) is a challenging and practical research topic where Graph Neural Networks (GNNs) have recently shown promising results. The effectiveness of existing GNNs in GAD has been mainly attributed to the simultaneous learning of node representations and the classifier in an end-to-end manner. Meanwhile, graph pre-training, the two-stage learning paradigm such as DGI and GraphMAE, has shown potential in leveraging unlabeled graph data to enhance downstream tasks, yet its impact on GAD remains under-explored. In this work, we show that graph pre-training models are strong graph anomaly detectors. Specifically, we demonstrate that pre-training is highly competitive, markedly outperforming the state-of-the-art end-to-end training models when faced with limited supervision. To understand this phenomenon, we further uncover pre-training enhances the detection of distant, under-represented, unlabeled anomalies that go beyond 2-hop neighborhoods of known anomalies, shedding light on its superior performance against end-to-end models. Moreover, we extend our examination to the potential of pre-training in graph-level anomaly detection. We envision this work to stimulate a re-evaluation of pre-training's role in GAD and offer valuable insights for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。