arXiv:2501.14694cs.LGcs.AI2025-01中稿 · Data Mining and Kn…被引 2

提出自动选择自监督学习超参数的方法,避免用标签信息导致性能虚高。

Towards Automated Self-Supervised Learning for Truly Unsupervised Graph Anomaly Detection

  • 基于内部评估策略自动选择自监督学习超参数
  • 实验证明现有方法因随意选参导致性能偏差
  • 适合追求公平评估的图异常检测研究者

自监督学习(SSL)通过数据自身生成监督信号,近年被广泛用于图异常检测。但我们发现三个关键因素显著影响不同数据集上的检测效果:1)采用的具体SSL策略;2)策略超参数的调优;3)多策略组合时的权重分配。现有方法常任意或依赖标签信息选择策略、参数与权重,这不仅可能导致性能不佳,更在无监督场景中引入标签泄露,严重夸大模型表现。标签泄露已被列为十大数据挖掘错误之一,但许多近期研究仍如此操作。为此,我们提出一种基于理论分析的内部评估策略,用于无监督异常检测中的SSL超参数选择。我们在10个主流SSL图异常检测算法上进行大量实验,验证了原有选择方式的问题及所提策略的有效性。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) is an emerging paradigm that exploits supervisory signals generated from the data itself, and many recent studies have leveraged SSL to conduct graph anomaly detection. However, we empirically found that three important factors can substantially impact detection performance across datasets: 1) the specific SSL strategy employed; 2) the tuning of the strategy's hyperparameters; and 3) the allocation of combination weights when using multiple strategies. Most SSL-based graph anomaly detection methods circumvent these issues by arbitrarily or selectively (i.e., guided by label information) choosing SSL strategies, hyperparameter settings, and combination weights. While an arbitrary choice may lead to subpar performance, using label information in an unsupervised setting is label information leakage and leads to severe overestimation of a method's performance. Leakage has been criticized as "one of the top ten data mining mistakes", yet many recent studies on SSL-based graph anomaly detection have been using label information to select hyperparameters. To mitigate this issue, we propose to use an internal evaluation strategy (with theoretical analysis) to select hyperparameters in SSL for unsupervised anomaly detection. We perform extensive experiments using 10 recent SSL-based graph anomaly detection algorithms on various benchmark datasets, demonstrating both the prior issues with hyperparameter selection and the effectiveness of our proposed strategy.

自监督学习图异常检测超参数优化无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。