不同评测设置下异常检测算法排名差异大,结果不可靠。
Why Ranking Anomaly Detection Algorithms Isn't as Reliable as You May Think
- 用7种算法、690个数据集测试排名稳定性
- 多数情况下无固定最优算法,排名常变
- 数据集选择和超参影响最大,适合严谨评估者
异常检测是安全关键的机器学习任务,应用涵盖欺诈检测、网络入侵防御和工业监控。尽管已有大量算法提出,许多新方法声称达到顶尖性能,但这些结论常基于互不兼容的评测设置,导致可复现性和可靠性存疑。本文研究常见评测选择对算法排名稳定性的影响。基于OddBench基准套件中的7种代表性算法和690个数据集,分析在不同数据集选取、评估指标、超参数配置及随机种子下的排名变化。引入排名不稳定性度量以量化这种波动。结果表明,异常检测算法排名高度不稳定:在多数情况下,几乎所有竞争性算法都可能在某种配置下表现最佳。其中,数据集选择和超参数选择是排名不确定性的主要来源,而随机种子和评估指标影响相对较小。我们还发现,可靠评测需远超以往研究使用的更大更多样化的数据集集合。
原文摘要 · Abstract (English)
Anomaly detection is a safety-critical machine learning problem with applications ranging from fraud detection to network intrusion prevention and industrial monitoring. Despite the large number of proposed anomaly detection algorithms, many novel methods claim state-of-the-art performance. However, many authors do so under benchmark settings that are not aligned with one another. This lack of comparability raises concerns regarding the reproducibility and reliability of anomaly detection benchmarks. In this work, we study the impact of common benchmarking choices on the stability of algorithm rankings. Using seven representative anomaly detection algorithms and 690 datasets from the OddBench benchmark suite, we analyze how rankings change under varying dataset selections, evaluation metrics, hyperparameter configurations, and random seeds. To quantify this effect, we introduce a rank instability metric measuring the variability of algorithm rankings across benchmark settings. Our results show that algorithm rankings in anomaly detection are highly unstable. In many cases, almost every competitive algorithm can appear as the best-performing method under some benchmark configuration. Among the studied factors, dataset selection and hyperparameter choice contribute most strongly to ranking uncertainty, while random seeds and evaluation metrics have a comparatively limited impact. We also observe that reliable benchmarking requires substantially larger and more diverse dataset collections than the ones commonly used in prior work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。