分析异常检测中不平衡数据下常用评估指标的表现差异。
An Insight on Evaluation Metrics Under the Imbalanced Case of Anomaly Detection

- 通过可视化指标景观,揭示不同指标在各类别不平衡下的偏好与稳定性。
- 发现AUROC、AUPR等指标值随异常比例变化而显著波动,解释需谨慎。
- 为跨数据集比较结果提供实用指导,适合关注评估可靠性的研究者。
异常检测本质上存在严重的类别不平衡,导致评估指标的解读变得困难。尽管AUROC、AUPR、F1-score和MCC等指标被广泛使用,但其数值含义会随异常比例变化而改变。本文分析了这四种常见指标在不同不平衡程度下的行为表现。重点研究了指标景观(metric landscapes),即指标值与真正例率、真负例率之间的关系图,提供对指标偏好和稳定性的直观理解。分析结果为在不同不平衡比率的数据集间解释和比较异常检测结果提供了实用指导。
原文摘要 · Abstract (English)
Anomaly detection is inherently characterised by severe class imbalance, making the interpretation of evaluation metrics challenging. Although metrics such as AUROC, AUPR, F1-score, and MCC are widely used, their values convey different meanings depending on the anomaly ratio. In this work, we analyse the behaviour of those four common anomaly detection metrics under varying levels of imbalance. We focus on the study of metric landscapes, visualisations that relate metric values to true positive and true negative rates, providing an intuitive view of metric preferences and stability. Our analysis offers practical guidance for interpreting and comparing anomaly detection results across datasets with different imbalance ratios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。