数据量太小,再罕见的异常也证明不了存在。
Rare anomalies require large datasets: About proving the existence of anomalies
- 提出判断异常是否存在需满足的样本量阈值公式
- 发现当样本数不足时,无法确认异常存在
- 适合做异常检测可靠性评估的研究者参考
检测数据集中是否存在异常是有效异常检测的关键,但该问题在异常检测文献中仍被严重忽视。本文通过超过三百万次统计测试,在多种异常检测任务与算法上展开全面研究,揭示了数据集规模、异常比例与算法相关常数 α_algo 之间的关系。结果表明,对于大小为 N、异常比例为 ν 的无标签数据集,只有当满足条件 $ N \ge \frac{α_{\text{algo}}}{ν^2} $ 时,才能确证异常存在。该阈值意味着:当异常过于稀少时,证明其存在将变得不可行,存在理论下限。
原文摘要 · Abstract (English)
Detecting whether any anomalies exist within a dataset is crucial for effective anomaly detection, yet it remains surprisingly underexplored in anomaly detection literature. This paper presents a comprehensive study that addresses the fundamental question: When can we conclusively determine that anomalies are present? Through extensive experimentation involving over three million statistical tests across various anomaly detection tasks and algorithms, we identify a relationship between the dataset size, contamination rate, and an algorithm-dependent constant $ α_{\text{algo}} $. Our results demonstrate that, for an unlabeled dataset of size $ N $ and contamination rate $ ν$, the condition $ N \ge \frac{α_{\text{algo}}}{ν^2} $ represents a lower bound on the number of samples required to confirm anomaly existence. This threshold implies a limit to how rare anomalies can be before proving their existence becomes infeasible.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。