arXiv:2511.21378cs.LGcs.AI2025-11

新方法动态剔除异常数据,提升噪声环境下异常检测精度。

Anomaly Detection with Adaptive and Aggressive Rejection for Contaminated Training Data

  • 用改进z-score与高斯混合模型动态设定剔除阈值
  • 在30个表格数据集上比当前最佳方法高0.041 AUROC
  • 适合噪声大、正常异常重叠的医疗安防场景

异常检测面临训练数据污染的核心挑战,传统模型假设训练数据全为正常样本。现有方法依赖固定污染比例,但假设与实际不符时性能严重下降,尤其在正常与异常分布重叠的噪声环境中。为此,本文提出自适应激进剔除(AAR)方法,通过改进z-score与基于高斯混合模型的阈值动态排除异常。AAR融合硬剔除与软剔除策略,在保留正常数据与剔除异常之间取得良好平衡。在两个图像数据集和三十个表格数据集上的大量实验表明,AAR相比当前最优方法提升0.041 AUROC。该方法具有可扩展性与可靠性,显著增强对污染数据集的鲁棒性,为安全、医疗等领域的实际应用提供支持。

原文摘要 · Abstract (English)

Handling contaminated data poses a critical challenge in anomaly detection, as traditional models assume training on purely normal data. Conventional methods mitigate contamination by relying on fixed contamination ratios, but discrepancies between assumed and actual ratios can severely degrade performance, especially in noisy environments where normal and abnormal data distributions overlap. To address these limitations, we propose Adaptive and Aggressive Rejection (AAR), a novel method that dynamically excludes anomalies using a modified z-score and Gaussian mixture model-based thresholds. AAR effectively balances the trade-off between preserving normal data and excluding anomalies by integrating hard and soft rejection strategies. Extensive experiments on two image datasets and thirty tabular datasets demonstrate that AAR outperforms the state-of-the-art method by 0.041 AUROC. By providing a scalable and reliable solution, AAR enhances robustness against contaminated datasets, paving the way for broader real-world applications in domains such as security and healthcare.

异常检测数据污染鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。