arXiv:2601.00005cs.LGcs.AI2026-01被引 2

在极端不平衡工业数据上测试14种异常检测方法,发现故障样本数量决定最佳算法选择。

Evaluating Anomaly Detectors for Simulated Highly Imbalanced Industrial Classification Problems

  • 用2维和10维超球面模拟异常分布,测试14种检测器在0.05%~20%异常率下的表现。
  • 少于20个故障样本时,无监督方法(kNN/LOF)最优;30-50个时,半监督/有监督方法显著提升。
  • 特征维度从2升到10时,半监督方法优势明显,小样本下泛化能力下降严重。

机器学习在工业质量控制与预测性维护中具有潜力,但面临极端类别不平衡的挑战,主要源于训练时故障数据稀缺。本文通过一个与问题无关的模拟数据集,评估了异常检测算法的表现,该数据集基于2维和10维超球面构造异常分布。在异常率0.05%至20%、训练样本量1,000至10,000(测试集40,000)的条件下,对14种检测器进行基准测试,以评估其性能与泛化误差。结果表明,最优检测器高度依赖训练集中故障样本总数:当故障样本少于20个时,无监督方法(kNN/LOF)表现最佳;当故障样本达30-50个时,半监督(XGBOD)和有监督(SVM/CatBoost)方法性能大幅提升。尽管在仅2个特征时半监督方法无显著优势,但在10个特征下改进明显。研究揭示了小样本下异常检测方法泛化能力下降的问题,为工业场景部署提供了实用指导。

原文摘要 · Abstract (English)

Machine learning offers potential solutions to current issues in industrial systems in areas such as quality control and predictive maintenance, but also faces unique barriers in industrial applications. An ongoing challenge is extreme class imbalance, primarily due to the limited availability of faulty data during training. This paper presents a comprehensive evaluation of anomaly detection algorithms using a problem-agnostic simulated dataset that reflects real-world engineering constraints. Using a synthetic dataset with a hyper-spherical based anomaly distribution in 2D and 10D, we benchmark 14 detectors across training datasets with anomaly rates between 0.05% and 20% and training sizes between 1 000 and 10 000 (with a testing dataset size of 40 000) to assess performance and generalization error. Our findings reveal that the best detector is highly dependant on the total number of faulty examples in the training dataset, with additional healthy examples offering insignificant benefits in most cases. With less than 20 faulty examples, unsupervised methods (kNN/LOF) dominate; but around 30-50 faulty examples, semi-supervised (XGBOD) and supervised (SVM/CatBoost) detectors, we see large performance increases. While semi-supervised methods do not show significant benefits with only two features, the improvements are evident at ten features. The study highlights the performance drop on generalization of anomaly detection methods on smaller datasets, and provides practical insights for deploying anomaly detection in industrial environments.

异常检测工业应用不平衡数据小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。