针对工业故障检测中罕见故障类型多样的问题,提出多生成器对抗学习方法提升识别效果。
Breaking the Homogeneity Assumption: Specialized Multi-Generator Adversarial Learning for Rare Failure Detection in Predictive Maintenance
- 为不同故障类型设计独立生成器,突破传统生成方法的同质化假设。
- 在AI4I 2020数据集上实现更高PR-AUC与召回率,优于单一生成器和传统重采样方法。
- 适合需要精准识别多种罕见故障的工业预测性维护场景。
预测性维护中的监督学习模型通常在高度不平衡的工业数据上训练:设备故障发生频率低,但对运营影响巨大。此外,故障数据通常非同质,不同故障模式源于不同的物理过程,在少数类中呈现多模态分布。传统不平衡处理方法(如欠采样、SMOTE插值、代价敏感学习)通常假设少数类是同质的,因此在工业实践中复杂多样的情形下效果受限。本文探索了基于故障类型感知的生成增强方案,以提升预测性维护系统对罕见故障的识别能力。采用防泄露实验设计,对比五种不平衡处理方法:代价敏感学习、随机欠采样、SMOTE过采样、单生成器GAN增强,以及一个具有独立生成器、分别学习各故障子类型的专用多生成器GAN架构。使用精确率/召回率导向指标评估性能,主要衡量标准为PR-AUC。在AI4I 2020预测性维护数据集上的实验表明,所提出的多生成器GAN框架能生成更真实的少数类样本,相比传统重采样方法和单生成器GAN,显著提升PR-AUC与召回率。
原文摘要 · Abstract (English)
Supervised learning models in the predictive maintenance field are regularly trained on highly imbalanced industrial datasets: machine failures occur rarely but have a disproportionate effect on operations. In addition to the clear class disparity, failure data are typically non-homogeneous, with different failure modes arising from distinct physical processes and exhibiting a multimodal distribution across minorities and classes. Traditional imbalance-management methods, e.g., undersampling, SMOTE-based interpolation, or cost-sensitive learning, typically assume that the minority population is homogeneous. This means their effectiveness is severely limited in the multifaceted conditions encountered in industrial practice. This paper determines the possibility of a failure-type-conscious generative augmentation program to improve the identification of infrequent failures in predictive maintenance systems. An experimental design that is leakage-safe is used to compare five imbalance-handling methods: cost-sensitive learning, random undersampling, SMOTE oversampling, single-generator GAN augmentation, and a specialized multi-generator GAN architecture that has independent generators that are asked to learn individual failure subtypes. Precision/Recall-oriented measures are used to quantify model performance; the main evaluation measure is the PR-AUC. Experiments conducted on the AI4I 2020 predictive maintenance dataset indicate that the proposed multi-generator GAN framework produces more realistic minority samples, yielding higher PR-AUC and recall scores compared to traditional resampling methods and individual-generator GAN augmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。