提出防泄漏评估方法,让轴承故障诊断模型更真实可信。
Towards a more realistic evaluation of machine learning models for bearing fault diagnosis
- 按物理轴承划分数据,杜绝训练测试重叠
- 多故障共存检测提升真实场景适应性
- 强调数据多样性对泛化能力的关键作用
可靠检测轴承故障对旋转机械的安全与效率至关重要。尽管机器学习(尤其是深度学习)在受控环境下表现优异,但许多研究因方法缺陷难以推广至实际应用,主要问题在于数据泄露。本文研究基于振动信号的轴承故障诊断中的数据泄露问题及其对模型评估的影响。我们证明,常见的分段或工况划分策略会引入虚假相关性,虚高性能指标。为此,提出以轴承为单位的数据划分方法,确保训练与测试使用不同物理部件,实现无泄露评估。同时将分类任务重构为多标签问题,支持共现故障检测,并采用基于ROC曲线的非依赖先验概率的评价指标。此外,研究发现唯一训练轴承数量是决定模型鲁棒性的关键因素。我们在四大常用数据集(CWRU、PU、UORED-VAFCLS、HUST bearing)上验证了该方法。本研究强调了泄漏感知评估协议的重要性,为数据划分、模型选择与验证提供了实用指南,推动工业故障诊断中更可信的机器学习系统发展。
原文摘要 · Abstract (English)
Reliable detection of bearing faults is essential for maintaining the safety and operational efficiency of rotating machinery. While recent advances in machine learning (ML), particularly deep learning, have shown strong performance in controlled settings, many studies fail to generalize to real-world applications due to methodological flaws, most notably data leakage. This paper investigates the issue of data leakage in vibration-based bearing fault diagnosis and its impact on model evaluation. We demonstrate that common dataset partitioning strategies, such as segment-wise and condition-wise splits, introduce spurious correlations that inflate performance metrics. To address this, we propose a rigorous, leakage-free evaluation methodology centered on bearing-wise data partitioning, ensuring no overlap between the physical components used for training and testing. Additionally, we reformulate the classification task as a multi-label problem, enabling the detection of co-occurring fault types and the use of prevalence-independent metrics based on the ROC curve. Beyond preventing leakage, we also examine the effect of dataset diversity on generalization, showing that the number of unique training bearings is a decisive factor for achieving robust performance. We evaluate our methodology on four widely adopted datasets: Case Western Reserve University (CWRU), Paderborn University (PU), University of Ottawa (UORED-VAFCLS) and Hanoi University of Science and Technology (HUST bearing). This study highlights the importance of leakage-aware evaluation protocols and provides practical guidelines for dataset partitioning, model selection, and validation, fostering the development of more trustworthy ML systems for industrial fault diagnosis applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。