提出防数据泄露的多任务故障诊断评估方法,解决模型性能虚高问题。
Leakage-Robust Evaluation and Data-Scale Sensitivity of Attention-Enhanced Multi-Task Learning for Joint Fault Diagnosis and Remaining Useful Life Estimation

- 设计分块泄漏审计分割协议,避免时间序列数据泄露
- 在C-MAPSS上分类准确率达84.12%,回归R2达0.86,显著优于基线
- 揭示标签来源影响任务稳定性,指导小样本下是否联合训练
多任务深度学习模型在联合故障分类与剩余寿命(RUL)回归中应用日益广泛,但性能常受滑动窗口划分方式影响。本文以注意力增强的多任务网络AMTLNet为基础,在NASA C-MAPSS、NASA IMS和UCI液压系统三个公开数据集上研究该问题。发现不当划分可使分类准确率从真实20-60%虚增至99.9%,或因类别退化降至0%。为此提出基于分块的泄漏审计分割方案,采用五次随机种子、单因素方差分析及Tukey HSD检验进行评估。在C-MAPSS数据集上,使用19,976个无泄漏训练窗口,AMTLNet分类准确率84.12±0.96%,与单任务CNN-LSTM相当(Tukey p=1.0),R2达0.86±0.01,显著优于朴素多任务基线。在轴承与液压系统小数据集上,多任务训练不稳定:轴承任务分类下降,液压任务回归下降。差异源于标签来源特性,据此提出适用性判断框架。消融实验表明,多头注意力分支是回归稳定主因,移除后R2由0.861降至0.766,分类方差翻倍;卷积分支虽占约1/3参数,对回归贡献甚微。
原文摘要 · Abstract (English)
Multi-task deep learning models that jointly perform fault classification and remaining useful life (RUL) regression are increasingly used in predictive maintenance, yet reported performance can be strongly affected by how sliding-window sequences are split into training and test sets. We investigate this issue using AMTLNet, an attention-enhanced multi-task architecture, on three public benchmarks: NASA C-MAPSS, NASA IMS, and the UCI Hydraulic System dataset. We show that naive splitting can inflate classification accuracy from a genuine 20-60 percent to 99.9 percent, or reduce it to 0 percent through degenerate class representation. To address this, we introduce a chunk-based, leakage-audited splitting protocol and evaluate all models using five seeds, one-way ANOVA, and Tukey HSD tests. On C-MAPSS, with 19,976 leakage-free training windows, AMTLNet matches a single-task CNN-LSTM baseline in classification, achieving 84.12 +/- 0.96 percent accuracy with Tukey p = 1.0, and reaches an R2 of 0.86 +/- 0.01 while significantly outperforming a naive multi-task baseline. On the smaller Bearing and Hydraulic datasets, multi-task training is unstable, but the failure mode differs: classification degrades for Bearing, whereas regression degrades for Hydraulic. We relate this asymmetry to label provenance and propose a practical framework for deciding when joint training is appropriate under data scarcity. Ablation results show that the multi-head attention branch is the main contributor to regression stability. Removing it reduces R2 from 0.861 to 0.766 and more than doubles classification variance, whereas the convolutional branch contributes little to regression despite using about one-third of the parameters. This study contributes a reusable leakage-audit protocol, seed-transparent evaluation, and evidence that task-specific stability depends more on label provenance than on task type.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。