融合图像与时频特征,提升轴承剩余寿命预测精度与可解释性。
A Novel Multimodal RUL Framework for Remaining Useful Life Estimation with Layer-wise Explanations
- 用图像和时频图双模态表示振动信号,提取退化特征。
- 在两个数据集上用更少数据达到或超越顶尖模型性能。
- 引入分层可解释技术,可视化关键故障特征,适合工业部署。
机械系统剩余使用寿命(RUL)估计是故障预测与健康管理(PHM)的核心。滚动轴承是设备失效的主要原因,亟需稳健的RUL估计方法。现有方法普遍存在泛化能力差、鲁棒性弱、数据需求高、可解释性不足等问题。本文提出一种新型多模态RUL框架,联合利用多通道非平稳振动信号的图像表示(ImR)与时频表示(TFR)。该架构包含三个分支:(1)ImR分支与(2)TFR分支均采用带残差连接的多空洞卷积块提取空间退化特征;(3)融合分支将两者特征拼接后输入LSTM以建模时间退化模式。随后通过多头注意力机制突出关键特征,经线性层完成最终RUL回归。为实现有效多模态学习,振动信号通过Bresenham算法生成ImR,用连续小波变换生成TFR。本文还提出多模态分层相关性传播(multimodal-LRP),显著提升模型透明度。在XJTU-SY与PRONOSTIA基准数据集上验证表明,本方法在已知与未知工况下均表现优异,较先进方法在XJTU-SY上减少约28%训练数据,在PRONOSTIA上减少约48%。模型具备强抗噪能力,且multimodal-LRP可视化证实预测结果具可解释性与可信度,适用于实际工业场景。
原文摘要 · Abstract (English)
Estimating the Remaining Useful Life (RUL) of mechanical systems is pivotal in Prognostics and Health Management (PHM). Rolling-element bearings are among the most frequent causes of machinery failure, highlighting the need for robust RUL estimation methods. Existing approaches often suffer from poor generalization, lack of robustness, high data demands, and limited interpretability. This paper proposes a novel multimodal-RUL framework that jointly leverages image representations (ImR) and time-frequency representations (TFR) of multichannel, nonstationary vibration signals. The architecture comprises three branches: (1) an ImR branch and (2) a TFR branch, both employing multiple dilated convolutional blocks with residual connections to extract spatial degradation features; and (3) a fusion branch that concatenates these features and feeds them into an LSTM to model temporal degradation patterns. A multi-head attention mechanism subsequently emphasizes salient features, followed by linear layers for final RUL regression. To enable effective multimodal learning, vibration signals are converted into ImR via the Bresenham line algorithm and into TFR using Continuous Wavelet Transform. We also introduce multimodal Layer-wise Relevance Propagation (multimodal-LRP), a tailored explainability technique that significantly enhances model transparency. The approach is validated on the XJTU-SY and PRONOSTIA benchmark datasets. Results show that our method matches or surpasses state-of-the-art baselines under both seen and unseen operating conditions, while requiring ~28 % less training data on XJTU-SY and ~48 % less on PRONOSTIA. The model exhibits strong noise resilience, and multimodal-LRP visualizations confirm the interpretability and trustworthiness of predictions, making the framework highly suitable for real-world industrial deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。