用强化学习让数字孪生数据更贴近真实设备,解决故障诊断数据少的难题。
Digital Twin-Driven Adaptive Sim-to-Real Alignment via Reinforcement Learning for Vibration-Based Bearing Health Monitoring Under Data Scarcity
- 设计动态策略,按故障类型逐次调整仿真数据特征空间。
- 在多个数据集上实现92.8%的跨设备诊断准确率,无需重新训练编码器。
- 针对正常状态保留真实数据,故障类用对齐后的仿真数据增强。
基于振动的旋转机械健康监测在运行数据稀缺下仍面临挑战,主要源于故障事件结构稀疏及数字孪生生成信号中的仿真到现实差距。不同故障类型产生的冲击具有不同的周期性、幅值调制和频谱特性,导致特征空间差异在各类故障间本质异质。现有领域自适应方法采用无类别全局变换,无法在不破坏类别可分性的情况下弥合所有故障特异性差距,而统一源-目标混合则向数据丰富的正常类引入分布噪声。这些局限源于将序列化的、状态依赖的对齐问题误当作一次性优化处理。每次修正变换同时重塑所有类别分布,产生静态梯度下降无法解决的状态依赖性。本文将特征对齐建模为连续动作马尔可夫决策过程,通过近端策略优化求解,所学策略根据当前特征空间配置发出故障类型专属的仿射校正,奖励函数双目标平衡间隙最小化与可分性保持。提出不对称策略:为正常类保留真实数据,用策略对齐的仿真样本增强故障类。在XJTU-SY、CWRU及自建回转轴承测试平台验证表明,强化学习驱动对齐带来显著提升,跨设备线性探查达92.8%准确率,无需编码器重训练,验证了可迁移监测能力。
原文摘要 · Abstract (English)
Vibration-based health monitoring of rotating machinery requires reliable fault diagnosis under operational data constraints, yet condition assessment remains challenged by structural scarcity of fault events and heterogeneous sim-to-real gaps in digital twin-generated signals. Each fault type generates impulses with distinct periodicity, amplitude modulation, and spectral character, making feature-space discrepancies fundamentally heterogeneous across fault classes. Existing domain adaptation methods apply a class-agnostic global transformation that cannot close all fault-specific gaps without distorting inter-class separability, while uniform source-target mixing introduces distributional noise into the data-abundant Normal class. These limitations stem from treating a sequential, state-dependent alignment problem as a one-shot optimization. Each corrective transformation simultaneously reshapes all class distributions, creating state dependencies that static gradient descent cannot resolve. We formulate feature alignment as a continuous-action Markov decision process solved via Proximal Policy Optimization, where the learned policy issues fault-type-specific affine corrections responsive to the current feature-space configuration, with a dual-objective reward balancing gap minimization against separability preservation. An asymmetry-aware strategy reserves real data for the Normal class while augmenting fault classes with policy-aligned simulated samples. Validation across XJTU-SY, CWRU, and a self-built slewing bearing testbed confirms the dominant gain from reinforcement learning-driven alignment, and cross-equipment linear probing achieves 92.8% without encoder retraining, demonstrating transferable monitoring capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。