用量子退火增强强化学习,提升设备剩余寿命预测精度。
Quantum Annealing Enhanced Reinforcement Learning for Accurate Remaining Useful Lifetime Prediction

- 将Q值更新转为量子退火可解的二元优化问题
- 在多个指标上优于经典与量子基线模型
- 适合需要高精度预测的工业维护场景
剩余使用寿命(RUL)估计是预测性维护的核心,意外故障的代价远高于资产本身。传统统计退化模型难以捕捉真实系统的强非线性,而数据驱动模型在高维非凸空间中常收敛于次优解。本文提出量子退火增强的Q学习(QAQL)框架,将量子退火的采样行为与Q学习的序列决策结合:每次Q值更新被编码为小型无约束二次二值优化(QUBO)问题,其基态对应贪婪动作;退火器不作为确定性优化器,而是通过多次读取返回近优动作分布,该随机动作选择有效抑制了非线性退化轨迹上的过早收敛。QUBO在D-Wave Advantage系统上通过小嵌入求解,退火器嵌入强化学习循环而非训练后附加。在NASA C-MAPSS涡轮风扇数据集和一个设备集群预测维护数据集上验证,跨六种误差指标、多次独立运行平均,QAQL显著优于所考虑的经典与量子基线。结果表明,量子退火可在工业预测性维护应用中作为实际可用的优化器,而非仅理论概念。
原文摘要 · Abstract (English)
Remaining useful life (RUL) estimation is central to predictive maintenance, where an unplanned failure can cost far more than the asset itself. Statistical degradation models miss the strong nonlinearity of real systems, and data-driven models often converge to suboptimal solutions in high-dimensional, non-convex search spaces. We propose a Quantum Annealing enhanced Q-Learning (QAQL) framework that couples the sampling behaviour of quantum annealing with the sequential decision making of Q-learning. Each Q-value update is encoded as a small quadratic unconstrained binary optimization (QUBO) whose ground state is the greedy action; rather than acting as a deterministic optimizer, the annealer returns a distribution over near-optimal actions across many reads, and this stochastic action selection supplies the exploration that curbs premature convergence on nonlinear degradation trajectories. The QUBO is solved on the D-Wave Advantage system using minor embedding, with the annealer woven into the reinforcement-learning loop rather than bolted on after training. We validate QAQL on two public benchmarks: the NASA C-MAPSS turbofan engine datasets and a device-fleet predictive maintenance dataset. Averaged over many independent runs and across six error metrics, QAQL outperforms the classical and quantum baselines considered in this study, with statistically significant improvements. The results indicate that quantum annealing is a usable, not merely theoretical, optimizer inside a reinforcement-learning loop for industrial predictive-maintenance applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。