用强化学习优化退化系统维护,修复效果随次数递减。
A reinforcement learning agent for maintenance of deteriorating systems with increasingly imperfect repairs
- 基于双深度Q网络构建维护代理,无需预设阈值。
- 修复效果随次数增加而减弱,更贴近真实系统退化。
- 在连续退化空间中灵活决策,长期成本显著降低。
高效的维护对工程系统的成功应用至关重要。然而,工业4.0的实施带来了新挑战,亟需新的维护优化范式。机器学习技术在工程与维护领域日益普及,其中强化学习最具前景。本文提出一种伽马退化过程,并引入一种新型维护模型,其中修复效果随维修次数增加而逐渐减弱,更真实地反映现实系统退化行为。为生成维护策略,我们采用双深度Q网络(Double Deep Q-Network)架构开发强化学习代理。该代理具有两大优势:无需预设预防性阈值,且可在连续退化状态空间中运行。代理能适应不同场景,表现出高度灵活性。此外,我们分析了环境主要参数变化对维护策略的影响。实验表明,所提方法在长期成本方面显著优于其他常见维护策略。
原文摘要 · Abstract (English)
Efficient maintenance has always been essential for the successful application of engineering systems. However, the challenges to be overcome in the implementation of Industry 4.0 necessitate new paradigms of maintenance optimization. Machine learning techniques are becoming increasingly used in engineering and maintenance, with reinforcement learning being one of the most promising. In this paper, we propose a gamma degradation process together with a novel maintenance model in which repairs are increasingly imperfect, i.e., the beneficial effect of system repairs decreases as more repairs are performed, reflecting the degradational behavior of real-world systems. To generate maintenance policies for this system, we developed a reinforcement-learning-based agent using a Double Deep Q-Network architecture. This agent presents two important advantages: it works without a predefined preventive threshold, and it can operate in a continuous degradation state space. Our agent learns to behave in different scenarios, showing great flexibility. In addition, we performed an analysis of how changes in the main parameters of the environment affect the maintenance policy proposed by the agent. The proposed approach is demonstrated to be appropriate and to significatively improve long-run cost as compared with other common maintenance strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。