用强化学习安全控制内燃机,实现高效低排放运行。
Safe Reinforcement Learning for Real-World Engine Control
- 基于DDPG算法与k近邻实时监控,保障发动机安全
- 控制误差仅0.1374巴,媲美现有神经网络控制器
- 适合想在真实设备上安全应用强化学习的研究者
本文提出一套用于安全关键型现实环境的强化学习(RL)工具链,以单缸内燃机在均质压燃(HCCI)模式下的瞬态负载控制为例。HCCI具有非线性、自回归和随机性特征,传统控制方法难以应对。尽管强化学习提供了可行方案,但需解决如压力上升速率过高等安全风险问题。单次不当控制可能损坏发动机或导致失火停机。此外,工作极限未知,须实验确定。为此,采用基于k近邻算法的实时安全监控机制,确保与试验台的安全交互。实验表明,该方法使指示平均有效压力的均方根误差降至0.1374 bar,性能与文献中神经网络控制器相当。工具链还成功适配策略以提升乙醇能量占比,推动可再生能源使用同时保持安全。该方法解决了强化学习在真实安全关键场景应用的长期难题,其灵活性和安全性机制为未来在发动机试验台及其他类似场景的应用铺平道路。
原文摘要 · Abstract (English)
This work introduces a toolchain for applying Reinforcement Learning (RL), specifically the Deep Deterministic Policy Gradient (DDPG) algorithm, in safety-critical real-world environments. As an exemplary application, transient load control is demonstrated on a single-cylinder internal combustion engine testbench in Homogeneous Charge Compression Ignition (HCCI) mode, that offers high thermal efficiency and low emissions. However, HCCI poses challenges for traditional control methods due to its nonlinear, autoregressive, and stochastic nature. RL provides a viable solution, however, safety concerns, such as excessive pressure rise rates, must be addressed when applying to HCCI. A single unsuitable control input can severely damage the engine or cause misfiring and shut down. Additionally, operating limits are not known a priori and must be determined experimentally. To mitigate these risks, real-time safety monitoring based on the k-nearest neighbor algorithm is implemented, enabling safe interaction with the testbench. The feasibility of this approach is demonstrated as the RL agent learns a control policy through interaction with the testbench. A root mean square error of 0.1374 bar is achieved for the indicated mean effective pressure, comparable to neural network-based controllers from the literature. The toolchain's flexibility is further demonstrated by adapting the agent's policy to increase ethanol energy shares, promoting renewable fuel use while maintaining safety. This RL approach addresses the longstanding challenge of applying RL to safety-critical real-world environments. The developed toolchain, with its adaptability and safety mechanisms, paves the way for future applicability of RL in engine testbenches and other safety-critical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。