用强化学习+梯度搜索,让控制器在系统变化时仍稳定高效。
Improved Robustness of Deep Reinforcement Learning for Control of Time-Varying Systems by Bounded Extremum Seeking
- DRL结合有界极值搜索,实现无模型自适应控制。
- 在粒子加速器实验中,控制精度提升且对时变系统鲁棒性增强。
- 适合需要快速响应、参数多变的工业控制场景。
本文研究了将鲁棒、无模型的有界极值搜索(ES)反馈控制应用于一类非线性时变系统,以提升深度强化学习(DRL)控制器的鲁棒性。尽管DRL能从大量数据中快速学习并控制多参数系统,但当系统模型随时间快速变化时,其性能会急剧下降。有界ES可处理未知控制方向的时变系统,但随着调参数量增加,收敛速度变慢,且易陷入局部最优。我们证明,两者结合形成的混合控制器性能超越各自之和:DRL利用历史数据快速学习并逼近目标设定点,而有界ES确保对时变扰动的鲁棒性。通过一个通用时变系统的数值实验,以及对洛斯阿拉莫斯中子科学中心线性粒子加速器低能束传输段的自动调参,验证了该方法的有效性。
原文摘要 · Abstract (English)
In this paper, we study the use of robust model independent bounded extremum seeking (ES) feedback control to improve the robustness of deep reinforcement learning (DRL) controllers for a class of nonlinear time-varying systems. DRL has the potential to learn from large datasets to quickly control or optimize the outputs of many-parameter systems, but its performance degrades catastrophically when the system model changes rapidly over time. Bounded ES can handle time-varying systems with unknown control directions, but its convergence speed slows down as the number of tuned parameters increases and, like all local adaptive methods, it can get stuck in local minima. We demonstrate that together, DRL and bounded ES result in a hybrid controller whose performance exceeds the sum of its parts with DRL taking advantage of historical data to learn how to quickly control a many-parameter system to a desired setpoint while bounded ES ensures its robustness to time variations. We present a numerical study of a general time-varying system and a combined ES-DRL controller for automatic tuning of the Low Energy Beam Transport section at the Los Alamos Neutron Science Center linear particle accelerator.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。