用批评者当李雅普诺夫函数,让强化学习自动稳定动态系统。
Critic as Lyapunov function (CALF): a model-free, stability-ensuring agent
- 用批评者函数充当李雅普诺夫函数,实现无需模型的在线系统稳定
- 在移动机器人仿真中显著提升学习性能,优于SARSA-m等方法
- 适合需要实时稳定性的强化学习应用,如机器人控制
本文提出一种新型无模型强化学习智能体CALF(Critic As Lyapunov Function),可确保在线环境稳定,即动态系统在每次学习过程中均被稳定。在移动机器人模拟器的案例研究中,该方法显著提升了整体学习性能。与经典SARSA相比,其改进版本SARSA-m虽在部分场景成功,但仍不及CALF表现。此外,CALF还能增强预先提供的基本稳定控制器。总体而言,该方法为经典控制与强化学习融合提供了一种可行路径。现有类似方法多为离线或基于模型,例如将模型预测控制融入智能体的设计。
原文摘要 · Abstract (English)
This work presents and showcases a novel reinforcement learning agent called Critic As Lyapunov Function (CALF) which is model-free and ensures online environment, in other words, dynamical system stabilization. Online means that in each learning episode, the said environment is stabilized. This, as demonstrated in a case study with a mobile robot simulator, greatly improves the overall learning performance. The base actor-critic scheme of CALF is analogous to SARSA. The latter did not show any success in reaching the target in our studies. However, a modified version thereof, called SARSA-m here, did succeed in some learning scenarios. Still, CALF greatly outperformed the said approach. CALF was also demonstrated to improve a nominal stabilizer provided to it. In summary, the presented agent may be considered a viable approach to fusing classical control with reinforcement learning. Its concurrent approaches are mostly either offline or model-based, like, for instance, those that fuse model-predictive control into the agent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。