用控制理论证明强化学习机器人走路稳定,理论验证可行。
Stability of Control Lyapunov Function Guided Reinforcement Learning

- 用控制李雅普诺夫函数设计奖励,引导强化学习
- 证明连续与离散时间下均具指数稳定性
- 适用于需高可靠性的机器人控制场景
强化学习已成为人形机器人实现行走的主流方法,但其控制策略的稳定性分析仍不足。近期工作尝试将控制理论与强化学习结合,其中一种重要方法是使用控制李雅普诺夫函数(CLF)构造强化学习奖励,即CLF-RL,已在实践中取得成功。本文研究了基于CLF-RL的最优控制器的稳定性特性,旨在连接实验观察到的稳定性与理论保证。将RL问题视为最优控制问题,分别在连续与离散时间下证明了指数稳定性,涵盖核心CLF奖励项及实际中使用的附加项。理论边界通过双积分器与倒立摆系统进行数值验证。最终,将CLF引导奖励应用于行走人形机器人,成功生成稳定的周期性运动轨迹。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has become the de facto method for achieving locomotion on humanoid robots in practice, yet stability analysis of the corresponding control policies is lacking. Recent work has attempted to merge control theoretic ideas with reinforcement learning through control guided learning. A notable example of this is the use of a control Lyapunov function (CLF) to synthesize the reinforcement learning rewards, a technique known as CLF-RL, which has shown practical success. This paper investigates the stability properties of optimal controllers using CLF-RL with the goal of bridging experimentally observed stability with theoretical guarantees. The RL problem is viewed as an optimal control problem and exponential stability is proven in both continuous and discrete time using both core CLF reward terms and the additional terms used in practice. The theoretical bounds are numerically verified on systems such as the double integrator and cart-pole. Finally, the CLF guided rewards are implemented for a walking humanoid robot to generate stable periodic orbits.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。