arXiv:2508.09354cs.RO2025-08被引 7

用控制李雅普诺夫函数引导强化学习,提升双足机器人运动稳定性。

CLF-RL: Control Lyapunov Function Guided Reinforcement Learning

  • 结合模型预测与李雅普诺夫函数设计中间奖励,避免繁琐调参。
  • 在仿真和真实G1机器人上均显著提升鲁棒性与运动性能。
  • 训练时用参考轨迹,部署时无额外开销,适合实际机器人应用。

强化学习在生成双足机器人稳健运动策略方面展现出潜力,但常面临奖励设计繁琐及目标函数不佳时敏感的问题。本文提出一种结构化的奖励塑造框架,利用基于模型的轨迹生成与控制李雅普诺夫函数(CLFs)指导策略学习。采用两种基于模型的规划器生成参考轨迹:针对速度条件运动规划的简化线性倒立摆(LIP)模型,以及基于全阶动力学的混合零动态(HZD)预计算步态库。这些规划器定义了期望的末端执行器与关节轨迹,用于构建基于CLF的奖励,惩罚跟踪误差并促进快速收敛。该方法提供有意义的中间奖励,一旦获得参考轨迹即可直接实现。参考轨迹与CLF塑造仅用于训练阶段,部署时策略轻量高效。我们在仿真和真实世界中对Unitree G1机器人进行了广泛实验,结果表明,相比基线强化学习策略,CLF-RL显著提升了鲁棒性,且优于经典跟踪奖励强化学习方法。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has shown promise in generating robust locomotion policies for bipedal robots, but often suffers from tedious reward design and sensitivity to poorly shaped objectives. In this work, we propose a structured reward shaping framework that leverages model-based trajectory generation and control Lyapunov functions (CLFs) to guide policy learning. We explore two model-based planners for generating reference trajectories: a reduced-order linear inverted pendulum (LIP) model for velocity-conditioned motion planning, and a precomputed gait library based on hybrid zero dynamics (HZD) using full-order dynamics. These planners define desired end-effector and joint trajectories, which are used to construct CLF-based rewards that penalize tracking error and encourage rapid convergence. This formulation provides meaningful intermediate rewards, and is straightforward to implement once a reference is available. Both the reference trajectories and CLF shaping are used only during training, resulting in a lightweight policy at deployment. We validate our method both in simulation and through extensive real-world experiments on a Unitree G1 robot. CLF-RL demonstrates significantly improved robustness relative to the baseline RL policy and better performance than a classic tracking reward RL formulation.

强化学习机器人控制李雅普诺夫函数双足机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。