arXiv:2509.19573cs.RO2025-09被引 4

用控制李雅普诺夫函数提升人形机器人跑步的稳定性与鲁棒性

Chasing Stability: Humanoid Running via Control Lyapunov Function Guided Reinforcement Learning

  • 将控制李雅普诺夫函数融入强化学习,自动设计奖励信号
  • 实现包含腾空与单足支撑阶段的稳定跑步,支持室外与跑台运行
  • 仅靠机载传感器即可精准跟踪全局参考轨迹,适合全自主系统

在人形机器人上实现高度动态行为(如跑步)需要兼具鲁棒性与精度的控制器,传统方法难以应对非线性与混合动力学的实时控制挑战。近期强化学习因能处理复杂动力学而受到关注。本文提出CLF-RL方法,将控制李雅普诺夫函数(CLFs)与优化的动态参考轨迹嵌入强化学习训练过程以构建奖励函数。该方法无需手动设计和调参启发式奖励,同时促进可证明的稳定性,并提供有意义的中间奖励引导学习。通过基于动力学可行轨迹进行策略学习,显著拓展了机器人的动态能力,实现了包含飞行相与单支撑相的稳定跑步。所获策略在跑台上及户外环境中均表现可靠,对躯干与足部扰动具有鲁棒性;且仅使用机载传感器即可实现精确的全局参考轨迹跟踪,为实现完整自主系统迈出关键一步。

原文摘要 · Abstract (English)

Achieving highly dynamic behaviors on humanoid robots, such as running, requires controllers that are both robust and precise, and hence difficult to design. Classical control methods offer valuable insight into how such systems can stabilize themselves, but synthesizing real-time controllers for nonlinear and hybrid dynamics remains challenging. Recently, reinforcement learning (RL) has gained popularity for locomotion control due to its ability to handle these complex dynamics. In this work, we embed ideas from nonlinear control theory, specifically control Lyapunov functions (CLFs), along with optimized dynamic reference trajectories into the reinforcement learning training process to shape the reward. This approach, CLF-RL, eliminates the need to handcraft and tune heuristic reward terms, while simultaneously encouraging certifiable stability and providing meaningful intermediate rewards to guide learning. By grounding policy learning in dynamically feasible trajectories, we expand the robot's dynamic capabilities and enable running that includes both flight and single support phases. The resulting policy operates reliably on a treadmill and in outdoor environments, demonstrating robustness to disturbances applied to the torso and feet. Moreover, it achieves accurate global reference tracking utilizing only on-board sensors, making a critical step toward integrating these dynamic motions into a full autonomy stack.

人形机器人强化学习稳定性控制运动规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。