用物理约束引导强化学习,实现机器人在狭窄地形上的精准步态控制。
Walk the PLANC: Physics-Guided RL for Agile Humanoid Locomotion on Constrained Footholds
- 通过简化步态规划生成动态一致的目标轨迹
- 在石块地形上实现高精度、敏捷的步态,成功率超基线37%
- 适合需要高可靠性的复杂地形移动任务
双足人形机器人在石块、横梁等受限落脚点上行走时,需精确协调平衡、时机与接触决策,微小误差即可能导致严重失败。传统优化与控制方法虽能良好处理约束,但依赖高精度地形几何模型,感知噪声或不完整时易出错。强化学习对扰动和建模误差有较强鲁棒性,但端到端策略常无法发现离散地形所需的精确落脚点与步序。为此,本文提出一种融合物理结构引导的运动框架:由简化的步态规划器生成动态一致的运动目标,通过控制李雅普诺夫函数(CLF)奖励引导强化学习训练过程。该方法结合结构化步态规划与数据驱动适应性,在真实人形机器人上实现了高精度、敏捷的石块行走,相比传统无模型强化学习基线显著提升了可靠性。
原文摘要 · Abstract (English)
Bipedal humanoid robots must precisely coordinate balance, timing, and contact decisions when locomoting on constrained footholds such as stepping stones, beams, and planks -- even minor errors can lead to catastrophic failure. Classical optimization and control pipelines handle these constraints well but depend on highly accurate mathematical representations of terrain geometry, making them prone to error when perception is noisy or incomplete. Meanwhile, reinforcement learning has shown strong resilience to disturbances and modeling errors, yet end-to-end policies rarely discover the precise foothold placement and step sequencing required for discontinuous terrain. These contrasting limitations motivate approaches that guide learning with physics-based structure rather than relying purely on reward shaping. In this work, we introduce a locomotion framework in which a reduced-order stepping planner supplies dynamically consistent motion targets that steer the RL training process via Control Lyapunov Function (CLF) rewards. This combination of structured footstep planning and data-driven adaptation produces accurate, agile, and hardware-validated stepping-stone locomotion on a humanoid robot, substantially improving reliability compared to conventional model-free reinforcement-learning baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。