arXiv:2606.10288cs.RO2026-06

用模型辅助强化学习,让机器人在稀疏地形上安全稳定地行走

MARCH: Model-Assisted Reinforcement Learning for the Perceptive Control of Humanoids over Sparse Footholds

论文配图:MARCH: Model-Assisted Reinforcement Learning for the Perceptive Control of Humanoids over Sparse Footholds
图 1 · 摘自论文原文
  • 先用简化模型生成安全路径,再训练带控制李雅普诺夫函数的教师策略
  • 通过知识蒸馏让视觉学生策略实现平滑步态,步石表现媲美纯强化学习
  • 适合需要高安全性的足式机器人控制,尤其在复杂稀疏地形中

在稀疏地形上实现感知型双足步行仍是难题:基于模型的方法精度高但对不确定性敏感,无模型方法鲁棒性强却难以发现安全所需的精确运动。本文提出一种模型辅助强化学习框架,分三步进行:(1) 利用简化模型生成安全参考轨迹;(2) 训练一个受控制李雅普诺夫函数(CLF)奖励驱动的特权教师策略,围绕该轨迹优化;(3) 将教师策略蒸馏为基于视觉的学生策略。实验表明,该方法生成物理合理的运动行为,提升样本效率,减少复杂学习课程依赖,实现更平滑的步态,且在步石任务上的表现与纯模型自由基线相当。我们在仿真中验证了该方法,并成功部署于Unitree G1人形机器人,在具有横向约束的稀疏脚点环境中完成导航。

原文摘要 · Abstract (English)

Perceptive bipedal locomotion over sparse terrain remains a difficult challenge: model-based methods are precise but brittle to uncertainty, while model-free methods are robust but struggle to discover the precise, constrained motions required for safety-critical locomotion where small errors can cause catastrophic failures. We propose a model-assisted reinforcement learning (RL) framework that combines both perspectives in three steps: (1) generate a safe reference trajectory using simplified models; (2) train a privileged teacher policy guided by a control Lyapunov function (CLF) reward built around the safe reference trajectory; and (3) distill the teacher into a vision-based student policy. We show that this model-assistance procedure produces physically grounded locomotion, improving sample efficiency, reducing the need for a complex learning curriculum, and achieving smoother locomotion behavior alongside stepping stone performance comparable to model-free baselines. We validate our approach in simulation and demonstrate successful deployment on a Unitree G1 humanoid robot navigating sparse footholds with lateral constraints.

人形机器人强化学习步态控制模型辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。