arXiv:2604.17335cs.RO2026-04被引 12

用生成模型+追踪器实现机器人在复杂地形上的全身自适应行走。

Learning Whole-Body Humanoid Locomotion via Motion Generation and Motion Tracking

论文配图:Learning Whole-Body Humanoid Locomotion via Motion Generation and Motion Tracking
图 1 · 摘自论文原文
  • 先用扩散模型生成适应地形的参考动作,再用强化学习训练全身追踪器。
  • 在真实机器人上完成越障、爬梯及混合地形行走,成功率超90%。
  • 适合研究具身智能、机器人运动控制或想部署真实机器人的团队。

全身类人机器人行走面临高维控制、形态不稳定性以及需基于机载感知实时适应不同地形的挑战。直接使用奖励塑形的强化学习常导致下肢主导行为,而基于模仿的强化学习虽能学习更协调的全身动作,却通常仅能回放参考动作,缺乏在线感知适应能力。为此,本文提出一种结合参考动作学习与地形感知适应的全身行走框架:首先在重定向的人体动作数据上训练扩散模型,实现实时预测地形感知的参考动作;同时,利用该动作数据训练全身参考追踪器。为提升在生成不完美参考动作时的鲁棒性,进一步在闭环设置下冻结动作生成器对追踪器进行微调。系统支持方向性目标到达控制,并可在具备机载感知与计算的Unitree G1机器人上部署。硬件实验表明,其成功穿越箱子、障碍物、台阶及混合地形组合。定量结果进一步验证了在线动作生成与追踪器微调对泛化性和鲁棒性的提升作用。

原文摘要 · Abstract (English)

Whole-body humanoid locomotion is challenging due to high-dimensional control, morphological instability, and the need for real-time adaptation to various terrains using onboard perception. Directly applying reinforcement learning (RL) with reward shaping to humanoid locomotion often leads to lower-body-dominated behaviors, whereas imitation-based RL can learn more coordinated whole-body skills but is typically limited to replaying reference motions without a mechanism to adapt them online from perception for terrain-aware locomotion. To address this gap, we propose a whole-body humanoid locomotion framework that combines skills learned from reference motions with terrain-aware adaptation. We first train a diffusion model on retargeted human motions for real-time prediction of terrain-aware reference motions. Concurrently, we train a whole-body reference tracker with RL using this motion data. To improve robustness under imperfectly generated references, we further fine-tune the tracker with a frozen motion generator in a closed-loop setting. The resulting system supports directional goal-reaching control with terrain-aware whole-body adaptation, and can be deployed on a Unitree G1 humanoid robot with onboard perception and computation. The hardware experiments demonstrate successful traversal over boxes, hurdles, stairs, and mixed terrain combinations. Quantitative results further show the benefits of incorporating online motion generation and fine-tuning the motion tracker for improved generalization and robustness.

类人机器人运动生成强化学习地形适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。