arXiv:2505.20619cs.RO2025-05被引 11

一个统一框架让机器人学会站立、行走、跑步及自然过渡。

Gait-Conditioned Reinforcement Learning with Multi-Phase Curriculum for Humanoid Locomotion

  • 用单个策略同时控制多种步态,通过奖励路由动态激活对应目标。
  • 在仿真和真实机器人上实现稳定行走、跑步及站-走转换。
  • 无需动作捕捉数据,适合需要自然运动的复杂场景控制。

我们提出一种统一的步态条件强化学习框架,使类人机器人能够在单一循环策略下完成站立、行走、跑步及平滑过渡。采用紧凑的奖励路由机制,根据独热编码的步态标识动态激活特定目标,减轻奖励干扰,支持稳定的多步态学习。引入类人奖励项,促进生物力学自然的动作,如直膝支撑与协调的臂腿摆动,无需动作捕捉数据。设计分阶段结构化课程,逐步引入步态复杂性并扩展命令空间。在仿真中,该策略成功实现稳健的站立、行走、跑步及步态切换;在真实 Unitree G1 类人机器人上,验证了站立、行走及走-站转换,展现出稳定协调的运动能力。本工作为多样化模式与环境下的灵活、自然类人控制提供了一种可扩展、无参考的解决方案。

原文摘要 · Abstract (English)

We present a unified gait-conditioned reinforcement learning framework that enables humanoid robots to perform standing, walking, running, and smooth transitions within a single recurrent policy. A compact reward routing mechanism dynamically activates gait-specific objectives based on a one-hot gait ID, mitigating reward interference and supporting stable multi-gait learning. Human-inspired reward terms promote biomechanically natural motions, such as straight-knee stance and coordinated arm-leg swing, without requiring motion capture data. A structured curriculum progressively introduces gait complexity and expands command space over multiple phases. In simulation, the policy successfully achieves robust standing, walking, running, and gait transitions. On the real Unitree G1 humanoid, we validate standing, walking, and walk-to-stand transitions, demonstrating stable and coordinated locomotion. This work provides a scalable, reference-free solution toward versatile and naturalistic humanoid control across diverse modes and environments.

强化学习步态控制类人机器人多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。