arXiv:2603.25902cs.RO2026-03被引 4

让机器人像人一样动态奔跑,还能自主避障。

Chasing Autonomy: Dynamic Retargeting and Control Guided RL for Performant and Controllable Humanoid Running

  • 通过优化人类动作生成周期性参考轨迹,实现动态重定向。
  • 实测达到3.3米/秒速度,连续跑数百米不中断。
  • 适合需要高速、长时自主行走的机器人研究者。

类人机器人有望实现类似人类的运动,包括快速而动态的奔跑。近年来,基于强化学习(RL)的控制器因其能生成高度动态的行为而受到关注,但通常仅限于单一动作回放,限制了其在长时间自主运动中的应用。本文提出一种管道方法,通过带硬约束的优化过程,从单次人类示范中动态重定向并生成改进的周期性参考轨迹库。我们研究了参考运动与奖励结构对参考速度和指令速度跟踪的影响,发现基于目标条件且受控制引导的奖励函数,结合动态优化的人类数据,可取得最佳性能。该策略已在硬件上部署,实现在Unitree G1机器人上最高达3.3米/秒的跑步速度,并在真实环境中连续运行数百米。此外,为验证运动可控性,将该控制器集成至完整感知与规划自主系统中,在户外奔跑时成功实现障碍物避让。

原文摘要 · Abstract (English)

Humanoid robots have the promise of locomoting like humans, including fast and dynamic running. Recently, reinforcement learning (RL) controllers that can mimic human motions have become popular as they can generate very dynamic behaviors, but they are often restricted to single motion play-back which hinders their deployment in long duration and autonomous locomotion. In this paper, we present a pipeline to dynamically retarget human motions through an optimization routine with hard constraints to generate improved periodic reference libraries from a single human demonstration. We then study the effect of both the reference motion and the reward structure on the reference and commanded velocity tracking, concluding that a goal-conditioned and control-guided reward which tracks dynamically optimized human data results in the best performance. We deploy the policy on hardware, demonstrating its speed and endurance by achieving running speeds of up to 3.3 m/s on a Unitree G1 robot and traversing hundreds of meters in real-world environments. Additionally, to demonstrate the controllability of the locomotion, we use the controller in a full perception and planning autonomy stack for obstacle avoidance while running outdoors.

人形机器人强化学习自主运动动态行走

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。