arXiv:2508.14098cs.ROcs.AI2025-08被引 1

让机器人短距离移动更自然高效,告别拖沓走路

No More Marching: Learning Humanoid Locomotion for Short-Range SE(2) Targets

  • 用新型奖励函数直接优化到目标姿态的运动
  • 能耗更低、到达更快、步数更少,性能全面超越现有方法
  • 适合需要精准短距移动的机器人应用,如工厂作业

在真实工作场景中,人形机器人常需执行以SE(2)为目标姿态的短距离移动任务。为实用,这些动作必须快速、鲁棒且节能。尽管基于学习的行走方法已取得进展,但多数方法优化的是速度追踪而非直接姿态抵达,导致应用于短距离任务时出现低效的“踱步”行为。本文提出一种强化学习方法,直接优化人形机器人对SE(2)目标的姿态到达。核心是设计一种基于星座的奖励函数,鼓励自然且高效的定向运动。我们构建了一个基准测试框架,评估能量消耗、到达时间与步数,在一系列SE(2)目标上进行测试。结果表明,该方法持续优于标准方法,并成功实现从仿真到硬件的迁移,凸显了针对性奖励设计对实际短距人形机器人运动的重要性。

原文摘要 · Abstract (English)

Humanoids operating in real-world workspaces must frequently execute task-driven, short-range movements to SE(2) target poses. To be practical, these transitions must be fast, robust, and energy efficient. While learning-based locomotion has made significant progress, most existing methods optimize for velocity-tracking rather than direct pose reaching, resulting in inefficient, marching-style behavior when applied to short-range tasks. In this work, we develop a reinforcement learning approach that directly optimizes humanoid locomotion for SE(2) targets. Central to this approach is a new constellation-based reward function that encourages natural and efficient target-oriented movement. To evaluate performance, we introduce a benchmarking framework that measures energy consumption, time-to-target, and footstep count on a distribution of SE(2) goals. Our results show that the proposed approach consistently outperforms standard methods and enables successful transfer from simulation to hardware, highlighting the importance of targeted reward design for practical short-range humanoid locomotion.

人形机器人强化学习运动控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。