arXiv:2509.13239cs.RO2025-09被引 7

用动态奖励课程训练双足机器人协作完成抓放任务

Collaborative Loco-Manipulation for Pick-and-Place Tasks with Dynamic Reward Curriculum

  • 设计动态奖励课程,分阶段引导机器人完成长时序抓放
  • 仿真中训练效率提升55%,执行时间减少18.6%
  • 支持双机器人自主注意力切换,实现高效协作

我们提出一种分层强化学习流水线,用于训练单臂双足机器人在单一和双机器人协作场景下端到端完成抓放任务——从接近目标物体到将其放置至指定区域。引入一种新型动态奖励课程,使单一策略能通过逐步引导代理完成以物体为中心的子目标,高效学习长时序抓放操作。相比现有长时序强化学习方法,本方法在仿真中训练效率提升55%,执行时间减少18.6%。在双机器人情况下,我们的策略使每个机器人能在不同任务阶段关注观察空间的不同部分,通过自主注意力转移促进有效协作。我们在ANYmal D平台上通过真实世界实验验证了该方法在单机器人和双机器人场景下的有效性。据我们所知,这是首个针对双足机械臂协作抓放任务的完整强化学习流水线。

原文摘要 · Abstract (English)

We present a hierarchical RL pipeline for training one-armed legged robots to perform pick-and-place (P&P) tasks end-to-end -- from approaching the payload to releasing it at a target area -- in both single-robot and cooperative dual-robot settings. We introduce a novel dynamic reward curriculum that enables a single policy to efficiently learn long-horizon P&P operations by progressively guiding the agents through payload-centered sub-objectives. Compared to state-of-the-art approaches for long-horizon RL tasks, our method improves training efficiency by 55% and reduces execution time by 18.6% in simulation experiments. In the dual-robot case, we show that our policy enables each robot to attend to different components of its observation space at distinct task stages, promoting effective coordination via autonomous attention shifts. We validate our method through real-world experiments using ANYmal D platforms in both single- and dual-robot scenarios. To our knowledge, this is the first RL pipeline that tackles the full scope of collaborative P&P with two legged manipulators.

强化学习机器人协作抓放任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。