arXiv:2410.01030cs.RO2024-10被引 2

用单一策略实现机器人动态走位与操作的平滑衔接,无需手动切换模式。

Preferenced Oracle Guided Multi-mode Policies for Dynamic Bipedal Loco-Manipulation

  • 通过混合自动机生成带连续动态和离散状态跳变的参考轨迹,指导策略学习。
  • 单个策略在足球和搬运任务中实现从追球到盘球再到射门的完整动作链。
  • 适用于不同体型机器人,同一奖励设置即可完成多种任务,通用性强。

动态走位操作需要有效的全身控制和与物体及环境的丰富接触交互。现有基于学习的控制合成方法依赖低层技能策略并配合高层策略或手工设计的有限状态机进行切换,导致行为接近静态。而像足球这样的动态任务要求机器人能跑向球、减速至最佳接球位置、持续盘球并最终射门——一系列连续流畅的动作。为此,我们提出偏好引导的多模式策略(OGMP),学习一个能掌握所有所需模式及期望转换顺序的单一策略,以解决单物体走位操作任务。我们设计混合自动机作为“预言者”,生成具有连续动力学和离散模式跳跃的参考轨迹,通过受限探索实现策略的引导优化。为强化期望的模式转换序列,引入任务无关的偏好奖励机制以提升性能。该方法在足球和全向搬运箱子等任务中成功实现全身控制下的走位操作。在足球任务中,单一策略自主学习从最优接球、过渡到接触丰富的盘球,到完成射门与停球的全过程。借助预言者的抽象能力,仅用相同的奖励定义和权重,即在不同形态的机器人(HECTOR V1、Berkeley Humanoid、Unitree G1 和 H1)上成功完成各类走位操作任务。

原文摘要 · Abstract (English)

Dynamic loco-manipulation calls for effective whole-body control and contact-rich interactions with the object and the environment. Existing learning-based control synthesis relies on training low-level skill policies and explicitly switching with a high-level policy or a hand-designed finite state machine, leading to quasi-static behaviors. In contrast, dynamic tasks such as soccer require the robot to run towards the ball, decelerate to an optimal approach to dribble, and eventually kick a goal - a continuum of smooth motion. To this end, we propose Preferenced Oracle Guided Multi-mode Policies (OGMP) to learn a single policy mastering all the required modes and preferred sequence of transitions to solve uni-object loco-manipulation tasks. We design hybrid automatons as oracles to generate references with continuous dynamics and discrete mode jumps to perform a guided policy optimization through bounded exploration. To enforce learning a desired sequence of mode transitions, we present a task-agnostic preference reward that enhances performance. The proposed approach demonstrates successful loco-manipulation for tasks like soccer and moving boxes omnidirectionally through whole-body control. In soccer, a single policy learns to optimally reach the ball, transition to contact-rich dribbling, and execute successful goal kicks and ball stops. Leveraging the oracle's abstraction, we solve each loco-manipulation task on robots with varying morphologies, including HECTOR V1, Berkeley Humanoid, Unitree G1, and H1, using the same reward definition and weights.

机器人控制多模式策略全身协调动态任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。