将动作空间离散化,让机器人学会多任务操作
Discrete Policy: Learning Disentangled Action Space for Multi-Task Robotic Manipulation
- 用向量量化把动作序列映射到离散潜在空间,分离不同任务动作
- 在5个任务下比扩散策略高26%成功率,12个任务时差距达32.5%
- 适合想训练通用机器人的研究者和工程师
多任务机器人操作的视觉-运动策略学习是机器人领域的长期挑战。问题在于动作空间多样性:一个目标可通过多种方式达成,导致单任务对应多模态动作分布,任务越多,分布越复杂。本文提出「Discrete Policy」,一种可训练通用机器人的多任务操作方法。该方法采用向量量化将动作序列映射至离散潜在空间,促进任务专属代码的学习,并基于观测与语言指令重建动作。我们在仿真环境及多个真实机器人平台(单臂与双臂)上评估,结果表明,Discrete Policy 在五个任务的真实世界训练中,平均成功率比扩散策略高26%,比OpenVLA高15%;当任务数增至12时,与扩散策略的性能差距扩大至32.5%,充分展现其优势。实验证明,在潜在空间中学习多任务策略是实现通用机器人的关键一步。
原文摘要 · Abstract (English)
Learning visuomotor policy for multi-task robotic manipulation has been a long-standing challenge for the robotics community. The difficulty lies in the diversity of action space: typically, a goal can be accomplished in multiple ways, resulting in a multimodal action distribution for a single task. The complexity of action distribution escalates as the number of tasks increases. In this work, we propose \textbf{Discrete Policy}, a robot learning method for training universal agents capable of multi-task manipulation skills. Discrete Policy employs vector quantization to map action sequences into a discrete latent space, facilitating the learning of task-specific codes. These codes are then reconstructed into the action space conditioned on observations and language instruction. We evaluate our method on both simulation and multiple real-world embodiments, including both single-arm and bimanual robot settings. We demonstrate that our proposed Discrete Policy outperforms a well-established Diffusion Policy baseline and many state-of-the-art approaches, including ACT, Octo, and OpenVLA. For example, in a real-world multi-task training setting with five tasks, Discrete Policy achieves an average success rate that is 26\% higher than Diffusion Policy and 15\% higher than OpenVLA. As the number of tasks increases to 12, the performance gap between Discrete Policy and Diffusion Policy widens to 32.5\%, further showcasing the advantages of our approach. Our work empirically demonstrates that learning multi-task policies within the latent space is a vital step toward achieving general-purpose agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。