arXiv:2410.11584cs.ROcs.AI2024-10ICRA被引 9

用偏好学习提升柔性物体长程操作的数据效率

DeformPAM: Data-Efficient Learning for Long-horizon Deformable Object Manipulation via Preference-based Action Alignment

  • 分解任务为动作基元,结合点云与扩散模型建模动作分布
  • 仅用少量数据即实现更高任务完成率与效率
  • 适合数据稀缺的机器人柔性操作场景

近年来,模仿学习在机器人操作领域取得进展,但在处理具有复杂动态和多模态动作分布的长时序柔性物体操作任务时仍面临挑战。传统方法需大量数据,易出现分布偏移和误差累积。为此,我们提出基于偏好学习与奖励引导动作选择的数据高效通用框架DeformPAM。该方法将长时序任务分解为多个动作基元,采用3D点云输入与扩散模型建模动作分布,并利用人类偏好数据训练隐式奖励模型。推理阶段,奖励模型对多个候选动作评分并选择最优动作执行,有效减少异常动作,提升任务完成质量。在三个真实世界长时序柔性物体操作任务上的实验表明,即使数据有限,DeformPAM仍显著优于基线方法,在任务完成率和效率上均有提升。代码与数据将公开于https://deform-pam.robotflow.ai。

原文摘要 · Abstract (English)

In recent years, imitation learning has made progress in the field of robotic manipulation. However, it still faces challenges when addressing complex long-horizon tasks with deformable objects, such as high-dimensional state spaces, complex dynamics, and multimodal action distributions. Traditional imitation learning methods often require a large amount of data and encounter distributional shifts and accumulative errors in these tasks. To address these issues, we propose a data-efficient general learning framework (DeformPAM) based on preference learning and reward-guided action selection. DeformPAM decomposes long-horizon tasks into multiple action primitives, utilizes 3D point cloud inputs and diffusion models to model action distributions, and trains an implicit reward model using human preference data. During the inference phase, the reward model scores multiple candidate actions, selecting the optimal action for execution, thereby reducing the occurrence of anomalous actions and improving task completion quality. Experiments conducted on three challenging real-world long-horizon deformable object manipulation tasks demonstrate the effectiveness of this method. Results show that DeformPAM improves both task completion quality and efficiency compared to baseline methods even with limited data. Code and data will be available at https://deform-pam.robotflow.ai.

机器人操作模仿学习偏好学习柔性物体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。