arXiv:2409.00499cs.ROcs.CV2024-09中稿 · IROS2024被引 2

用扩散模型解决多模态物品存放的精准定位问题。

DAP: Diffusion-based Affordance Prediction for Multi-modality Storage

  • 基于扩散模型分两步预测物体在容器中的放置区域与相对位姿。
  • 在基准测试中性能超越现有方法,训练效率更高。
  • 适合需要真实场景适应性的机器人抓取与存储任务。

解决物品精确放置于容器中的存储问题,需实现精细的6D操作,并面对解空间的多模态特性——同一容器存在多个有效放置配置。本文提出一种基于扩散模型的仿生预判(DAP)流程,采用两阶段方法:首先识别容器上的可放置区域,再精确计算物体与该区域的相对位姿。现有方法或难以处理多模态问题,或训练成本过高。实验表明,DAP在RPDiff基准上表现优于当前最优方法RPDiff,且具备更高的训练效率;同时在真实场景中展现出优异的数据效率,优于依赖仿真的现有方案。本工作填补了机器人操作研究中对高效、鲁棒多模态存储解决方案的空白。代码与补充材料见:https://github.com/changhaonan/DPS.git。

原文摘要 · Abstract (English)

Solving storage problem: where objects must be accurately placed into containers with precise orientations and positions, presents a distinct challenge that extends beyond traditional rearrangement tasks. These challenges are primarily due to the need for fine-grained 6D manipulation and the inherent multi-modality of solution spaces, where multiple viable goal configurations exist for the same storage container. We present a novel Diffusion-based Affordance Prediction (DAP) pipeline for the multi-modal object storage problem. DAP leverages a two-step approach, initially identifying a placeable region on the container and then precisely computing the relative pose between the object and that region. Existing methods either struggle with multi-modality issues or computation-intensive training. Our experiments demonstrate DAP's superior performance and training efficiency over the current state-of-the-art RPDiff, achieving remarkable results on the RPDiff benchmark. Additionally, our experiments showcase DAP's data efficiency in real-world applications, an advancement over existing simulation-driven approaches. Our contribution fills a gap in robotic manipulation research by offering a solution that is both computationally efficient and capable of handling real-world variability. Code and supplementary material can be found at: https://github.com/changhaonan/DPS.git.

机器人操作扩散模型多模态存储规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。