arXiv:2409.09725cs.ROcs.CV2024-09被引 3

用扩散模型提升机器人抓取旋转精度,实现实时高精度操作

Precise Pick-and-Place using Score-Based Diffusion Networks

  • 分步扩散网络从图像中连续估计物体位姿
  • 仿真与真实场景下抓放成功率显著提升
  • 适合对旋转精度要求高的机器人操作任务

本文提出一种新型的粗到精连续位姿扩散方法,用于提升机器人操作中抓取与放置任务的精度。利用扩散网络的能力,实现对物体位姿的精确感知,从而提高抓放成功率与整体操作精度。该方法基于来自RGB-D相机的俯视RGB图像,采用粗到精架构,有效学习粗粒度与细粒度模型。其关键特点是关注连续位姿估计,尤其在旋转角度上表现更优。此外,通过引入位姿与颜色增强技术,在数据有限条件下实现有效训练。在模拟与真实场景中进行了大量实验及消融研究,全面评估所提方法的有效性。结果表明,该方法能有效实现高精度抓放任务。

原文摘要 · Abstract (English)

In this paper, we propose a novel coarse-to-fine continuous pose diffusion method to enhance the precision of pick-and-place operations within robotic manipulation tasks. Leveraging the capabilities of diffusion networks, we facilitate the accurate perception of object poses. This accurate perception enhances both pick-and-place success rates and overall manipulation precision. Our methodology utilizes a top-down RGB image projected from an RGB-D camera and adopts a coarse-to-fine architecture. This architecture enables efficient learning of coarse and fine models. A distinguishing feature of our approach is its focus on continuous pose estimation, which enables more precise object manipulation, particularly concerning rotational angles. In addition, we employ pose and color augmentation techniques to enable effective training with limited data. Through extensive experiments in simulated and real-world scenarios, as well as an ablation study, we comprehensively evaluate our proposed methodology. Taken together, the findings validate its effectiveness in achieving high-precision pick-and-place tasks.

机器人操作扩散模型位姿估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。