arXiv:2410.15879cs.RO2024-10被引 1

仅用一张彩色图实现快速精准6自由度抓取

Triplane Grasping: Efficient 6-DoF Grasping with Single RGB Images

  • 构建点与三平面混合3D表示,高效重建目标物体
  • 端到端网络直接从点云生成6-自由度抓取分布
  • 在两个数据集上验证了实时性与强泛化能力

可靠的物体抓取是机器人领域的一项基础任务。然而,仅依赖单张图像确定抓取姿态长期面临视觉信息有限和真实世界物体复杂性的挑战。本文提出Triplane Grasping,一种仅需单张RGB图像输入的快速抓取决策方法。该方法通过点解码器与三平面解码器联合生成混合的Triplane-Gaussian 3D表示,实现高效且高质量的待抓取物体重建,满足实时抓取需求。我们设计端到端网络,直接从点云中的3D点生成6-自由度平行爪抓取分布,并将抓取姿态锚定于观测数据中。在OmniObject3D和GraspNet-1Billion数据集上的实验表明,该方法能对日常物体实现快速建模与抓取姿态决策,并具备强泛化能力。

原文摘要 · Abstract (English)

Reliable object grasping is one of the fundamental tasks in robotics. However, determining grasping pose based on single-image input has long been a challenge due to limited visual information and the complexity of real-world objects. In this paper, we propose Triplane Grasping, a fast grasping decision-making method that relies solely on a single RGB-only image as input. Triplane Grasping creates a hybrid Triplane-Gaussian 3D representation through a point decoder and a triplane decoder, which produce an efficient and high-quality reconstruction of the object to be grasped to meet real-time grasping requirements. We propose to use an end-to-end network to generate 6-DoF parallel-jaw grasp distributions directly from 3D points in the point cloud as potential grasp contacts and anchor the grasp pose in the observed data. Experiments on the OmniObject3D and GraspNet-1Billion datasets demonstrate that our method achieves rapid modeling and grasping pose decision-making for daily objects, and strong generalization capability.

6-DoF抓取单图像3D重建机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。