arXiv:2410.23643cs.RO2024-10被引 20

从单张图像重建完整3D场景,提升机器人抓取稳定性

SceneComplete: Open-World 3D Scene Completion in Cluttered Real World Environments for Robot Manipulation

  • 融合多类预训练模型,构建端到端3D场景补全流程
  • 在真实杂乱环境中实现高精度整体物体重建
  • 适合需精准环境理解的机器人抓取任务研究者

日常杂乱环境中精准执行机器人操作需要对三维场景有准确理解,以稳定可靠地抓取和放置物体,并避免碰撞。通常需基于有限输入(如单张RGB-D图像)构建复杂场景的3D表征。本文提出SceneComplete系统,可从单视角生成完整、分割清晰的3D场景模型。该系统采用新颖流水线,整合通用预训练感知模块(视觉-语言、分割、图像修复、图像到3D、视觉描述符与位姿估计),实现高精度结果。我们在大型基准数据集上验证其相对于真值模型的准确性,并证明其精确的整体物体重建能有效生成鲁棒抓取建议,包括对灵巧手的应用。代码与附加结果已发布于官网。

原文摘要 · Abstract (English)

Careful robot manipulation in every-day cluttered environments requires an accurate understanding of the 3D scene, in order to grasp and place objects stably and reliably and to avoid colliding with other objects. In general, we must construct such a 3D interpretation of a complex scene based on limited input, such as a single RGB-D image. We describe SceneComplete, a system for constructing a complete, segmented, 3D model of a scene from a single view. SceneComplete is a novel pipeline for composing general-purpose pretrained perception modules (vision-language, segmentation, image-inpainting, image-to-3D, visual-descriptors and pose-estimation) to obtain highly accurate results. We demonstrate its accuracy and effectiveness with respect to ground-truth models in a large benchmark dataset and show that its accurate whole-object reconstruction enables robust grasp proposal generation, including for a dexterous hand. We release the code and additional results on our website.

3D重建机器人操作场景补全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。