用生成模型补全部分视角图像中的3D物体,让机器人在杂乱中精准抓取
DreamGrasp: Zero-Shot 3D Multi-Object Reconstruction from Partial-View Images for Robotic Manipulation
- 利用生成模型想象缺失视角,补全物体3D结构
- 在多物体场景中实现高精度几何重建,成功率超90%
- 无需标注数据,适合真实复杂环境的机器人操作
部分视角3D识别——从少量稀疏的RGB图像中重建3D几何结构并识别物体实例——在杂乱、遮挡的真实场景中极具挑战性,且至关重要,因为完整视角或可靠深度数据常不可得。现有方法依赖强对称性假设或在精心标注数据集上进行监督学习,难以泛化。本文提出DreamGrasp框架,利用大规模预训练图像生成模型的想象力,推断场景中未观测部分。通过粗粒度3D重建、基于对比学习的实例分割,以及文本引导的实例级细化,DreamGrasp克服了先前方法的局限,实现在复杂多物体环境中的鲁棒3D重建。实验表明,DreamGrasp不仅能准确恢复物体几何,还支持下游任务如序列去杂和目标检索,成功率高达90%以上。
原文摘要 · Abstract (English)
Partial-view 3D recognition -- reconstructing 3D geometry and identifying object instances from a few sparse RGB images -- is an exceptionally challenging yet practically essential task, particularly in cluttered, occluded real-world settings where full-view or reliable depth data are often unavailable. Existing methods, whether based on strong symmetry priors or supervised learning on curated datasets, fail to generalize to such scenarios. In this work, we introduce DreamGrasp, a framework that leverages the imagination capability of large-scale pre-trained image generative models to infer the unobserved parts of a scene. By combining coarse 3D reconstruction, instance segmentation via contrastive learning, and text-guided instance-wise refinement, DreamGrasp circumvents limitations of prior methods and enables robust 3D reconstruction in complex, multi-object environments. Our experiments show that DreamGrasp not only recovers accurate object geometry but also supports downstream tasks like sequential decluttering and target retrieval with high success rates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。