用生成先验和接触约束,让机器人在遮挡下也能准确重建物体形状。
Object Reconstruction under Occlusion with Generative Priors and Contact-induced Constraints
- 利用生成模型学习常见物体的形状先验,推测被遮挡部分
- 通过视频与物理交互获取接触信息,提供几何边界约束
- 结合接触引导生成,提升遮挡下的重建精度,适合机器人抓取场景
物体几何是机器人操作的关键信息,但因遮挡导致观测不完整,重建困难。本文引入两种额外信息源:一是生成模型学习常见物体的形状先验,可合理推测未见部分;二是从视频和物理交互中获取接触信息,提供几何边界的稀疏约束。通过接触引导的3D生成方法融合二者,其指导机制受拖拽式图像编辑启发。我们探索了不同引导策略,强调短梯度路径对生成效果的重要性。合成数据与真实世界实验表明,该方法优于纯3D生成和基于接触的优化方法,显著提升了遮挡下的物体重建性能。
原文摘要 · Abstract (English)
Object geometry is key information for robot manipulation. Yet, object reconstruction is a challenging task because camera observations are partial due to occlusions. The scene may not offer the flexibility for a robot to alter its viewpoint to obtain a full observation of the object of interest. In this paper, we leverage two extra sources of information to reduce the ambiguity of vision signals under occlusion. First, generative models learn priors of the shapes of commonly seen objects, allowing us to make reasonable guesses of the unseen part of geometry. Second, contact information, which can be obtained from videos and physical interactions, provides sparse constraints on the boundary of the geometry. We combine the two sources of information through contact-guided 3D generation. The guidance formulation is inspired by drag-based generative image editing. We explore different guidance strategies and highlight the importance of short gradient paths for guided generation. Experiments on synthetic and real-world data show that our approach improves the object reconstruction compared to pure 3D generation and contact-based optimization methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。