arXiv:2512.16449cs.RO2025-12被引 2

用扩散模型补全单视角下的物体形状,提升杂乱环境中的抓取成功率。

Single-View Shape Completion for Robotic Grasping in Clutter

  • 基于单视角深度图,用扩散模型生成完整3D形状。
  • 在杂乱场景中抓取成功率比基线高23%,比现有方法高19%。
  • 专为家庭常见物品设计,适合真实机器人抓取任务。

在视觉驱动的机器人操作中,单摄像头只能获取物体部分外观,杂乱场景中的遮挡进一步限制可见性,导致几何信息不完整,影响抓取算法性能。为此,本文利用扩散模型,从单视图的局部深度观测中进行类别级3D形状补全,重建完整物体几何,为抓取规划提供更丰富上下文。方法聚焦于具有多样几何结构的家庭常见物品,生成的完整3D形状作为下游抓取推理网络输入。与以往主要关注孤立物体或少量遮挡的工作不同,本文在真实杂乱场景下评估形状补全与抓取表现。初步实验表明,在杂乱场景中,该方法抓取成功率比无形状补全的基线高出23%,优于近期先进形状补全方法19%。代码已公开于https://amm.aass.oru.se/shape-completion-grasping/。

原文摘要 · Abstract (English)

In vision-based robot manipulation, a single camera view can only capture one side of objects of interest, with additional occlusions in cluttered scenes further restricting visibility. As a result, the observed geometry is incomplete, and grasp estimation algorithms perform suboptimally. To address this limitation, we leverage diffusion models to perform category-level 3D shape completion from partial depth observations obtained from a single view, reconstructing complete object geometries to provide richer context for grasp planning. Our method focuses on common household items with diverse geometries, generating full 3D shapes that serve as input to downstream grasp inference networks. Unlike prior work, which primarily considers isolated objects or minimal clutter, we evaluate shape completion and grasping in realistic clutter scenarios with household objects. In preliminary evaluations on a cluttered scene, our approach consistently results in better grasp success rates than a naive baseline without shape completion by 23% and over a recent state of the art shape completion approach by 19%. Our code is available at https://amm.aass.oru.se/shape-completion-grasping/.

3D补全机器人抓取扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。