arXiv:2606.08440cs.ROcs.CV2026-06

用3D基础模型提升机器人抓取,同时重建物体形状。

GraspFoM: Towards Reconstruction-Driven Robotic Grasping with 3D Foundation Priors

论文配图:GraspFoM: Towards Reconstruction-Driven Robotic Grasping with 3D Foundation Priors
图 1 · 摘自论文原文
  • 共享3D隐空间统一重建与抓取预测
  • 仅增少量参数即实现顶尖抓取与重建效果
  • 适合需要高精度抓取和3D重建的机器人任务

机器人抓取是机器人操作的基础能力,但在部分观测下仍具挑战。可靠抓取依赖局部接触线索和物体级3D结构。现有几何感知抓取方法虽重视重建,但通常将几何视为中间预测而非可复用的对象先验。本文提出GraspFoM,一个统一框架,利用3D基础先验(SAM3D)构建共享3D物体隐空间,用于重建与抓取姿态预测。基于该隐空间,设计锚点初始化的截断姿态推理扩散器,无需依赖离散抓取候选即可预测连续且多模态抓取姿态。进一步通过重建感知评分器与残差隐空间更新器,探究重建与抓取的交互:重建提供几何线索,抓取监督则引导隐空间向抓取相关属性优化。GraspFoM联合预测抓取姿态并以网格和3DGS形式重建高保真3D资产。大量实验表明,该方法在重建与抓取任务上均达到当前最优,且仅增加少量可训练参数。组件消融实验验证了各模块的有效性。

原文摘要 · Abstract (English)

Robotic grasping is a fundamental capability in robotic manipulation. Yet grasping remains challenging under partial observations. Reliable grasping depends on both local contact cues and object-level 3D structure. Existing geometry-aware grasping methods recognize the value of reconstruction, but they typically treat geometry as an intermediate prediction rather than a reusable object prior for grasping. In this paper, we present GraspFoM, a unified framework that leverages 3D foundation priors (SAM3D) to build a shared 3D object latent for both reconstruction and grasp pose prediction. Built on this shared object latent, we introduce an anchor-initialized truncated pose-reasoning diffuser that predicts continuous and multimodal grasp poses without directly relying on discrete grasp candidates. We further investigate the interaction between reconstruction and grasping through a reconstruction-aware scorer and a residual latent updater. Reconstruction provides grounded geometric cues, while grasp supervision refines the shared object latent toward grasp-relevant affordances. GraspFoM jointly predicts grasp poses and reconstructs high-fidelity 3D assets in mesh and 3DGS forms. Comprehensive experiments demonstrate that GraspFoM achieves state-of-the-art results on both reconstruction and grasping. Notably, these improvements require only a small number of additional trainable parameters. Component-wise ablation studies also demonstrate the contribution of each component.

机器人抓取3D重建扩散模型基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。