arXiv:2504.10857cs.ROcs.CV2025-04CVPR被引 21

零样本重建与抓取联合优化,提升机器人抓握精度与安全性

ZeroGrasp: Zero-Shot Shape Reconstruction Enabled Robotic Grasping

  • 联合3D重建与抓取预测,利用遮挡推理增强空间理解
  • 在GraspNet-1B上达顶尖性能,可泛化至真实世界新物体
  • 基于百万级合成数据训练,支持零样本场景适应

机器人抓取是具身系统的核心能力。现有方法常直接从局部信息输出抓取动作,忽略场景几何建模,导致运动不佳甚至碰撞。为此,我们提出ZeroGrasp框架,实现近实时的3D重建与抓取姿态联合预测。关键洞察在于:遮挡推理与物体间空间关系建模对重建与抓取均有帮助。我们构建了一个大规模合成数据集,包含100万张照片级真实感图像、高分辨率3D重建以及来自Objaverse-LVIS数据集的12,000个物体的113亿条物理合理抓取标注。在GraspNet-1B基准和真实机器人实验中,ZeroGrasp达到当前最优性能,并通过合成数据实现对新实物的零样本泛化。

原文摘要 · Abstract (English)

Robotic grasping is a cornerstone capability of embodied systems. Many methods directly output grasps from partial information without modeling the geometry of the scene, leading to suboptimal motion and even collisions. To address these issues, we introduce ZeroGrasp, a novel framework that simultaneously performs 3D reconstruction and grasp pose prediction in near real-time. A key insight of our method is that occlusion reasoning and modeling the spatial relationships between objects is beneficial for both accurate reconstruction and grasping. We couple our method with a novel large-scale synthetic dataset, which comprises 1M photo-realistic images, high-resolution 3D reconstructions and 11.3B physically-valid grasp pose annotations for 12K objects from the Objaverse-LVIS dataset. We evaluate ZeroGrasp on the GraspNet-1B benchmark as well as through real-world robot experiments. ZeroGrasp achieves state-of-the-art performance and generalizes to novel real-world objects by leveraging synthetic data.

机器人抓取3D重建零样本学习合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。