arXiv:2409.02084cs.ROcs.CV2024-09CoRL被引 59

用3D点云特征图实现快速零样本抓取,支持动态物体操作。

GraspSplats: Efficient Manipulation with 3D Feature Splatting

论文配图:GraspSplats: Efficient Manipulation with 3D Feature Splatting
图 1 · 摘自论文原文
  • 基于深度监督和新参考特征计算,60秒内生成高质量3D场景表示
  • 显式优化的几何结构支持实时抓取采样与可动物体精准操控
  • 适合需要快速适应新物体的机器人抓取任务,尤其在动态场景中表现优

机器人实现高效且零样本的物体部件抓取能力对实际应用至关重要,近年来视觉语言模型(VLMs)的发展推动了这一趋势。为弥合2D到3D表征之间的差距,现有方法依赖神经场(NeRFs)通过可微渲染或基于点的投影方式。然而,我们证明了NeRF因隐式特性不适用于场景变化,而点基方法在无渲染优化时难以准确进行部件定位。为此,我们提出GraspSplats:利用深度监督和一种新颖的参考特征计算方法,在60秒内生成高质量场景表示。进一步验证表明,基于高斯的表示形式使显式且优化的几何结构足以原生支持(1)实时抓取采样,(2)结合点追踪器实现动态与可动物体操作。在Franka机械臂上进行的大量实验显示,GraspSplats在多种任务设置下显著优于现有方法,尤其超越基于NeRF的F3RM和LERF-TOGO,以及2D检测方法。

原文摘要 · Abstract (English)

The ability for robots to perform efficient and zero-shot grasping of object parts is crucial for practical applications and is becoming prevalent with recent advances in Vision-Language Models (VLMs). To bridge the 2D-to-3D gap for representations to support such a capability, existing methods rely on neural fields (NeRFs) via differentiable rendering or point-based projection methods. However, we demonstrate that NeRFs are inappropriate for scene changes due to their implicitness and point-based methods are inaccurate for part localization without rendering-based optimization. To amend these issues, we propose GraspSplats. Using depth supervision and a novel reference feature computation method, GraspSplats generates high-quality scene representations in under 60 seconds. We further validate the advantages of Gaussian-based representation by showing that the explicit and optimized geometry in GraspSplats is sufficient to natively support (1) real-time grasp sampling and (2) dynamic and articulated object manipulation with point trackers. With extensive experiments on a Franka robot, we demonstrate that GraspSplats significantly outperforms existing methods under diverse task settings. In particular, GraspSplats outperforms NeRF-based methods like F3RM and LERF-TOGO, and 2D detection methods.

3D抓取高斯表示机器人操作动态物体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。