arXiv:2509.15733cs.ROcs.AI2025-09被引 12

用多视角图像实现无需深度传感器的3D抓取

GP3: A 3D Geometry-Aware Policy with Multi-View Images for Robotic Manipulation

  • 通过多视角图像构建紧凑的3D场景表示
  • 在模拟和真实环境中均超越现有方法
  • 适合无深度传感器的机器人部署

有效的机器人操作依赖于对3D场景几何的精确理解,而多视角观测是获取此类几何信息的直接方式。为此,我们提出GP3——一种基于多视图输入的3D几何感知操作策略。GP3采用空间编码器从RGB观测中推断密集空间特征,从而估计深度与相机参数,生成针对操作任务定制的紧凑且表达性强的3D场景表征。该表征与语言指令融合,并通过轻量级策略头转化为连续动作。大量实验表明,GP3在仿真基准上持续优于当前最优方法。此外,GP3可有效迁移至无深度传感器或预先建图环境的真实机器人,仅需极少微调。这些结果表明,GP3是一种实用、传感器无关的几何感知机器人操作方案。

原文摘要 · Abstract (English)

Effective robotic manipulation relies on a precise understanding of 3D scene geometry, and one of the most straightforward ways to acquire such geometry is through multi-view observations. Motivated by this, we present GP3 -- a 3D geometry-aware robotic manipulation policy that leverages multi-view input. GP3 employs a spatial encoder to infer dense spatial features from RGB observations, which enable the estimation of depth and camera parameters, leading to a compact yet expressive 3D scene representation tailored for manipulation. This representation is fused with language instructions and translated into continuous actions via a lightweight policy head. Comprehensive experiments demonstrate that GP3 consistently outperforms state-of-the-art methods on simulated benchmarks. Furthermore, GP3 transfers effectively to real-world robots without depth sensors or pre-mapped environments, requiring only minimal fine-tuning. These results highlight GP3 as a practical, sensor-agnostic solution for geometry-aware robotic manipulation.

机器人操作多视角3D感知无深度传感器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。