arXiv:2411.14869cs.CVcs.AI2024-11CVPR被引 9

用图像特征提升3D感知,让智能体更懂环境。

BIP3D: Bridging 2D Images and 3D Perception for Embodied Intelligence

  • 以图像为中心,融合预训练视觉模型与3D位置编码
  • 在EmbodiedScan上3D检测提升5.69%,视觉定位提升15.25%
  • 适合做机器人导航、虚拟助手等需要空间理解的任务

在具身智能系统中,3D感知是核心组件,使智能体能够理解周围环境。以往方法主要依赖点云,尽管几何信息精确,但受限于固有的稀疏性、噪声和数据稀缺性。本文提出一种新型图像中心的3D感知模型BIP3D,利用丰富的图像特征与显式的3D位置编码,克服点云方法的局限。具体而言,我们借助预训练的2D视觉基础模型增强语义理解,并引入空间增强模块提升空间感知能力。二者结合实现多视角、多模态特征融合与端到端3D感知。实验表明,BIP3D在EmbodiedScan基准上优于当前最先进方法,在3D检测任务中提升5.69%,在3D视觉定位任务中提升15.25%。

原文摘要 · Abstract (English)

In embodied intelligence systems, a key component is 3D perception algorithm, which enables agents to understand their surrounding environments. Previous algorithms primarily rely on point cloud, which, despite offering precise geometric information, still constrain perception performance due to inherent sparsity, noise, and data scarcity. In this work, we introduce a novel image-centric 3D perception model, BIP3D, which leverages expressive image features with explicit 3D position encoding to overcome the limitations of point-centric methods. Specifically, we leverage pre-trained 2D vision foundation models to enhance semantic understanding, and introduce a spatial enhancer module to improve spatial understanding. Together, these modules enable BIP3D to achieve multi-view, multi-modal feature fusion and end-to-end 3D perception. In our experiments, BIP3D outperforms current state-of-the-art results on the EmbodiedScan benchmark, achieving improvements of 5.69% in the 3D detection task and 15.25% in the 3D visual grounding task.

3D感知具身智能图像特征视觉定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。