arXiv:2411.17030cs.CVcs.AI2024-11CVPR被引 29

让机器人在新环境中实时理解语言指令并导航

g3D-LF: Generalizable 3D-Language Feature Fields for Embodied Tasks

  • 用多尺度编码融合语义与空间信息,生成可泛化的3D语言特征场
  • 在全景和单目设置下,导航任务成功率显著提升,零样本物体导航效果优异
  • 适合需要跨场景理解语言指令的智能体应用,如机器人导航与交互

我们提出通用3D语言特征场(g3D-LF),一种在大规模3D-语言数据集上预训练的3D表示模型,用于具身任务。该模型处理智能体的带姿态RGB-D图像,编码特征场以实现:1)从3D场景任意位置预测新视角;2)生成以智能体为中心的鸟瞰图(BEV);3)在上述表示中通过多粒度语言查询目标。该表示可泛化至未见环境,支持实时构建与动态更新。通过沿采样射线进行体渲染并利用多尺度编码器整合语义与空间关系,g3D-LF通过多层次对比学习生成多尺度、多视角且对齐多粒度语言的表示。我们还构建了一个大规模3D-语言数据集,以对齐特征场与语言表示。在全景与单目设置下的视觉-语言导航、零样本物体导航及情境问答任务中的大量实验,验证了g3D-LF在具身任务中的显著优势与有效性。

原文摘要 · Abstract (English)

We introduce Generalizable 3D-Language Feature Fields (g3D-LF), a 3D representation model pre-trained on large-scale 3D-language dataset for embodied tasks. Our g3D-LF processes posed RGB-D images from agents to encode feature fields for: 1) Novel view representation predictions from any position in the 3D scene; 2) Generations of BEV maps centered on the agent; 3) Querying targets using multi-granularity language within the above-mentioned representations. Our representation can be generalized to unseen environments, enabling real-time construction and dynamic updates. By volume rendering latent features along sampled rays and integrating semantic and spatial relationships through multiscale encoders, our g3D-LF produces representations at different scales and perspectives, aligned with multi-granularity language, via multi-level contrastive learning. Furthermore, we prepare a large-scale 3D-language dataset to align the representations of the feature fields with language. Extensive experiments on Vision-and-Language Navigation under both Panorama and Monocular settings, Zero-shot Object Navigation, and Situated Question Answering tasks highlight the significant advantages and effectiveness of our g3D-LF for embodied tasks.

具身智能3D表示语言对齐导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。