arXiv:2505.11383cs.CVcs.RO2025-05NeurIPS被引 28

动态分层3D标记提升视觉语言导航的环境理解与长期记忆能力

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation

  • 构建动态分层3D表示,融合语义与几何信息,支持实时更新
  • 在单目设置下刷新R2R-CE、REVERIE-CE等基准性能纪录
  • 适用于机器人长时探索与真实场景部署,具备强实用性

视觉-语言导航(VLN)是智能体基于自然语言指令,在3D环境中移动至目标位置的核心任务。尽管具备强泛化能力的视频-语言大模型(Video-VLM)已在该任务上表现优异,但其在真实3D导航中仍面临三大挑战:对3D几何与空间语义理解不足;大规模探索与长期环境记忆能力有限;对动态变化环境适应性差。为此,我们提出Dynam3D,一种动态分层3D表征模型,将语言对齐、可泛化、分层的3D表示作为视觉输入,用于训练3D-VLM进行导航动作预测。给定单目RGB-D图像,Dynam3D将2D CLIP特征投影至3D空间,构建多层级3D patch-instance-zone表示,实现对3D几何与语义的深度理解,并采用动态分层更新策略。该模型可在线编码与定位3D实例,动态更新以应对环境变化,从而支持大规模探索与长期记忆。通过大规模3D-语言预训练与特定任务微调,Dynam3D在单目设置下于R2R-CE、REVERIE-CE和NavRAG-CE等基准上达到新SOTA。预探索、终身记忆及真实机器人实验验证了其实际部署的有效性。

原文摘要 · Abstract (English)

Vision-and-Language Navigation (VLN) is a core task where embodied agents leverage their spatial mobility to navigate in 3D environments toward designated destinations based on natural language instructions. Recently, video-language large models (Video-VLMs) with strong generalization capabilities and rich commonsense knowledge have shown remarkable performance when applied to VLN tasks. However, these models still encounter the following challenges when applied to real-world 3D navigation: 1) Insufficient understanding of 3D geometry and spatial semantics; 2) Limited capacity for large-scale exploration and long-term environmental memory; 3) Poor adaptability to dynamic and changing environments.To address these limitations, we propose Dynam3D, a dynamic layered 3D representation model that leverages language-aligned, generalizable, and hierarchical 3D representations as visual input to train 3D-VLM in navigation action prediction. Given posed RGB-D images, our Dynam3D projects 2D CLIP features into 3D space and constructs multi-level 3D patch-instance-zone representations for 3D geometric and semantic understanding with a dynamic and layer-wise update strategy. Our Dynam3D is capable of online encoding and localization of 3D instances, and dynamically updates them in changing environments to provide large-scale exploration and long-term memory capabilities for navigation. By leveraging large-scale 3D-language pretraining and task-specific adaptation, our Dynam3D sets new state-of-the-art performance on VLN benchmarks including R2R-CE, REVERIE-CE and NavRAG-CE under monocular settings. Furthermore, experiments for pre-exploration, lifelong memory, and real-world robot validate the effectiveness of practical deployment.

视觉语言导航3D表示动态建模机器人导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。