arXiv:2602.12159cs.ROcs.AI2026-02被引 2

用3D高斯点云构建环境记忆,提升视觉语言模型导航推理能力

3DGSNav: Enhancing Vision-Language Model Reasoning for Object Navigation via Active 3D Gaussian Splatting

  • 将3D高斯点云作为视觉语言模型的持久记忆,实现动态环境建模
  • 在多个基准上达到领先性能,真实机器人实验验证了可靠性
  • 适合研究视觉语言导航、具身智能与主动感知的学者参考

物体导航是具身智能的核心能力,使智能体能在未知环境中定位目标物体。近年来,视觉语言模型(VLMs)推动了零样本物体导航(ZSON)的发展。然而,现有方法多依赖场景抽象,将环境转换为语义地图或文本表示,导致高层决策受限于底层感知精度。本文提出3DGSNav,一种新型的ZSON框架,将3D高斯点云(3DGS)嵌入为VLM的持久记忆,以增强空间推理能力。通过主动感知,3DGSNav增量构建环境的3DGS表示,支持轨迹引导的自由视角渲染和前沿感知的第一人称视图。我们设计结构化视觉提示,并与思维链(CoT)提示结合,进一步提升VLM推理能力。导航过程中,实时目标检测器筛选潜在目标,而由VLM驱动的主动视角切换进行目标再验证,确保高效可靠的识别。在多个基准上的大量评估及在四足机器人上的真实世界实验表明,该方法在鲁棒性和性能上均优于现有先进方法。

原文摘要 · Abstract (English)

Object navigation is a core capability of embodied intelligence, enabling an agent to locate target objects in unknown environments. Recent advances in vision-language models (VLMs) have facilitated zero-shot object navigation (ZSON). However, existing methods often rely on scene abstractions that convert environments into semantic maps or textual representations, causing high-level decision making to be constrained by the accuracy of low-level perception. In this work, we present 3DGSNav, a novel ZSON framework that embeds 3D Gaussian Splatting (3DGS) as persistent memory for VLMs to enhance spatial reasoning. Through active perception, 3DGSNav incrementally constructs a 3DGS representation of the environment, enabling trajectory-guided free-viewpoint rendering of frontier-aware first-person views. Moreover, we design structured visual prompts and integrate them with Chain-of-Thought (CoT) prompting to further improve VLM reasoning. During navigation, a real-time object detector filters potential targets, while VLM-driven active viewpoint switching performs target re-verification, ensuring efficient and reliable recognition. Extensive evaluations across multiple benchmarks and real-world experiments on a quadruped robot demonstrate that our method achieves robust and competitive performance against state-of-the-art approaches.The Project Page:https://aczheng-cai.github.io/3dgsnav.github.io/

视觉语言导航3D建模具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。