arXiv:2605.22036cs.CVcs.AI2026-05

用3D几何信息构建紧凑视觉语言导航表示,提升效率与空间推理能力。

GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation

论文配图:GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation
图 1 · 摘自论文原文
  • 基于RGB-D输入构建代理中心的3D BEV地图,融合显式深度与隐式结构先验。
  • 仅用导航数据即达顶尖性能,无需DAgger或混合训练,降低数据依赖。
  • 适合追求高效、高精度视觉语言导航系统的研究者与开发者。

尽管视觉语言导航(VLN)取得显著进展,现有方法仍依赖密集的RGB视频,产生大量补丁标记且缺乏显式空间结构,导致计算开销大、空间推理能力有限。为此,我们提出几何感知鸟瞰图(GA-BEV)——一种紧凑的、基于3D的多模态大语言模型(MLLM)导航特征表示。通过将视觉特征投影至3D空间并聚合为以代理为中心的布局,构建保持几何一致性的BEV空间地图,有效减少标记冗余。为进一步丰富几何理解,引入预训练3D基础模型的特征,注入从大规模3D重建任务中学习到的结构先验。显式深度投影与隐式学习先验相结合,生成紧凑而空间表达力强的表示,显著提升导航效率与性能。实验表明,该方法仅使用导航数据即可达到当前最优结果,无需DAgger增强或混合视觉问答训练,验证了所提GA-VLN框架的鲁棒性与数据效率。

原文摘要 · Abstract (English)

Despite significant progress in Vision-Language Navigation (VLN), existing approaches still rely on dense RGB videos that produce excessive patch tokens and lack explicit spatial structure, resulting in substantial computational overhead and limited spatial reasoning. To address these issues, we introduce the Geometry-Aware BEV (GA-BEV) - a compact, 3D-grounded feature representation that integrates both explicit and implicit geometric cues into multimodal large language model (MLLM) - based navigation systems. We construct BEV spatial maps from RGB-D inputs by projecting visual features into 3D space and aggregating them into an agent-centric layout that preserves geometric consistency while reducing token redundancy. To further enrich geometric understanding, we incorporate features from a pretrained 3D foundation model into the BEV space, injecting structural priors learned from large-scale 3D reconstruction tasks. Together, these complementary cues - explicit depth-based projection and implicit learned priors - yield compact yet spatially expressive representations that substantially improve navigation efficiency and performance. Experiments show that our method achieves state-of-the-art results using only navigation data, without DAgger augmentation or mixed VQA training, demonstrating the robustness and data efficiency of the proposed GA-VLN framework.

视觉导航3D表示多模态高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。