arXiv:2605.26503cs.CV2026-05被引 3

让智能体主动感知视觉语言导航中的不确定性,提升决策可靠性。

Uncertainty-Aware Gaussian Map for Vision-Language Navigation

论文配图:Uncertainty-Aware Gaussian Map for Vision-Language Navigation
图 1 · 摘自论文原文
  • 构建语义高斯地图,用可微3D高斯体编码环境结构与语义
  • 量化几何、语义、外观三类感知不确定性,提升导航鲁棒性
  • 适合关注复杂场景下智能体决策可靠性的研究者

视觉-语言导航(VLN)要求智能体根据自然语言指令在3D环境中导航。现有方法常忽略感知不确定性,如缺乏可靠定位证据或空间线索歧义,却仍直接预测动作。本文显式建模几何、语义和外观三类感知不确定性,并将其融入智能体观测空间以支持更明智的决策。具体地,智能体首先基于全景观测构建可微的语义高斯地图(SGM),由3D高斯原语组成,编码环境的几何结构与语义内容;在此基础上,通过高斯位置与尺度的变分扰动估计几何不确定性;通过扰动高斯语义属性捕捉语义不确定性;利用Fisher信息衡量渲染观测对高斯级变化的敏感度,表征外观不确定性。这些不确定性被融合进SGM,扩展为统一的3D价值图,将其作为可操作的提示与约束,支持可靠导航。在多个VLN基准上的综合评估验证了该方法的有效性。

原文摘要 · Abstract (English)

Vision-Language Navigation (VLN) requires an agent to navigate 3D environments following natural language instructions. During navigation, existing agents commonly encounter perceptual uncertainty, such as insufficient evidence for reliable grounding or ambiguity in interpreting spatial cues, yet they typically ignore such information when predicting actions. In this work, we explicitly model three forms of perceptual uncertainty (i.e., geometric, semantic, and appearance uncertainty) and integrate them into the agent's observation space to enable informed decision-making. Concretely, our agent first constructs a Semantic Gaussian Map (SGM), composed of differentiable 3D Gaussian primitives initialized from panoramic observations, that encodes both the geometric structure and semantic content of the environment. On top of SGM, geometric uncertainty is estimated through variational perturbations of Gaussian position and scale to assess structural reliability; semantic uncertainty is captured by perturbing Gaussian semantic attributes to reveal ambiguous interpretations; and appearance uncertainty is characterized by Fisher Information, which measures the sensitivity of rendered observations to Gaussian-level variations. These uncertainties are incorporated into SGM, extending it into a unified 3D Value Map, which grounds them as affordances and constraints that support reliable navigation. Comprehensive evaluations across multiple VLN benchmarks show the effectiveness of our agent.

视觉语言导航不确定性建模高斯地图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。