用高斯分布建模视觉语言模型的语义模糊性,提升户外导航安全性。
Estimating Semantic Ambiguity via Gaussian Context Distributions for VLM-Driven Traversability Analysis

- 将VLM输出转化为高斯分布,量化场景理解的不确定性
- 在GOOSE数据集上验证,不确定性与视觉干扰、地形重叠相关
- 适合需要安全决策的自动驾驶系统使用
在非结构化环境中实现自主导航需强大的场景理解能力,但视觉语言模型(VLM)常因语义模糊导致冲突预测,引发危险故障。为此,本文提出一种基于视觉的可通行性估计新方法,显式建模上下文不确定性。该方法通过概念锚定(Conceptual Anchoring),将开放词汇VLM的预测映射到连续的物理可通行性尺度上。将模型响应形式化为高斯上下文分布(GCD),基于分布的统计特性生成密集的可通行性图与不确定性图。在真实世界GOOSE数据集上的实验表明,所提出的不确定性度量能有效关联视觉伪影和混合地形重叠等模糊来源。该方法性能具有竞争力,且能提供统计不确定性估计,缓解语义模糊问题,从而在复杂室外环境中实现更安全可靠的自主行为。
原文摘要 · Abstract (English)
Autonomous navigation in unstructured environments requires robust scene understanding, yet Vision-Language Models (VLMs) often suffer from semantic ambiguity, where conflicting predictions can lead to dangerous failures. To address this, we present a novel pipeline for vision-based traversability estimation that explicitly models contextual uncertainty. Our approach utilizes Conceptual Anchoring to ground open-vocabulary VLM predictions onto a continuous physical traversability scale. By formulating the model's responses as a Gaussian Context Distribution (GCD), we derive both a dense traversability map and a dense uncertainty map based on the statistical properties of the distribution. Experimental validation on the real-world GOOSE dataset demonstrates that our proposed uncertainty metric effectively correlates with sources of ambiguity, such as visual artifacts and mixed terrain overlap. The method exhibits competitive performance while offering the distinct advantage of providing statistical uncertainty estimates to address semantic ambiguity, enabling safer and more reliable autonomous behavior in complex outdoor settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。