arXiv:2606.31407cs.CVcs.AI2026-06中稿 · ECCV

提出视觉语义熵,更准确衡量模型对模糊图像的不确定性。

Visual Semantic Entropy: Do Vision Language Models Recognize Visual Ambiguity?

论文配图:Visual Semantic Entropy: Do Vision Language Models Recognize Visual Ambiguity?
图 1 · 摘自论文原文
  • 仅扰动图像,固定文本提示,避免提示敏感性干扰。
  • 通过聚类答案并计算语义原型分散度,量化视觉不确定性。
  • 在5个模型和5个数据集上表现最优,适合评估视觉模糊场景。

视觉语言模型在视觉模糊输入上仍会给出自信预测,导致偏差。现有基于熵的方法(如语义熵)依赖输出多样性,但过度自信的视觉特征会抑制随机解码下的输出多样性,造成不确定性低估。近期方法通过文本改写或图文联合扰动探测多样性,效果提升,但其变化常由文本差异主导,反映的是提示敏感性而非视觉模糊。为此,本文提出视觉语义熵(VSE),仅扰动图像以探测邻近视觉变化,保持文本查询不变。VSE通过将生成答案聚类为语义原型,并计算其质量加权的分散程度来衡量不确定性。在五个现代视觉语言模型和五个多样化VQA基准上的广泛评估表明,VSE能有效捕捉视觉模糊性,建立了视觉语言模型不确定性估计的新基准。

原文摘要 · Abstract (English)

Vision-language models can produce confident answers on visually ambiguous inputs, resulting in biased predictions. Common entropy-based methods, such as Semantic Entropy (SE), rely on output diversity. Yet our analysis shows that overconfident visual embeddings suppress output diversity under stochastic decoding, causing SE to underestimate uncertainty in such cases. Recent methods instead probe output diversity through input perturbations, including textual paraphrasing or joint text-image perturbations, and show improved performance. We study these approaches and reveals that the resulting variability is often dominated by textual changes rather than visual evidence, causing uncertainty estimates to reflect prompt sensitivity rather than visual ambiguity. We therefore propose Visual Semantic Entropy (VSE), which perturbs only the image to probe nearby visual variations while keeping the text query fixed. VSE measures uncertainty by clustering generated answers into semantic prototypes and computing the mass-weighted dispersion among them. Extensive evaluation across five modern vision-language models and five diverse VQA benchmarks demonstrates that VSE effectively captures visual ambiguity, establishing a new state-of-the-art for VLM uncertainty estimation.

视觉语言模型不确定性估计多模态视觉模糊

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。