arXiv:2602.09214cs.CV2026-02被引 5

评测视觉语言模型的模态特异性与跨模态不确定性,发现现有方法仍有明显短板。

VLM-UQBench: A Benchmark for Modality-Specific and Cross-Modality Uncertainties in Vision Language Models

  • 构建600个真实样本的基准,区分图像、文本及跨模态不确定性
  • 发现现有不确定性量化方法对特定模态敏感且易产生幻觉但信号弱
  • 适合关注VLM安全部署与可靠性评估的研究者

不确定性量化(UQ)对确保视觉语言模型(VLMs)安全可靠至关重要。核心挑战在于定位不确定性来源,判断其源于图像、文本,还是两者间的错配。我们提出VLM-UQBench,一个用于评估VLM中模态特异性与跨模态不确定性的基准。该基准包含从VizWiz数据集选取的600个真实世界样本,分为纯净、图像、文本和跨模态不确定性子集,并设计了可扩展的扰动流程:8种视觉扰动、5种文本扰动和3种跨模态扰动。我们还提出两个简单指标,量化UQ得分对扰动的敏感性及其与幻觉的相关性,并在此基础上评估四种VLM在三个数据集上的多种UQ方法。实证发现:(i) 现有UQ方法表现出显著的模态特异性,且高度依赖底层VLM;(ii) 模态特异性不确定性常与幻觉共现,但当前UQ分数提供的风险信号弱且不一致;(iii) 尽管在群体层面的明显模糊性上能媲美基于推理的链式思维基线,但在我们扰动流程引入的细微实例级模糊性检测上表现不佳。这些结果揭示了当前UQ实践与可靠VLM部署所需的细粒度、模态感知不确定性之间存在显著差距。

原文摘要 · Abstract (English)

Uncertainty quantification (UQ) is vital for ensuring that vision-language models (VLMs) behave safely and reliably. A central challenge is to localize uncertainty to its source, determining whether it arises from the image, the text, or misalignment between the two. We introduce VLM-UQBench, a benchmark for modality-specific and cross-modal data uncertainty in VLMs, It consists of 600 real-world samples drawn from the VizWiz dataset, curated into clean, image-, text-, and cross-modal uncertainty subsets, and a scalable perturbation pipeline with 8 visual, 5 textual, and 3 cross-modal perturbations. We further propose two simple metrics that quantify the sensitivity of UQ scores to these perturbations and their correlation with hallucinations, and use them to evaluate a range of UQ methods across four VLMs and three datasets. Empirically, we find that: (i) existing UQ methods exhibit strong modality-specific specialization and substantial dependence on the underlying VLM, (ii) modality-specific uncertainty frequently co-occurs with hallucinations while current UQ scores provide only weak and inconsistent risk signals, and (iii) although UQ methods can rival reasoning-based chain-of-thought baselines on overt, group-level ambiguity, they largely fail to detect the subtle, instance-level ambiguity introduced by our perturbation pipeline. These results highlight a significant gap between current UQ practices and the fine-grained, modality-aware uncertainty required for reliable VLM deployment.

不确定性量化视觉语言模型基准测试多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。