arXiv:2509.13379cs.AIcs.CV2025-09Conference of the …被引 3

为视觉语言模型的不确定性评估建立全面基准,揭示大模型更懂自己不知道什么。

The Art of Saying "Maybe": A Conformal Lens for Uncertainty Benchmarking in VLMs

  • 用三种评分函数在6个数据集上评测18个先进VLM的不确定性表现。
  • 大模型不确定性判断更准,数学推理任务表现最差。
  • 提出指令引导的似然代理方法,适用于无日志概率的闭源模型。

视觉语言模型(VLMs)在科学与推理类复杂视觉理解任务中取得显著进展。尽管性能评估已推动对模型能力的理解,但不确定性量化这一关键维度仍缺乏关注。为此,不同于以往局限于特定场景的置信度预测研究,本文开展了一项全面的不确定性基准评估,覆盖18个前沿VLM(开源与闭源),在6个多模态数据集上使用3种不同的评分函数进行测试。针对无法访问词级别对数概率的闭源模型,我们提出并验证了基于指令的似然代理方法。结果表明:模型越大,不确定性判断越准确;知道得越多,越清楚自己不知道什么。更确信的模型往往准确率更高,而数学与推理类任务中的不确定性表现普遍低于其他领域。本工作为多模态系统中的可靠性不确定性评估奠定了基础。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have achieved remarkable progress in complex visual understanding across scientific and reasoning tasks. While performance benchmarking has advanced our understanding of these capabilities, the critical dimension of uncertainty quantification has received insufficient attention. Therefore, unlike prior conformal prediction studies that focused on limited settings, we conduct a comprehensive uncertainty benchmarking study, evaluating 18 state-of-the-art VLMs (open and closed-source) across 6 multimodal datasets with 3 distinct scoring functions. For closed-source models lacking token-level logprob access, we develop and validate instruction-guided likelihood proxies. Our findings demonstrate that larger models consistently exhibit better uncertainty quantification; models that know more also know better what they don't know. More certain models achieve higher accuracy, while mathematical and reasoning tasks elicit poorer uncertainty performance across all models compared to other domains. This work establishes a foundation for reliable uncertainty evaluation in multimodal systems.

视觉语言模型不确定性量化基准评测置信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。