分析视觉语言模型的口语化置信度,发现普遍不准且需改进。
Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models

- 用自然语言表达信心,评估视觉语言模型的可靠性
- 多数模型在多任务中严重校准偏差,视觉推理模型更准
- 提出双阶段提示法提升多模态信心对齐,适合可信AI研究者
不确定性量化对于评估现代AI系统的可靠性和可信度至关重要。现有方法中,通过自然语言表达置信度的口语化不确定性,在大语言模型中已展现出轻量且可解释的优势。然而其在视觉语言模型(VLMs)中的有效性尚未充分研究。本文对三类VLM、四个任务领域和三种评估场景下的口语化置信度进行了全面评估。结果表明,当前VLM在多样任务与设置下普遍存在显著的校准偏差。值得注意的是,视觉推理模型(即‘以图像思考’)始终表现出更好的校准能力,表明模态特定推理对可靠不确定性估计至关重要。为进一步应对校准挑战,我们提出视觉置信感知提示(Visual Confidence-Aware Prompting),一种两阶段提示策略,可提升多模态环境中的信心对齐。总体而言,本研究揭示了跨模态下VLM固有的校准问题,强调了模态对齐与模型忠实性在构建可靠多模态系统中的根本重要性。
原文摘要 · Abstract (English)
Uncertainty quantification is essential for assessing the reliability and trustworthiness of modern AI systems. Among existing approaches, verbalized uncertainty, where models express their confidence through natural language, has emerged as a lightweight and interpretable solution in large language models (LLMs). However, its effectiveness in vision-language models (VLMs) remains insufficiently studied. In this work, we conduct a comprehensive evaluation of verbalized confidence in VLMs, spanning three model categories, four task domains, and three evaluation scenarios. Our results show that current VLMs often display notable miscalibration across diverse tasks and settings. Notably, visual reasoning models (i.e., thinking with images) consistently exhibit better calibration, suggesting that modality-specific reasoning is critical for reliable uncertainty estimation. To further address calibration challenges, we introduce Visual Confidence-Aware Prompting, a two-stage prompting strategy that improves confidence alignment in multimodal settings. Overall, our study highlights the inherent miscalibration in VLMs across modalities. More broadly, our findings underscore the fundamental importance of modality alignment and model faithfulness in advancing reliable multimodal systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。