医学视觉问答模型常过度自信,新方法结合幻觉检测提升可信度。
Overconfidence and Calibration in Medical VQA: Empirical Findings and Hallucination-Aware Mitigation
- 通过实证研究发现模型即使扩大规模或改用思维链提示仍过度自信。
- 后处理校准可显著降低误差,但对区分能力改善有限。
- 引入幻觉检测信号能同时改进校准效果和预测区分度,尤其在开放问题上。
随着视觉语言模型(VLMs)在临床决策支持中的广泛应用,不仅需要高准确率,更需知道何时信任其预测。然而,医学领域对模型置信度校准的系统性研究仍很缺乏。本研究在三个VLM家族(Qwen3-VL、InternVL3、LLaVA-NeXT)、三种模型规模(2B–38B)及多个医疗视觉问答(VQA)基准上,系统评估了不同置信度估计提示策略的校准表现。发现:1)过度自信现象普遍存在,不受模型规模或提示方式(如思维链、显式置信度)影响;2)简单后处理校准(如Platt缩放)可有效降低校准误差,优于提示法;3)由于严格单调性,后处理方法无法提升预测区分能力,AUROC保持不变。为此,我们提出幻觉感知校准(HAC),利用视觉锚定的幻觉检测信号作为补充输入以优化置信度。结果显示,该方法在开放问题上显著提升校准效果与AUROC。研究建议将后处理校准作为医学VLM部署的标准做法,并证明幻觉信号对提升模型可靠性具有实际价值。
原文摘要 · Abstract (English)
As vision-language models (VLMs) are increasingly deployed in clinical decision support, more than accuracy is required: knowing when to trust their predictions is equally critical. Yet, a comprehensive and systematic investigation into the overconfidence of these models remains notably scarce in the medical domain. We address this gap through a comprehensive empirical study of confidence calibration in VLMs, spanning three model families (Qwen3-VL, InternVL3, LLaVA-NeXT), three model scales (2B--38B), and multiple confidence estimation prompting strategies, across three medical visual question answering (VQA) benchmarks. Our study yields three key findings: First, overconfidence persists across model families and is not resolved by scaling or prompting, such as chain-of-thought and verbalized confidence variants. Second, simple post-hoc calibration approaches, such as Platt scaling, reduce calibration error and consistently outperform the prompt-based strategy. Third, due to their (strict) monotonicity, these post-hoc calibration methods are inherently limited in improving the discriminative quality of predictions, leaving AUROC at the same level. Motivated by these findings, we investigate hallucination-aware calibration (HAC), which incorporates vision-grounded hallucination detection signals as complementary inputs to refine confidence estimates. We find that leveraging these hallucination signals improves both calibration and AUROC, with the largest gains on open-ended questions. Overall, our findings suggest post-hoc calibration as standard practice for medical VLM deployment over raw confidence estimates, and highlight the practical usefulness of hallucination signals to enable more reliable use of VLMs in medical VQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。