提升医疗多模态模型置信度校准,减少误诊风险
Confidence Calibration for Multimodal LLMs: An Empirical Study through Medical VQA

- 用多策略提问+专家模型评估融合,改进置信度判断
- 在三个医疗VQA数据集上平均降低40%预期校准误差
- 适合医疗AI诊断系统可靠性提升的从业者参考
多模态大语言模型在医疗任务中潜力巨大,但其输出置信度常与实际准确率不匹配,可能引发误诊或忽略正确建议。本研究首次系统分析了医疗多模态大模型中准确率与置信度的关系,提出一种结合多策略融合式提问(MS-FBI)与辅助专家大模型评估的新方法,旨在提升医学视觉问答(Medical VQA)中的置信度校准效果。实验表明,该方法在三个医疗VQA数据集上平均将预期校准误差(ECE)降低40%,显著增强模型可靠性。研究强调了医疗领域专用校准的重要性,为AI辅助诊断提供了更可信的解决方案。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) show great potential in medical tasks, but their elicited confidence often misaligns with actual accuracy, potentially leading to misdiagnosis or overlooking correct advice. This study presents the first comprehensive analysis of the relationship between accuracy and confidence in medical MLLMs. It proposes a novel method that combines Multi-Strategy Fusion-Based Interrogation (MS-FBI) with auxiliary expert LLM assessment, aiming to improve confidence calibration in Medical Visual Question Answering (VQA). Experiments demonstrate that our method reduces the Expected Calibration Error (ECE) by an average of 40\% across three Medical VQA datasets, significantly enhancing MLLMs' reliability. The findings highlight the importance of domain-specific calibration for MLLMs in healthcare, offering a more trustworthy solution for AI-assisted diagnosis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。