提出医疗视觉语言模型校准评估基准,提升临床可靠性。
MVC-Bench: Benchmarking Calibration of Medical Vision-Language Models

- 构建多维度校准评测基准,覆盖多种模态与模型架构
- 发现模型校准误差主因,提出训练时正则化方法降低误差
- 适合医疗AI安全评估与模型优化研究者参考
可靠的视觉语言模型(VLMs)及医学视觉语言模型(Medical-VLMs)评估需具备校准的置信度,尤其在真实临床场景中。然而现有工作主要关注准确率提升,对医学领域的校准问题研究不足。为此,我们提出MVC-Bench,一个以校准为核心的医学图像分类评测基准。该基准从三个维度评估校准性能:(i)对模态、骨干网络和领域偏移的鲁棒性;(ii)校准策略与提示微调方法的有效性;(iii)提示模板与随机种子变化下的稳定性。基准涵盖八种骨干网络、三种医学模态(眼底成像、病理切片、胸部X光),在域内与域外设置下进行测试。对比后处理校准、训练时校准与零样本推理方法,以及六种提示微调方法。在超过1638次受控实验中,以准确率和期望校准误差(ECE)为主要指标,辅以最大校准误差(MCE)与自适应校准误差(ACE)。进一步分析了VLMs与Medical-VLMs校准偏差的根源,并提出一种简单训练时校准方法——多类边界(MCM)正则化,在12个设置中10个达到最低ECE,域外场景下仍具竞争力。整体提供结构化评估框架与医疗高风险应用中的可操作指导。
原文摘要 · Abstract (English)
Reliable evaluation of vision-language models (VLMs) and medical vision-language models (Medical-VLMs) requires calibrated confidence, particularly under realistic clinical conditions. However, existing efforts mainly focused on improving accuracy, leaving calibration in the medical domain underexplored. To this end, we propose MVC-Bench, a calibration-centric benchmark for medical image classification with VLMs and Medical-VLMs. MVC-Bench assesses the calibration across three axes: (i) robustness to modality, backbone, and domain shift (ii) effectiveness of calibration strategies and prompt-tuning methods (iii) stability under prompt-template and random-seed variations. The benchmark covers eight different backbones, three medical modalities, including fundus imaging, histopathology, and chest X-ray under in-domain and domain shift settings. It compares post-hoc calibration, train-time calibration, and zero-shot inference methods, together with six prompt-tuning methods. Across more than 1638 controlled experiments, we report accuracy and Expected Calibration Error (ECE) as primary metrics, and further report results with complementary calibration measures, including Maximum Calibration Error (MCE) and Adaptive Calibration Error (ACE). We further investigate the underlying causes of miscalibration in VLMs and Medical-VLMs and propose a simple train-time calibration method, Multi-Class Margin (MCM) regularization, which achieves lowest ECE on 10 out of 12 settings in in-domain and remains competitive under domain shifts. Collectively, MVC-Bench provides a structured evaluation framework and actionable guidance for improving calibration in safety-critical medical workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。