arXiv:2608.02790cs.CV2026-08

医学影像大模型看似准确,实则自信过度,自知能力差。

Confident but Unreliable: A Behavioral Safety Audit of Vision-Language Models on Brain MRI

  • 用4102张脑部MRI测试6个视觉语言模型,评估其自信度与正确性分离现象
  • 错误答案的平均置信度高达0.82-0.97,近半数高自信回答是错的
  • 建议医疗领域评估应同时报告置信度、幻觉和拒答情况

视觉语言模型(VLMs)在医学影像中的应用日益增多,但其宣称的置信度很少与正确性分开评估。本文以脑部MRI为高风险、可控测试场景,开展前沿多模态系统的系统性行为审计,揭示一个普遍存在的失效模式:模型可能表现得十分胜任,却缺乏可靠的自我认知能力。我们对六种指令微调的VLMs(五种通用型,一种医学专用)进行了自动评分的行为审计,涵盖4,102张图像(来自250名受试者的4,032张轴位/冠状位/矢状位MRI切片,以及70张非脑部/噪声控制图像),标签基于公开元数据和已发布的专家分割掩码,而非新的人工标注。各模型回答覆盖率接近100%,但置信度校准效果差:ECE值在0.27至0.40之间,错误答案的平均置信度为0.82至0.97,33%-46%的答题属于高自信错误。最准确的模型在错误上也最为自信;基线与专业模型对比显示,医学适配提升了肿瘤检测能力,但未改善置信度可靠性。开放性诊断分析进一步表明,幻觉与回避行为独立于多项选择准确性。研究呼吁,在医疗图像的VLM评估中,应同步报告置信度可靠性、自信错误、幻觉和拒绝回答等指标。

原文摘要 · Abstract (English)

Vision-language models (VLMs), including medical specialists, are increasingly proposed for medical imaging, yet their stated confidence is rarely evaluated separately from correctness. We use brain MRI as a controlled, high-stakes testbed for a broader failure mode in frontier multimodal systems: models can appear competent while lacking reliable self-knowledge. We present an automatically graded behavioral audit and pilot study of six instruction-tuned VLMs (five general-purpose and one medical specialist) on 4,102 images (4,032 axial/coronal/sagittal MRI slices from 250 subjects plus 70 non-brain/noise controls), with labels derived from public metadata and released expert segmentation masks rather than new human annotation. Across models, answer coverage is near-complete, but verbalized-confidence calibration is poor: ECE ranges from 0.27 to 0.40, mean confidence on incorrect answers ranges from 0.82 to 0.97, and 33-46% of answered items are high-confidence errors. The most accurate model is also the most confident on its errors, while a base/specialist family contrast suggests that medical adaptation improves tumor-presence detection without improving confidence reliability. Open-ended diagnostics further show that hallucination and abstention vary separately from multiple-choice accuracy. These findings argue that medical-image VLM evaluation should report verbalized-confidence reliability, confident error, hallucination, and abstention alongside accuracy.

视觉语言模型医学影像置信度评估大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。