arXiv:2605.17436cs.CVcs.CL2026-05

医学文本会扭曲视觉语言模型的诊断判断,影响临床可靠性。

Medical Context Distorts Decisions in Clinical Vision Language Models

论文配图:Medical Context Distorts Decisions in Clinical Vision Language Models
图 1 · 摘自论文原文
  • 通过控制图像-文本对齐、病史和提示,发现模型过度依赖文本。
  • 即使有明确影像证据,模型仍受无关病史干扰,正确率下降27%。
  • 微小提示变化可导致正确预测反转,适合医疗AI安全评估者关注。

视觉语言模型(VLMs)在临床决策支持中日益受到关注,但其在需整合医学记录中视觉与文本上下文的真实场景下的可靠性尚未充分验证。本文识别出三种失效模式:(1) 文本模态过度依赖,(2) 对无关临床病史的虚假依赖,(3) 对语义等价输入的提示敏感性。我们在MIMIC-CXR数据集上评估了多种通用领域及医学调优的开源与闭源VLM,在胸部X光任务中系统性地操纵图像-文本对齐、临床病史和提示形式。结果表明,模型决策主要受文本主导,即使存在可用的视觉证据。此外,模型严重受无关报告影响,而轻微提示变化即可逆转基于图像的正确预测。研究强调必须在临床应用前引入显式防护机制与压力测试。

原文摘要 · Abstract (English)

Vision-language models (VLMs) are increasingly proposed for clinical decision support, yet their reliability in real-world scenarios that require integrating both visual and textual context from medical records remains poorly characterized. This paper identifies three failure modes: (1) modality over-reliance on text over images, (2) spurious reliance on irrelevant clinical history, and (3) prompt sensitivity across semantically equivalent inputs. We evaluate a diverse set of general-domain and medically-tuned open and closed VLMs on chest x-ray tasks using MIMIC-CXR. By systematically manipulating image-text alignment, clinical history, and prompt formulations, we found that VLM decisions are dominated by the text modality, even when visual evidence is available. Moreover, we observed that VLMs are heavily influenced by irrelevant reports, while minor prompt changes can reverse correct image-based predictions. Our findings underscore the need for explicit safeguards and stress-testing before considering the use of these models in clinical practice.

视觉语言模型临床决策医疗AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。