诊断医学多模态模型对文本的依赖偏见,揭示其忽视影像的关键问题。
On the Risk of Misleading Reports: Diagnosing Textual Biases in Multimodal Clinical AI
- 通过替换图像或文本,检测模型对模态的依赖程度。
- 六款模型在两张数据集上均严重依赖文本,准确率下降超30%。
- 适合关注医疗AI公平性与多模态融合的研究者阅读。
临床决策依赖医学影像与关联报告的综合分析。尽管视觉语言模型(VLMs)可统一处理此类任务,但常表现出对某一模态的强烈偏见,频繁忽略关键视觉线索而偏向文本信息。本文提出选择性模态切换(SMS),一种基于扰动的方法,用于量化模型在二分类任务中对各模态的依赖程度。通过系统地在标签相反的样本间交换图像或文本,暴露模态特异性偏见。我们在两个不同模态的医学影像数据集MIMIC-CXR(胸部X光)和FairVLMed(扫描激光眼底成像)上评估六款开源VLMs——四款通用模型及两款针对医学数据微调的模型。通过对比模型在正常与扰动设置下的性能表现及校准情况,发现文本依赖显著,即使存在互补视觉信息也未改善。进一步的注意力可视化分析表明,图像内容常被文本细节掩盖。研究强调需设计并评估真正融合视觉与文本线索的多模态医疗模型,而非仅依赖单一模态信号。
原文摘要 · Abstract (English)
Clinical decision-making relies on the integrated analysis of medical images and the associated clinical reports. While Vision-Language Models (VLMs) can offer a unified framework for such tasks, they can exhibit strong biases toward one modality, frequently overlooking critical visual cues in favor of textual information. In this work, we introduce Selective Modality Shifting (SMS), a perturbation-based approach to quantify a model's reliance on each modality in binary classification tasks. By systematically swapping images or text between samples with opposing labels, we expose modality-specific biases. We assess six open-source VLMs-four generalist models and two fine-tuned for medical data-on two medical imaging datasets with distinct modalities: MIMIC-CXR (chest X-ray) and FairVLMed (scanning laser ophthalmoscopy). By assessing model performance and the calibration of every model in both unperturbed and perturbed settings, we reveal a marked dependency on text input, which persists despite the presence of complementary visual information. We also perform a qualitative attention-based analysis which further confirms that image content is often overshadowed by text details. Our findings highlight the importance of designing and evaluating multimodal medical models that genuinely integrate visual and textual cues, rather than relying on single-modality signals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。