医学决策中,纯文本比多模态更准,视觉信息反而拖后腿。
Why Text Prevails: Vision May Undermine Multimodal Medical Decision Making
- 用纯文本推理,比图文结合更可靠。
- 在阿尔茨海默病和X光片分类任务中,图文模型表现反不如纯文本。
- 适合医疗AI研究者关注多模态可靠性问题。
随着大语言模型的快速发展,先进的多模态大语言模型(MLLMs)在视觉-语言任务上展现出强大的零样本能力。然而,在生物医学领域,即使最先进的MLLMs也难以完成基础的医学决策(MDM)任务。我们通过两个挑战性数据集进行研究:(1)三阶段阿尔茨海默病(AD)分类(正常、轻度认知障碍、痴呆),类别间视觉差异微弱;(2)包含14种非互斥病状的MIMIC-CXR胸部X光片分类。实证结果表明,纯文本推理始终优于仅视觉或图文结合设置,且多模态输入常导致性能下降。为此,我们探索三种改进策略:(1)使用带推理标注示例的上下文学习;(2)先生成视觉描述再进行纯文本推理;(3)对视觉编码器进行少量样本微调并加入分类监督。这些发现揭示当前MLLM缺乏可靠的视觉理解能力,并指明了提升医疗领域多模态决策的可行方向。
原文摘要 · Abstract (English)
With the rapid progress of large language models (LLMs), advanced multimodal large language models (MLLMs) have demonstrated impressive zero-shot capabilities on vision-language tasks. In the biomedical domain, however, even state-of-the-art MLLMs struggle with basic Medical Decision Making (MDM) tasks. We investigate this limitation using two challenging datasets: (1) three-stage Alzheimer's disease (AD) classification (normal, mild cognitive impairment, dementia), where category differences are visually subtle, and (2) MIMIC-CXR chest radiograph classification with 14 non-mutually exclusive conditions. Our empirical study shows that text-only reasoning consistently outperforms vision-only or vision-text settings, with multimodal inputs often performing worse than text alone. To mitigate this, we explore three strategies: (1) in-context learning with reason-annotated exemplars, (2) vision captioning followed by text-only inference, and (3) few-shot fine-tuning of the vision tower with classification supervision. These findings reveal that current MLLMs lack grounded visual understanding and point to promising directions for improving multimodal decision making in healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。