arXiv:2609.03261cs.CVcs.CL2026-09

发现医学多模态答题中模型依赖文字线索而非图像,误导了评估结果。

MedQA-MM: Shortcuts Behind Medical Visual Reasoning

论文配图:MedQA-MM: Shortcuts Behind Medical Visual Reasoning
图 1 · 摘自论文原文
  • 通过审计和消融实验分离图像与文字线索,识别答题路径
  • 全输入准确率62.63%,仅用文字或选项时分别降至53.96%和29.71%
  • 构建1000题去捷径数据集,文本/选项单独作答准确率降至5.21%和12.33%

评估仅关注最终答案正确性,忽视作答路径。在医学多模态选择题中,正确答案可能源于图像真实发现,也可能依赖答案表述中的保留线索——如非视觉临床文本、可见图像文字、人工标注或设备/上下文伪影。这种现象导致评分层面的推理过拟合。我们通过提示与图像侧审计、模态消融及匹配修复,在六大数据集上分离候选线索与行为证据。13种配置的开放模型测试中,全输入准确率为62.63%,纯文本与仅选项设置分别为53.96%和29.71%。移除长度差、绝对/显著、空间/介词线索后,准确率分别下降6.58、3.50和4.77个百分点。我们构建了包含1000题的去捷径数据集MedQA-MM,其中纯文本与仅选项准确率降至5.21%和12.33%。这不意味着模型不用图像,而是强调医学图像推理需基于路径级证据。

原文摘要 · Abstract (English)

A benchmark score credits final answers, but not the route by which an item can be answered. In medical multimodal multiple-choice questions (MCQs), this distinction matters because a correct answer can be supported by the intended image finding or by benchmark-preserved cues in the wording of answers, non-visual clinical text, visible image text, artificial annotations, or device/context artifacts. We call the resulting score-level overinterpretation reasoning inflation. Here, a route is an observable input path that can support answer selection, not a claim about the model's hidden cognition. Across six medical multimodal MCQ datasets, we separate candidate cues from behavioral evidence through prompt- and image-side audits, modality ablations, and matched repairs that preserve the medical target and answer key. In a 13-configuration open-model panel, full-input accuracy is 62.63%, while text-only and options-only settings achieve 53.96% and 29.71%, respectively. Removing length-gap, absolute/conspicuous, and spatial/prepositional cues lowers accuracy by 6.58, 3.50, and 4.77 percentage points. We also construct MedQA-MM, a 1,000-item shortcut-mitigated subset, where text-only and options-only accuracy fall to 5.21% and 12.33%. This does not imply that models never use images; it shows that medical image-reasoning claims require route-level evidence.

医学视觉推理多模态评估捷径检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。