医学影像模型可能靠病名先验而非图像判断,真实能力需用因果审计验证。
Vision-language models for chest radiography do not always need the image

- 通过遮蔽、替换图像等干预手段,测试模型是否真正依赖图像信息。
- 纯文本模型最高仅比最优多模态模型低5.7分,大模型表现与小文本模型无显著差异。
- 多数模型仅对部分病灶使用图像,且报告置信度可反映其是否真看图。
医学视觉-语言模型在胸部X光片上表现优异,常被解读为它们真正理解了图像。然而这种推断不可靠:仅依赖病灶名称先验的模型也能达到类似准确率,且现有标准评估无法区分二者。本文提出一种因果审计方法,通过遮蔽相关/无关区域、替换同标签扫描等方式干预图像输入,并结合三种行为指标判断正确答案是否依赖图像。在九个系统中,完全不访问图像的文本模型准确率仅比最佳多模态模型低5.7个百分点;一个1190亿参数的多模态模型与70亿参数的文本基线在统计上无差异。审计将模型分为三类:三类忽略图像,一类不稳定,五类仅选择性使用图像处理特定发现。该分类在第二数据集、不同分辨率和提示语句下依然成立。与持证放射科医生对比,纯文本模型准确率与放射科医生相当,但接地率为零;而图像使用型模型的接地率可达医生水平。报告置信度仅在模型真正使用图像时才有效标记非接地回答。临床部署应以接地审计为准,而非仅看准确率。
原文摘要 · Abstract (English)
Medical vision-language models report strong chest radiograph accuracy, and this is increasingly read as evidence that they use the image. That inference is unsafe: a model exploiting finding-name priors scores like one that reads the scan, and no standard benchmark separates them. We introduce a causal audit that intervenes on the image, occluding the relevant region, occluding an irrelevant one, and swapping in another patient's same-label scan, and combines three behavioral metrics to test whether a correct answer depends on the image. Across nine systems, a text-only model with no image access reaches within 5.7 accuracy points of the best multimodal one, and a 119-billion-parameter multimodal model is statistically indistinguishable from a 7-billion text-only baseline. The audit splits the cohort into three models that ignore the image, one that is unstable, and five that use it selectively, for a subset of findings; the categories hold across a second dataset, resolution, and prompt phrasing. Against board-certified radiologists, a text-only model is statistically indistinguishable from a radiologist's accuracy while grounding at zero, whereas the image-using models ground at radiologist-comparable rates. Reported confidence flags ungrounded answers only when a model uses the image. Grounding audits, not accuracy, should gate clinical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。