测试视觉语言模型在医学影像中受误导性上下文干扰时的可靠性。
MC-CXR: A Multi-Context Chest X-ray Benchmark for Context-Induced Disruption in Vision-Language Models

- 构建多上下文胸部X光数据集,模拟文本与影像报告冲突场景。
- 模型在误导性文本下错误率高达78.1%,视觉干扰也导致35.7%~61.7%误判。
- 发现文本误导比视觉误导更易引发错误,适合医疗AI可靠性研究者使用。
视觉语言模型(VLMs)在临床流程中常需结合检索到的报告、初步笔记或既往影像解读胸部X光片。现有基准仅评估模型孤立回答的准确性,未衡量当合理上下文与图像矛盾时,模型是否仍能保持正确的图像决策。我们提出多上下文胸部X光(MC-CXR)基准,包含240个病例扩展为2,522个实例,通过成对扰动隔离上下文诱导的干扰。每个病例固定当前图像和目标病灶,同时呈现匹配的可靠与误导性文本及既往胸片上下文,必要时辅以视觉叠加。MC-CXR定义了三类任务与两种配对指标:错转率与上下文一致误差率。我们评估了十种VLMs,涵盖开源通用、医疗专用及闭源系统。图像独立准确率虽必要但不足。在误导性文本下,平均错转率范围为45.6%-78.1%;在误导性视觉上下文下为35.7%-61.7%。错转预测中,74.6%与文本误导标签一致,而视觉上下文仅为17.6%,差距达57.0点(95%置信区间50.9-62.8)。该文-视不对称现象在标准化直接回答协议下依然显著。数据集已发布于PhysioNet。
原文摘要 · Abstract (English)
Vision-language models (VLMs) are increasingly used in clinical pipelines where a chest X-ray is interpreted alongside retrieved reports, preliminary notes, or prior imaging. Existing benchmarks measure whether models answer correctly in isolation, but not whether they preserve a correct image-only decision when plausible context conflicts with the image. We introduce Multi-Context Chest X-ray (MC-CXR), a benchmark of 240 cases expanded into 2,522 instances that isolates context-induced disruption through paired perturbation. Each case fixes the current image and target finding while presenting matched reliable and misleading context across text and prior CXR, with visual overlays where available. MC-CXR defines three task families and two paired metrics, the switch-to-wrong rate and the context-aligned error rate. We evaluate ten VLMs spanning open-source general, medical-domain, and closed-source systems. Image-only accuracy is necessary but insufficient. Mean switch rates range from 45.6-78.1% across misleading textual sources and 35.7-61.7% across misleading visual sources. Among switched predictions, 74.6% align with the misleading label for text versus 17.6% for visual context, a 57.0-point gap (95% CI 50.9-62.8). This text-visual asymmetry is observed under the standardized direct-answer protocol. The dataset is available on PhysioNet.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。