发现医学图像检测中文本误导模型判断,导致真假误判。
Text Over Image: Auditing Multimodal Robustness in Synthetic Medical Image Detection

- 固定图像,更换文本记录,测试多模态模型判断变化
- 相同图像因文本不同,真伪判断翻转,真实图像识别率下降61.1%
- 提出无需重训练的推理阶段修复方案,效果优于直接提示抑制
随着生成式AI的广泛应用,合成医学图像带来诊断欺骗和保险欺诈等风险。尽管已有研究使用视觉语言模型(VLM)检测合成图像,但大多仅考虑图像单独输入。临床实践中,图像常与结构化记录和元数据一同分析,且VLM越来越多地在图像-记录联合输入下部署。我们发现一种此前未被充分关注的多模态脆弱性:当同时提供图像和文本时,VLM可能过度依赖文本上下文进行真实性判断,导致同一图像因伴随文本不同而获得不同预测结果。这引发了真实场景部署中的鲁棒性担忧。为系统刻画该现象,我们将合成医学图像检测重构为对图像-记录接口处多模态鲁棒性的审计,并引入一个配对基准,保持图像不变而交换受控的元数据变体。在多种成像模态下,评估了多种开源权重与前沿API VLM,发现仅改变元数据上下文即可引发真伪判断翻转,尤其在显式标注为AI生成时,真实图像识别准确率平均下降61.1%。我们进一步提出一种推理时缓解管道,可检测并消除来源捷径,无需模型重训练,在受影响子集上显著优于直接提示抑制。该基准为超越图像单一设置的多模态鲁棒性评估与改进提供了标准化工具。代码与数据将在录用后公开。
原文摘要 · Abstract (English)
With the rapid adoption of generative AI, synthetic medical images pose growing risks, including diagnostic deception and insurance fraud. Although prior work has explored vision-language model (VLM)-based synthetic image detection, these evaluations typically consider images in isolation. In clinical practice, however, images are interpreted alongside structured records and metadata, and VLMs are increasingly deployed under joint image-record inputs. We uncover a previously underexamined multimodal vulnerability: when given both modalities, VLMs may overweight record context in authenticity judgments, such that the same image receives different predictions solely due to changes in its accompanying text. This raises concerns about robustness in real-world deployment. To systematically characterize this effect, we reformulate synthetic medical image detection as an audit of multimodal robustness at the image-record interface and introduce a paired benchmark that holds the image fixed while swapping controlled metadata variants. Across multiple imaging modalities, we evaluate diverse open-weight and frontier API VLMs and find that changing the metadata context alone can flip authenticity judgments, with accuracy on authentic images dropping by 61.1% on average under an explicit AI-origin tag. We further propose an inference-time mitigation pipeline that detects and neutralizes provenance shortcuts without model retraining, substantially outperforming direct prompt-based suppression on the affected subset. Our benchmark provides a standardized tool for assessing and improving multimodal robustness beyond image-only settings. Code and data will be released upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。