通过隐蔽扰动背景诱导医疗视觉语言模型误诊,且攻击可跨模型迁移。
When Background Matters: Breaking Medical Vision Language Models by Transferable Attack

- 在非病灶背景区域注入协同扰动,用注意力干扰让模型忽略病变区。
- 六种医学影像模态下攻击成功率超现有方法,诊断结果仍具临床合理性。
- 提出新评估框架,揭示当前医疗多模态模型推理能力存在严重缺陷。
视觉语言模型(VLMs)在临床诊断中应用日益广泛,但其对对抗攻击的鲁棒性尚未被充分探索,存在重大风险。现有医疗攻击多聚焦于模型窃取或对抗微调等次要目标,而来自自然图像的可迁移攻击会产生明显失真,易被临床医生察觉。为此,我们提出 MedFocusLeak,一种高度可迁移的黑盒多模态攻击方法,在保持扰动不可感知的同时,诱导出错误却临床合理的诊断。该方法在非诊断性背景区域注入协调扰动,并采用注意力干扰机制,引导模型关注非病灶区域。在六种医学成像模态上的广泛评估显示,MedFocusLeak 达到当前最佳性能,能生成误导性但逼真的诊断输出,适用于多种临床 VLM。我们进一步构建统一评估框架,引入新指标联合衡量攻击成功率与图像保真度,揭示现代临床 VLM 在推理能力上的关键弱点。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) are increasingly used in clinical diagnostics, yet their robustness to adversarial attacks remains largely unexplored, posing serious risks. Existing medical attacks focus on secondary objectives such as model stealing or adversarial fine-tuning, while transferable attacks from natural images introduce visible distortions that clinicians can easily detect. To address this, we propose MedFocusLeak, a highly transferable black-box multimodal attack that induces incorrect yet clinically plausible diagnoses while keeping perturbations imperceptible. The method injects coordinated perturbations into non-diagnostic background regions and employs an attention distraction mechanism to shift the model's focus away from pathological areas. Extensive evaluations across six medical imaging modalities show that MedFocusLeak achieves state-of-the-art performance, generating misleading yet realistic diagnostic outputs across diverse VLMs. We further introduce a unified evaluation framework with novel metrics that jointly capture attack success and image fidelity, revealing a critical weakness in the reasoning capabilities of modern clinical VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。