提出跨模态攻击框架Medusa,可黑盒欺骗医疗多模态生成系统
Medusa: Cross-Modal Transferable Adversarial Attacks on Multimodal Medical Retrieval-Augmented Generation
- 用多正样本损失对齐视觉与恶意文本嵌入,劫持检索过程
- 在报告生成和疾病诊断任务中攻击成功率超90%,且抗多种防御
- 适用于评估医疗AI系统安全,尤其关注临床决策支持场景
随着检索增强型视觉-语言模型的快速发展,多模态医疗检索增强生成(MMed-RAG)系统在临床决策支持中日益普及。这类系统通过跨模态检索整合视觉与文本证据,用于报告生成和疾病诊断等任务。然而其复杂架构也带来了未充分研究的对抗性漏洞,尤其是针对视觉输入扰动的威胁。本文提出Medusa,一种在黑盒设置下针对MMed-RAG系统的新型跨模态可迁移对抗攻击框架。Medusa将攻击建模为扰动优化问题,利用多正样本InfoNCE损失(MPIL)使对抗性视觉嵌入与医学上合理但恶意的文本目标对齐,从而劫持检索流程。为提升迁移能力,采用代理模型集成,并设计结合不变风险最小化(IRM)的双循环优化策略。在两项真实世界医疗任务(医学报告生成与疾病诊断)上的大量实验表明,Medusa在适当参数配置下,对多种生成模型与检索器平均攻击成功率超过90%,且对四种主流防御手段保持鲁棒性,显著优于现有基准。结果揭示了MMed-RAG系统的重大安全隐患,凸显在高风险医疗应用中进行鲁棒性评测的必要性。代码与数据见:https://anonymous.4open.science/r/MMed-RAG-Attack-F05A。
原文摘要 · Abstract (English)
With the rapid advancement of retrieval-augmented vision-language models, multimodal medical retrieval-augmented generation (MMed-RAG) systems are increasingly adopted in clinical decision support. These systems enhance medical applications by performing cross-modal retrieval to integrate relevant visual and textual evidence for tasks, e.g., report generation and disease diagnosis. However, their complex architecture also introduces underexplored adversarial vulnerabilities, particularly via visual input perturbations. In this paper, we propose Medusa, a novel framework for crafting cross-modal transferable adversarial attacks on MMed-RAG systems under a black-box setting. Specifically, Medusa formulates the attack as a perturbation optimization problem, leveraging a multi-positive InfoNCE loss (MPIL) to align adversarial visual embeddings with medically plausible but malicious textual targets, thereby hijacking the retrieval process. To enhance transferability, we adopt a surrogate model ensemble and design a dual-loop optimization strategy augmented with invariant risk minimization (IRM). Extensive experiments on two real-world medical tasks, including medical report generation and disease diagnosis, demonstrate that Medusa achieves over 90% average attack success rate across various generation models and retrievers under appropriate parameter configuration, while remaining robust against four mainstream defenses, outperforming state-of-the-art baselines. Our results reveal critical vulnerabilities in the MMed-RAG systems and highlight the necessity of robustness benchmarking in safety-critical medical applications. The code and data are available at https://anonymous.4open.science/r/MMed-RAG-Attack-F05A.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。