arXiv:2502.02438cs.CRcs.AI2025-02AAAI被引 14

用普通图片加对抗噪声,黑盒盗取医学多模态模型。

Medical Multimodal Model Stealing Attacks via Adversarial Domain Alignment

  • 用自然图像+对抗噪声弥补与医疗数据分布差距
  • 在两个胸部X光数据集上成功复现模型功能
  • 无需医疗数据,适合研究模型安全的学者

医学多模态大语言模型(MLLM)正成为医疗系统的重要组成部分,辅助医疗人员进行决策和结果分析。放射科报告生成模型能解读医学影像,减轻放射科医生的工作负担。由于医学数据稀缺且受隐私法规保护,医学MLLM具有重要知识产权价值。然而,这些资产可能面临模型窃取攻击,即攻击者通过黑盒访问复制其功能。目前针对医疗领域的模型窃取主要集中在分类任务,对MLLM无效。本文提出首个针对医学MLLM的窃取攻击——对抗域对齐(ADA-STEAL)。该方法使用公开可得的自然图像,而非医疗图像。实验表明,通过对抗噪声的数据增强,足以克服自然图像与目标MLLM领域分布之间的差异。在IU X-RAY和MIMIC-CXR两个放射科数据集上的实验验证了该方法无需任何医疗数据即可成功窃取医学MLLM。

原文摘要 · Abstract (English)

Medical multimodal large language models (MLLMs) are becoming an instrumental part of healthcare systems, assisting medical personnel with decision making and results analysis. Models for radiology report generation are able to interpret medical imagery, thus reducing the workload of radiologists. As medical data is scarce and protected by privacy regulations, medical MLLMs represent valuable intellectual property. However, these assets are potentially vulnerable to model stealing, where attackers aim to replicate their functionality via black-box access. So far, model stealing for the medical domain has focused on classification; however, existing attacks are not effective against MLLMs. In this paper, we introduce Adversarial Domain Alignment (ADA-STEAL), the first stealing attack against medical MLLMs. ADA-STEAL relies on natural images, which are public and widely available, as opposed to their medical counterparts. We show that data augmentation with adversarial noise is sufficient to overcome the data distribution gap between natural images and the domain-specific distribution of the victim MLLM. Experiments on the IU X-RAY and MIMIC-CXR radiology datasets demonstrate that Adversarial Domain Alignment enables attackers to steal the medical MLLM without any access to medical data.

模型窃取医学AI对抗攻击多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。