提出跨模态攻击方法CrossFire,有效欺骗多模态模型
Adversarial Attacks to Multi-Modal Models
- 将目标输入转为原始媒体格式,优化角度偏差生成扰动
- 在6个数据集上显著干扰下游任务,超越现有攻击
- 验证6种防御无效,揭示当前防护体系不足
多模态模型因其强大能力受到广泛关注,能有效对齐不同模态的嵌入表示,在下游任务中表现优于单模态模型。近期研究发现,攻击者可通过修改图像或音频文件,使其嵌入与目标输入匹配,从而欺骗下游模型。然而,由于不同模态间固有差异,该方法常表现不佳。本文提出CrossFire,一种创新的多模态模型攻击方法。CrossFire首先将攻击者选定的目标输入转换为与原始媒体(图像或音频)相同的模态格式,随后将其建模为优化问题,旨在最小化转换后输入与修改后媒体之间的嵌入角度偏差。求解该问题可确定需添加的扰动。在六个真实世界基准数据集上的大量实验表明,CrossFire能显著操纵下游任务,性能超过现有攻击方法。此外,我们评估了六种防御策略,发现现有防御无法有效应对CrossFire。
原文摘要 · Abstract (English)
Multi-modal models have gained significant attention due to their powerful capabilities. These models effectively align embeddings across diverse data modalities, showcasing superior performance in downstream tasks compared to their unimodal counterparts. Recent study showed that the attacker can manipulate an image or audio file by altering it in such a way that its embedding matches that of an attacker-chosen targeted input, thereby deceiving downstream models. However, this method often underperforms due to inherent disparities in data from different modalities. In this paper, we introduce CrossFire, an innovative approach to attack multi-modal models. CrossFire begins by transforming the targeted input chosen by the attacker into a format that matches the modality of the original image or audio file. We then formulate our attack as an optimization problem, aiming to minimize the angular deviation between the embeddings of the transformed input and the modified image or audio file. Solving this problem determines the perturbations to be added to the original media. Our extensive experiments on six real-world benchmark datasets reveal that CrossFire can significantly manipulate downstream tasks, surpassing existing attacks. Additionally, we evaluate six defensive strategies against CrossFire, finding that current defenses are insufficient to counteract our CrossFire.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。