提升多模态模型对抗攻击的泛化性与隐蔽性,让攻击更难被发现且适应更多目标。
Improving Generalizability and Undetectability for Targeted Adversarial Attacks on Multimodal Pre-trained Models
- 用多源多目标代理优化对抗样本,增强跨模态泛化能力。
- 在多个目标上成功率超90%,且能避开多种异常检测方法。
- 适用于研究模型安全性的研究人员,尤其关注对抗攻击设计者。
多模态预训练模型(如ImageBind)通过将不同数据模态对齐到共享嵌入空间,在下游任务中表现卓越。然而其广泛应用也引发严重安全问题,尤其是针对特定目标的对抗攻击。本文指出,现有攻击在泛化性和隐蔽性方面仍存不足:生成的对抗样本在部分已知或语义相似目标上的迁移能力弱(泛化性差),且易被简单异常检测方法识别(隐蔽性差)。为此,我们提出新方法Proxy Targeted Attack(PTA),利用多源模态和目标模态的代理来优化对抗样本,使其在保持规避防御的同时,能对多个潜在目标实现对齐。我们还提供了理论分析,揭示泛化性与隐蔽性间的权衡关系,并确保在满足隐蔽性要求的前提下实现最优泛化。实验表明,PTA在多个目标上均达到高成功率(>90%),且对多种异常检测方法具有强鲁棒性。
原文摘要 · Abstract (English)
Multimodal pre-trained models (e.g., ImageBind), which align distinct data modalities into a shared embedding space, have shown remarkable success across downstream tasks. However, their increasing adoption raises serious security concerns, especially regarding targeted adversarial attacks. In this paper, we show that existing targeted adversarial attacks on multimodal pre-trained models still have limitations in two aspects: generalizability and undetectability. Specifically, the crafted targeted adversarial examples (AEs) exhibit limited generalization to partially known or semantically similar targets in cross-modal alignment tasks (i.e., limited generalizability) and can be easily detected by simple anomaly detection methods (i.e., limited undetectability). To address these limitations, we propose a novel method called Proxy Targeted Attack (PTA), which leverages multiple source-modal and target-modal proxies to optimize targeted AEs, ensuring they remain evasive to defenses while aligning with multiple potential targets. We also provide theoretical analyses to highlight the relationship between generalizability and undetectability and to ensure optimal generalizability while meeting the specified requirements for undetectability. Furthermore, experimental results demonstrate that our PTA can achieve a high success rate across various related targets and remain undetectable against multiple anomaly detection methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。