提出新方法攻击文本生成图像模型的安全机制,高效绕过内容过滤。
PLA: Prompt Learning Attack against Text-to-Image Generative Models
- 利用多模态相似性设计黑盒下的梯度学习攻击框架
- 在多个主流模型上实现高于现有方法的攻击成功率
- 适合研究模型安全漏洞或对抗样本的学者参考
文本到图像(T2I)模型已广泛应用于各类场景。尽管取得成功,其被滥用可能导致生成不当内容(NSFW)。为探究T2I模型的安全脆弱性,本文研究在黑盒设置下绕过安全机制的对抗攻击。以往方法多依赖词语替换搜索对抗提示,受限于搜索空间,性能不如基于梯度的训练方法。然而,黑盒环境无法访问T2I模型的内部结构与参数,难以直接应用梯度驱动的攻击。为此,本文提出新型提示学习攻击框架(PLA),通过利用多模态相似性,设计适用于黑盒T2I模型的梯度训练策略。实验表明,该方法能有效攻击包括提示过滤器和事后安全检查器在内的多种黑盒T2I模型安全机制,相比当前最优方法具有更高成功率。警告:本文可能包含由模型生成的不当内容。
原文摘要 · Abstract (English)
Text-to-Image (T2I) models have gained widespread adoption across various applications. Despite the success, the potential misuse of T2I models poses significant risks of generating Not-Safe-For-Work (NSFW) content. To investigate the vulnerability of T2I models, this paper delves into adversarial attacks to bypass the safety mechanisms under black-box settings. Most previous methods rely on word substitution to search adversarial prompts. Due to limited search space, this leads to suboptimal performance compared to gradient-based training. However, black-box settings present unique challenges to training gradient-driven attack methods, since there is no access to the internal architecture and parameters of T2I models. To facilitate the learning of adversarial prompts in black-box settings, we propose a novel prompt learning attack framework (PLA), where insightful gradient-based training tailored to black-box T2I models is designed by utilizing multimodal similarities. Experiments show that our new method can effectively attack the safety mechanisms of black-box T2I models including prompt filters and post-hoc safety checkers with a high success rate compared to state-of-the-art methods. Warning: This paper may contain offensive model-generated content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。