提出通用多模态越狱攻击,无需针对特定提示或图像优化即可突破文本和图像安全防护。
Universally Unfiltered and Unseen:Input-Agnostic Multimodal Jailbreaks against Text-to-Image Model Safeguards
- 在图像背景添加通用对抗补丁,绕过安全检查器
- 生成通用安全改写词集,突破提示过滤机制
- 对商用模型成功率提升4倍,适配性强
为应对文本到图像(T2I)模型生成不当内容的风险,各类文本提示过滤器与图像安全检查器已被部署。然而,现有越狱攻击局限于特定提示或图像的扰动,存在可扩展性差、优化耗时的问题。为此,本文提出通用无过滤、无见(U3-Attack)多模态越狱方法。该方法在图像背景上优化一个通用对抗补丁,以普遍绕过安全检查器;同时从敏感词生成通用安全改写词集,以普遍突破提示过滤器,并消除冗余计算。大量实验表明,U3-Attack在开源与商业T2I模型上均表现优越。例如,在集成提示过滤器与安全检查器的商用Runway-inpainting模型上,其成功率比当前最优的多模态越狱方法MMA-Diffusion高出约4倍。
原文摘要 · Abstract (English)
Various (text) prompt filters and (image) safety checkers have been implemented to mitigate the misuse of Text-to-Image (T2I) models in creating Not-Safe-For-Work (NSFW) content. In order to expose potential security vulnerabilities of such safeguards, multimodal jailbreaks have been studied. However, existing jailbreaks are limited to prompt-specific and image-specific perturbations, which suffer from poor scalability and time-consuming optimization. To address these limitations, we propose Universally Unfiltered and Unseen (U3)-Attack, a multimodal jailbreak attack method against T2I safeguards. Specifically, U3-Attack optimizes an adversarial patch on the image background to universally bypass safety checkers and optimizes a safe paraphrase set from a sensitive word to universally bypass prompt filters while eliminating redundant computations. Extensive experimental results demonstrate the superiority of our U3-Attack on both open-source and commercial T2I models. For example, on the commercial Runway-inpainting model with both prompt filter and safety checker, our U3-Attack achieves $~4\times$ higher success rates than the state-of-the-art multimodal jailbreak attack, MMA-Diffusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。