用隐喻构造提示词,绕过图像生成模型的安全防护。
Metaphor-based Jailbreak Attacks on Text-to-Image Models
- 通过多智能体系统生成基于隐喻的对抗性提示。
- 在多种防御机制下攻击成功率更高,且查询次数更少。
- 揭示隐喻导致语义模糊,是绕过安全机制的关键。
文本到图像(T2I)模型通常内置防御机制以防止生成敏感内容。然而,近期的越狱攻击表明,对抗性提示能有效绕过这些机制,诱导模型生成敏感图像,暴露了严重的安全漏洞。现有方法隐含假设攻击者知晓部署的防御类型,限制了其对未知或多样防御机制的攻击效果。本文揭示了一种未被充分研究的漏洞:基于隐喻的越狱攻击(MJA),可在不预先了解防御类型的情况下,攻击多种防御机制。MJA包含两个模块:基于大语言模型的多智能体生成模块(LMAG)和对抗性提示优化模块(APO)。LMAG将隐喻类对抗提示生成分解为隐喻检索、上下文匹配和提示生成三个子任务,并由三个基于LLM的智能体协作探索不同隐喻与上下文,生成多样化对抗提示。为提升攻击效率,APO首先训练一个代理模型预测对抗提示的效果,并设计自适应获取策略识别最优提示。在具备多种外部与内部防御机制的T2I模型上进行的大量实验表明,MJA相比六种基线方法,在更低的查询次数下实现了更强的攻击性能。此外,深入的漏洞分析表明,隐喻类对抗提示通过引入语义模糊性规避安全机制,而敏感图像源于模型对隐藏语义的概率性解读。
原文摘要 · Abstract (English)
Text-to-image (T2I) models commonly incorporate defense mechanisms to prevent the generation of sensitive images. Unfortunately, recent jailbreak attacks have shown that adversarial prompts can effectively bypass these mechanisms and induce T2I models to produce sensitive content, revealing critical safety vulnerabilities. However, existing attack methods implicitly assume that the attacker knows the type of deployed defenses, which limits their effectiveness against unknown or diverse defense mechanisms. In this work, we reveal an underexplored vulnerability of T2I models to metaphor-based jailbreak attacks (MJA), which aims to attack diverse defense mechanisms without prior knowledge of their type by generating metaphor-based adversarial prompts. Specifically, MJA consists of two modules: an LLM-based multi-agent generation module (LMAG) and an adversarial prompt optimization module (APO). LMAG decomposes the generation of metaphor-based adversarial prompts into three subtasks: metaphor retrieval, context matching, and adversarial prompt generation. Subsequently, LMAG coordinates three LLM-based agents to generate diverse adversarial prompts by exploring various metaphors and contexts. To enhance attack efficiency, APO first trains a surrogate model to predict the attack results of adversarial prompts and then designs an acquisition strategy to adaptively identify optimal adversarial prompts. Extensive experiments on T2I models with various external and internal defense mechanisms demonstrate that MJA achieves stronger attack performance while using fewer queries, compared with six baseline methods. Additionally, we provide an in-depth vulnerability analysis suggesting that metaphor-based adversarial prompts evade safety mechanisms by inducing semantic ambiguity, while sensitive images arise from the model's probabilistic interpretation of concealed semantics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。