arXiv:2503.01839cs.CRcs.AI2025-03Conference of the …被引 4

用微调大模型高效生成绕过图像生成安全过滤的恶意提示。

Jailbreaking Safeguarded Text-to-Image Models via Large Language Models

  • 用微调大模型生成一次就能绕过安全过滤的对抗性提示。
  • 在三个恶意提示数据集上,成功突破五种安全防护机制。
  • 可作为其他攻击的辅助工具,提升整体破解效率。

文本到图像模型可能生成有害内容(如色情图像),尤其在输入不安全提示时。为应对该问题,常在模型上添加安全过滤器或对模型进行对齐以减少有害输出。然而,当攻击者精心设计对抗性提示时,这些防御仍易被突破。本文提出 extsc{Alg},一种利用微调的大语言模型(AttackLLM)实现对带安全防护的文本到图像模型的越狱攻击。与需要多次查询目标模型的传统无框攻击不同,本方法在微调后可一次性生成有效对抗性提示。我们在三个不安全提示数据集上,针对五种安全防护机制进行了评估,结果表明该方法能有效绕过安全屏障,优于现有无框攻击,并能增强其他查询型攻击的效果。

原文摘要 · Abstract (English)

Text-to-Image models may generate harmful content, such as pornographic images, particularly when unsafe prompts are submitted. To address this issue, safety filters are often added on top of text-to-image models, or the models themselves are aligned to reduce harmful outputs. However, these defenses remain vulnerable when an attacker strategically designs adversarial prompts to bypass these safety guardrails. In this work, we propose \alg, a method to jailbreak text-to-image models with safety guardrails using a fine-tuned large language model. Unlike other query-based jailbreak attacks that require repeated queries to the target model, our attack generates adversarial prompts efficiently after fine-tuning our AttackLLM. We evaluate our method on three datasets of unsafe prompts and against five safety guardrails. Our results demonstrate that our approach effectively bypasses safety guardrails, outperforms existing no-box attacks, and also facilitates other query-based attacks.

越狱攻击安全防护大模型图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。