用自然语言生成恶意图像提示,自动突破安全防护。
Red-Teaming Text-to-Image Models via In-Context Experience Replay and Semantic-Preserving Prompt Rewriting
- 通过大模型重写提示词,生成流畅且保持恶意意图的攻击语句。
- 在六种安全机制下表现超越七种基线方法,30%以上提示可跨系统通用。
- 适合研究模型安全、对抗攻击或内容过滤的开发者使用。
理解文本到图像(T2I)模型生成有害内容的能力对安全与合规至关重要。然而,人工红队测试成本高且不一致,亟需能模拟真实滥用行为的自动化工具。现有方法或需白盒访问,或无法跨防御泛化,或产生难以解释的对抗性标记;而生成既流畅又保留原始恶意意图的提示仍鲜有探索。本文提出ICER框架,包含两个组件:基于大模型的提示重写器,生成自然语言攻击提示;以及上下文经验回放机制,将成功越狱模式积累为可复用先验。二者通过贝叶斯优化整合,实现高效利用已知攻击策略与探索新策略的平衡。在六种安全机制上的实验表明,ICER优于七种基线,在标准与语义保持评估中均表现更优,超过30%生成提示可迁移至DALL-E 3和Midjourney等商用系统。
原文摘要 · Abstract (English)
Understanding the capabilities of text-to-image (T2I) models in harmful content generation is essential to safety and compliance. However, human red-teaming is costly and inconsistent, driving the need for automatic tools that simulate realistic misuse attempts. Existing methods either require white-box access, fail to generalize across defenses, or produce uninterpretable adversarial tokens, while generating fluent prompts that preserve the original harmful intent remains underexplored despite its practical relevance. We propose ICER, a black-box framework that addresses this gap through two components: an LLM-based rewriter that produces fluent, natural-language adversarial prompts, and in-context experience replay that accumulates successful jailbreaking patterns into a reusable prior. These components are integrated via bandit optimization, enabling ICER to efficiently balance exploiting proven attack strategies with exploring new ones. Experiments across six safety mechanisms show that ICER outperforms seven baselines under both standard and semantics-preserving evaluation, with over 30% of generated prompts transferring to commercial systems like DALL-E 3 and Midjourney.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。