用大模型自动生成绕过安全检测的有害图像提示,揭露文本转图像模型的真实风险。
GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models
- 训练专用大模型,结合监督微调与强化学习生成高毒害提示。
- 在商业模型上实现黑盒攻击,成功触发大量违规图像生成。
- 能发现真实威胁,适合评估和提升AI内容安全系统可靠性。
文本到图像(T2I)模型如Stable Diffusion发展迅速,广泛用于内容创作,但可能被滥用于生成色情或暴力等有害内容,带来重大安全风险。尽管多数平台部署了内容审核系统,但仍有漏洞可被恶意利用。现有红队测试与对抗攻击研究存在局限:部分研究虽能生成高度有害图像,但依赖易被过滤的对抗性提示;另一些研究虽能绕过安全机制,却无法生成真正有害内容,忽视了高风险提示的发现。因此,缺乏可靠工具评估防御后T2I模型的安全性。为此,我们提出GenBreak框架,通过微调红队大语言模型(LLM),系统探索T2I生成器的潜在漏洞。方法结合在精选数据集上的监督微调与与代理T2I模型交互的强化学习,通过多奖励信号引导模型生成兼具逃避能力、高毒性、语义连贯性和多样性的对抗提示。这些提示在商业T2I生成器上表现出强大黑盒攻击效果,揭示了实际且令人担忧的安全弱点。
原文摘要 · Abstract (English)
Text-to-image (T2I) models such as Stable Diffusion have advanced rapidly and are now widely used in content creation. However, these models can be misused to generate harmful content, including nudity or violence, posing significant safety risks. While most platforms employ content moderation systems, underlying vulnerabilities can still be exploited by determined adversaries. Recent research on red-teaming and adversarial attacks against T2I models has notable limitations: some studies successfully generate highly toxic images but use adversarial prompts that are easily detected and blocked by safety filters, while others focus on bypassing safety mechanisms but fail to produce genuinely harmful outputs, neglecting the discovery of truly high-risk prompts. Consequently, there remains a lack of reliable tools for evaluating the safety of defended T2I models. To address this gap, we propose GenBreak, a framework that fine-tunes a red-team large language model (LLM) to systematically explore underlying vulnerabilities in T2I generators. Our approach combines supervised fine-tuning on curated datasets with reinforcement learning via interaction with a surrogate T2I model. By integrating multiple reward signals, we guide the LLM to craft adversarial prompts that enhance both evasion capability and image toxicity, while maintaining semantic coherence and diversity. These prompts demonstrate strong effectiveness in black-box attacks against commercial T2I generators, revealing practical and concerning safety weaknesses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。