用轻量分布优化突破图像生成模型安全限制
JANUS: A Lightweight Framework for Jailbreaking Text-to-Image Models via Distribution Optimization
- 通过优化提示分布实现端到端黑盒攻击
- 在Stable Diffusion 3.5上攻击成功率提升至43.15%
- 适用于开源与商用模型,暴露现有防御漏洞
文本到图像(T2I)模型如Stable Diffusion和DALLE在越狱攻击下仍可能生成有害或不适宜内容,尽管已部署安全过滤器。现有越狱方法或依赖代理损失而非真实端到端目标,或依赖大规模且昂贵的强化学习训练生成器。为此,我们提出JANUS,一种轻量级框架,将越狱建模为在黑盒、端到端奖励信号下优化结构化提示分布的问题。JANUS用低维混合策略替代高容量生成器,基于两个语义锚定提示分布进行高效探索,同时保持目标语义。在现代T2I模型上,JANUS超越现有最先进越狱方法,在Stable Diffusion 3.5 Large Turbo上将攻击成功率(ASR-8)从25.30%提升至43.15%,且一致获得更高CLIP和NSFW得分。JANUS在开源与商业模型上均有效。这些发现揭示了当前T2I安全管道的结构性缺陷,推动更强大、分布感知的防御机制发展。警告:本文包含可能令人不适的模型输出。
原文摘要 · Abstract (English)
Text-to-image (T2I) models such as Stable Diffusion and DALLE remain susceptible to generating harmful or Not-Safe-For-Work (NSFW) content under jailbreak attacks despite deployed safety filters. Existing jailbreak attacks either rely on proxy-loss optimization instead of the true end-to-end objective, or depend on large-scale and costly RL-trained generators. Motivated by these limitations, we propose JANUS , a lightweight framework that formulates jailbreak as optimizing a structured prompt distribution under a black-box, end-to-end reward from the T2I system and its safety filters. JANUS replaces a high-capacity generator with a low-dimensional mixing policy over two semantically anchored prompt distributions, enabling efficient exploration while preserving the target semantics. On modern T2I models, we outperform state-of-the-art jailbreak methods, improving ASR-8 from 25.30% to 43.15% on Stable Diffusion 3.5 Large Turbo with consistently higher CLIP and NSFW scores. JANUS succeeds across both open-source and commercial models. These findings expose structural weaknesses in current T2I safety pipelines and motivate stronger, distribution-aware defenses. Warning: This paper contains model outputs that may be offensive.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。