通过分析模型防御信号,智能生成能突破文本到图像模型安全限制的提示。
Dynamic Defense Profiling Enables Cognitive Jailbreak of Text-to-Image Models

- 将攻击视为对隐藏防御机制的信念推理,动态建模失败反馈
- 在六个防御设置下成功率高达95.62%,跨多平台通用性强
- 适合研究模型安全漏洞与对抗攻击的开发者或安全研究人员
文本到图像生成模型虽已取得显著进展,但仍易被滥用生成不适宜内容(NSFW)。现有越狱攻击多依赖启发式提示工程或黑盒优化,仅将模型反馈视为成功/失败的二元信号,忽略了文本拒绝、视觉阻断、语义净化等多种失败模式所蕴含的丰富信息,导致探索效率低且语义严重退化。本文提出MIND框架,将对抗提示生成重构为对隐含防御机制的信念状态推断问题。该框架通过多模态判别器细粒度分解反馈信号,利用防御剖析器迭代更新对防御机制的认知,并结合元记忆模块检索历史有效攻击策略。三者统一于推理驱动的演化优化流程中,实现自适应且语义一致的越狱生成。在I2P基准测试中,MIND在六种典型预处理与后处理防御设置下,对Stable Diffusion v1.5模型达到95.62%的攻击成功率(ASR),显著优于现有方法。同时,在四个主流商用T2I系统上验证有效性,于Wan-2.5上最高达成91.58%的ASR。
原文摘要 · Abstract (English)
Text-to-Image (T2I) generative models have achieved remarkable progress in synthesizing high-quality visual content, yet they remain vulnerable to adversarial misuse, particularly in generating Not-Safe-For-Work (NSFW) images. Most existing jailbreak attacks primarily rely on heuristic prompt engineering or black-box optimization, treating model feedback as a binary signal (success or failure). This coarse-grained paradigm overlooks the rich information embedded in diverse failure modes, such as textual refusal, visual blocking, and semantic sanitization, resulting in inefficient exploration and severe semantic collapse. In this paper, we propose MIND, a cognitive jailbreak framework that reframes adversarial prompt generation as a belief-state inference problem over latent defense mechanisms. Instead of blindly searching for bypass prompts, MIND actively models the target system's latent defense mechanisms by interpreting multi-modal feedback as high-density signals. Specifically, the framework integrates three core components: (1) a Multi-modal Judge for fine-grained feedback decomposition, (2) a Defense Profiler for iterative belief updating, and (3) a Meta-Memory module for retrieving historically effective attack strategies. These components are unified within a reasoning-driven evolutionary optimization process, enabling adaptive and semantically consistent jailbreak generation. Extensive experiments on the I2P benchmark demonstrate the effectiveness of MIND. Under six representative pre-processing and post-processing defense settings applied to the Stable Diffusion v1.5 model, MIND achieves an Attack Success Rate (ASR) of 95.62%, significantly outperforming existing methods. Additionally, the effectiveness of the proposed framework is validated across four widely used commercial T2I systems, achieving the highest ASR of 91.58% on Wan-2.5.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。