arXiv:2505.18979cs.LG2025-05被引 5

提出自动越狱框架OptJail,高效绕过图文安全过滤器。

Dynamic Optimization and Safety Indicator Injection for Jailbreaking Text-to-Image Models with Multimodal Safety Filters

  • 动态优化提示词并注入安全标识,结合文本与图像反馈迭代攻击。
  • 使ShieldLM-7B越狱成功率从8.9%提升至99.0%,图像语义质量提升。
  • 可泛化到未知过滤器,适用于DALL·E 3等主流模型。

文本到图像(T2I)模型可能生成不适宜工作场合(NSFW)内容,促使采用文本与图像双重过滤的多阶段安全防护流程。基于大语言模型的过滤器能检测关键词之外的潜在意图,使传统词级扰动攻击失效。我们的评估表明,现有越狱方法在绕过过滤器与保持语义一致性之间存在显著权衡,且需大量查询才能成功。我们提出 extbf{OptJail},一种结合动态提示优化与多模态反馈的自动化越狱框架。其核心包含两部分:(i) 动态优化,通过文本过滤器反馈与语义一致性约束,迭代重写提示词生成对抗样本;(ii) 自适应安全指标注入,将注入良性视觉线索建模为强化学习问题,以规避图像级过滤。OptJail实现当前最优性能,将ShieldLM-7B的越狱成功率从8.9%(Sneakyprompt)提升至99.0%,同时CLIP分数由0.2637升至0.2762。该方法还可泛化至未见过滤器,并成功攻破DALL·E 3。机制分析揭示防御失效原因:优化提示词被投影至过滤器表示空间的“安全”区域,但在生成模型语义空间中几乎保持不变;注入的安全指标引导图像检测器注意力远离NSFW内容,转而关注良性视觉线索。本研究揭示了当前多模态防御体系的系统性漏洞,呼吁构建更强的自适应防护机制。

原文摘要 · Abstract (English)

Text-to-image (T2I) models can generate not-safe-for-work (NSFW) content, motivating multi-stage safety pipelines with both text and image filters. Newer LLM-based filters detect latent intent beyond keywords, making token-level perturbation attacks unreliable. Our evaluation further shows that existing jailbreak methods exhibit a sharp trade-off between filter evasion and semantic fidelity, while also requiring excessive queries to succeed. We introduce \textbf{OptJail}, an automated jailbreak framework that combines dynamic prompt optimization with multimodal feedback. It consists of two key components: (i) \textit{Dynamic Optimization}, an iterative process that leverages text-filter feedback and semantic consistency to rewrite prompts into adversarial variants; and (ii) \textit{Adaptive Safety Indicator Injection}, which formulates the injection of benign visual cues as a reinforcement learning problem to bypass image-level filters. OptJail achieves state-of-the-art performance, increasing the ShieldLM-7B bypass rate from 8.9\% (Sneakyprompt) to 99.0\%, improving CLIP score from 0.2637 to 0.2762. Moreover, it generalizes to unseen filters and successfully jailbreaks DALL E 3 in our evaluation. Mechanistic analysis reveals why these defenses fail: optimized prompts are projected into the ``safe'' region of the filter's representation space yet remain nearly stationary in the generative model's semantic space, and injected safety indicators redirect image detectors' attention away from NSFW content toward benign visual cues. This study reveals systemic vulnerabilities in current multimodal defenses and motivates stronger adaptive defenses.

越狱攻击多模态安全扩散模型提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。