只需改写提示词,就能绕过图文模型的安全过滤
Low-Effort Jailbreak Attacks Against Text-to-Image Safety Filters
- 用艺术化、伪教育等语言伪装恶意意图
- 在多个主流模型上实现最高74.47%的攻击成功率
- 适合关注生成模型安全漏洞的研究者和开发者
文本到图像生成模型广泛应用于创意工具和在线平台。为防止滥用,这些系统依赖于安全过滤器和内容审核流程以阻止有害或违规内容。本文表明,现代文本到图像模型仍易受低投入劫持攻击,仅需自然语言提示即可实现突破。我们系统研究了无需模型访问、优化或对抗训练的基于提示的策略,提出视觉劫持技术分类:艺术重构、材料替换、伪教育框架、生活方式美学伪装、模糊动作替代。这些策略通过将不良意图隐藏在无害语义上下文中,利用提示审核与视觉安全过滤间的漏洞。我们在多个先进文本到图像系统上评估,发现简单语言修改可稳定绕过现有防护并生成受限内容。所有测试模型和攻击类别中,攻击成功率达最高74.47%,揭示表层提示过滤与语义理解间的关键差距。
原文摘要 · Abstract (English)
Text-to-image generative models are widely deployed in creative tools and online platforms. To mitigate misuse, these systems rely on safety filters and moderation pipelines that aim to block harmful or policy violating content. In this work we show that modern text-to-image models remain vulnerable to low-effort jailbreak attacks that require only natural language prompts. We present a systematic study of prompt-based strategies that bypass safety filters without model access, optimization, or adversarial training. We introduce a taxonomy of visual jailbreak techniques including artistic reframing, material substitution, pseudo-educational framing, lifestyle aesthetic camouflage, and ambiguous action substitution. These strategies exploit weaknesses in prompt moderation and visual safety filtering by masking unsafe intent within benign semantic contexts. We evaluate these attacks across several state-of-the-art text-to-image systems and demonstrate that simple linguistic modifications can reliably evade existing safeguards and produce restricted imagery. Our findings highlight a critical gap between surface-level prompt filtering and the semantic understanding required to detect adversarial intent in generative media systems. Across all tested models and attack categories we observe an attack success rate (ASR) of up to 74.47%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。