arXiv:2411.18699cs.CRcs.CL2024-11

用单轮渐进攻击测试文生图模型的防护能力,发现主流模型可被绕过。

An indicator for effectiveness of text-to-image guardrails utilizing the Single-Turn Crescendo Attack (STCA)

  • 设计单轮渐进攻击,通过逐步升级提示词和建立信任欺骗模型
  • 成功攻破DALL-E 3的伦理防护,生成内容接近未过滤的Flux Schnell输出
  • 为评估文生图模型安全防护提供可复现的评测框架,适合安全研究者使用

单轮渐进攻击(STCA)最初由Aqrawi和Abbasi [2024] 提出,是一种用于绕过文本生成AI模型伦理防护的新方法,能诱导模型生成有害内容。该技术通过在单一提示中逐步增强上下文,并结合信任构建机制,隐蔽地欺骗模型产生非预期输出。本文将STCA扩展至文生图模型,验证其有效性:成功攻破广泛应用的DALL-E 3的防护机制,生成结果与作为基准对照的未过滤模型Flux Schnell输出相当。本研究为研究人员提供了评估文生图模型防护机制鲁棒性的系统性框架,并可用于量化其对对抗攻击的抵御能力。

原文摘要 · Abstract (English)

The Single-Turn Crescendo Attack (STCA), first introduced in Aqrawi and Abbasi [2024], is an innovative method designed to bypass the ethical safeguards of text-to-text AI models, compelling them to generate harmful content. This technique leverages a strategic escalation of context within a single prompt, combined with trust-building mechanisms, to subtly deceive the model into producing unintended outputs. Extending the application of STCA to text-to-image models, we demonstrate its efficacy by compromising the guardrails of a widely-used model, DALL-E 3, achieving outputs comparable to outputs from the uncensored model Flux Schnell, which served as a baseline control. This study provides a framework for researchers to rigorously evaluate the robustness of guardrails in text-to-image models and benchmark their resilience against adversarial attacks.

安全评测文生图对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。