arXiv:2502.16423cs.CV2025-02TPAMI被引 7

统一攻击文本与图像防御,高效生成有害内容且难被发现。

Unified Prompt Attack Against Text-to-Image Generation Models

  • 联合攻击文本与视觉防御,支持梯度优化提升效率。
  • 仅需少量查询即可触发有害生成,规避检测。
  • 生成自然且语义精准的对抗提示,适合安全评估研究者。

文本到图像(T2I)模型发展迅速,但其生成有害内容的能力引发安全担忧。为此,我们提出UPAM框架,从攻击角度评估T2I模型鲁棒性。不同于仅关注文本防御的已有方法,UPAM统一攻击文本与视觉双重防御机制,并支持基于梯度的优化,突破了依赖枚举的效率瓶颈。针对模型因防御机制阻断图像输出的问题,引入球面探测学习(Sphere-Probing Learning, SPL),实现无图像输出情况下的优化。结合语义增强学习(Semantic-Enhancing Learning, SEL),确保攻击意图的语义一致性。通过上下文自然度增强(In-context Naturalness Enhancement, INE),提升对抗提示的自然性,降低人工识别概率。此外,为解决先前方法中迭代查询易被API防御系统检测的问题,提出可迁移攻击学习(Transferable Attack Learning, TAL),实现以最少查询完成有效攻击。大量实验验证了UPAM在效果、效率、自然度及低查询暴露率方面的优势。

原文摘要 · Abstract (English)

Text-to-Image (T2I) models have advanced significantly, but their growing popularity raises security concerns due to their potential to generate harmful images. To address these issues, we propose UPAM, a novel framework to evaluate the robustness of T2I models from an attack perspective. Unlike prior methods that focus solely on textual defenses, UPAM unifies the attack on both textual and visual defenses. Additionally, it enables gradient-based optimization, overcoming reliance on enumeration for improved efficiency and effectiveness. To handle cases where T2I models block image outputs due to defenses, we introduce Sphere-Probing Learning (SPL) to enable optimization even without image results. Following SPL, our model bypasses defenses, inducing the generation of harmful content. To ensure semantic alignment with attacker intent, we propose Semantic-Enhancing Learning (SEL) for precise semantic control. UPAM also prioritizes the naturalness of adversarial prompts using In-context Naturalness Enhancement (INE), making them harder for human examiners to detect. Additionally, we address the issue of iterative queries--common in prior methods and easily detectable by API defenders--by introducing Transferable Attack Learning (TAL), allowing effective attacks with minimal queries. Extensive experiments validate UPAM's superiority in effectiveness, efficiency, naturalness, and low query detection rates.

文本生成对抗攻击安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。