arXiv:2605.26332cs.CVcs.AI2026-05

黑盒攻击可绕过删去概念的图像生成模型,用自然语言提示生成高质量违规图像。

Erased but Exploitable: Black-box Embedding-Aware Prompting Against Unlearned Text-to-Image Diffusion Models

论文配图:Erased but Exploitable: Black-box Embedding-Aware Prompting Against Unlearned Text-to-Image Diffusion Models
图 1 · 摘自论文原文
  • 利用大模型迭代生成对抗性提示,基于嵌入空间搜索优化
  • 攻击成功率提升60%以上,平均仅需15个提示即成功
  • 提示难以被安全过滤器识别,适合研究模型漏洞与防御

机器退化旨在从预训练文本到图像扩散模型中移除特定概念,但已有白盒和黑盒攻击方法仍存在现实威胁模型缺陷:要么需要模型权重,要么生成的对抗提示为乱码,易被简单规则检测。本文提出BEAP,一种黑盒、嵌入感知的对抗提示攻击,通过大语言模型在文本空间中迭代生成有效提示,结合未学习概念存在性、文本-图像对齐度与图像质量多维度奖励信号进行优化。该方法在不触发安全过滤的前提下,生成高质量图像。大量实验表明,相比先前方法,BEAP将攻击成功率(ASR)提升超60%,每次成功攻击平均仅需15个提示。警告:本文包含可能令人不适或冒犯的模型输出。

原文摘要 · Abstract (English)

Machine unlearning aims to remove specific concepts from pretrained text-to-image diffusion models, yet several white- and black-box attacks have been introduced to make the model generate such unlearned concepts. These attacks, nevertheless, do not assume a realistic threat model, i.e. they either assume access to the model weights, or result in gibberish adversarial prompts that could be easily detected even through naive rule-based safeguarding. We aim to address this gap in this paper. We introduce BEAP, a black-box, embedding-aware adversarial prompting attack that leverages a large language model (LLM) to iteratively generate effective adversarial prompts and exploit such hidden vulnerabilities. BEAP performs an embedding-aware search in text space, combining multiple reward signals: unlearned concept presence, text-image alignment, and image quality, to refine generated prompts. Unlike previous attack methods, BEAP keeps its prompts undetectable to safety filters while producing high-quality images. Extensive experiments show that BEAP improves the Attack Success Rate (ASR) by more than 60% over prior methods, while requiring only an average of fifteen prompts per successful attack. Warning: This paper contains model outputs that may be offensive or upsetting in nature.

对抗攻击文本生成扩散模型安全漏洞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。