arXiv:2605.18150cs.AI2026-05

提出黑盒框架,让被擦除的概念在扩散模型中重新浮现。

Whispers in the Noise: Surrogate-Guided Concept Awakening via a Multi-Agent Framework

论文配图:Whispers in the Noise: Surrogate-Guided Concept Awakening via a Multi-Agent Framework
图 1 · 摘自论文原文
  • 通过代理框架从噪声状态启动去噪轨迹,绕过概念擦除
  • 无需参数、梯度或内部表示,即可在黑盒下精准唤醒概念
  • 揭示现有擦除方法的局限性,适合安全评估与模型审计

扩散模型(DMs)广泛用于文本到图像生成,但其强大生成能力也引发生成不当内容的风险。概念擦除旨在移除预训练模型中的特定概念,但近期研究发现此类方法常仅抑制而非彻底消除目标概念,使模型仍易受唤醒攻击。现有方法多依赖白盒访问(如优化或反演),而黑盒条件下的概念唤醒仍缺乏探索。本文从轨迹视角重审去噪过程,发现概念擦除主要干扰早期文本-语义对齐,但并未完全阻断语义信息沿去噪动态传播。随着生成进行,模型越来越依赖不断演化的噪声状态而非文本条件,从而产生绕过擦除映射的契机。为此,我们提出ConceptAgent:一种无需训练、黑盒、多代理框架,通过代理引导的噪声状态初始化去噪轨迹,实现被擦除概念的准确与可控唤醒。大量实验表明,ConceptAgent可在无模型参数、梯度或内部表示访问条件下,有效唤醒被擦除概念。结果揭示了当前概念擦除方法的根本局限,并为扩散模型中语义控制的动态特性提供了新见解。

原文摘要 · Abstract (English)

Diffusion models (DMs) are widely used for text-to-image generation, but their strong generative capabilities also raise concerns about unsafe or undesirable content. Concept erasure aims to mitigate these risks by removing specific concepts from pretrained models. However, recent studies show that such methods often suppress rather than fully eliminate target concepts, leaving models vulnerable to awakening attacks. Existing approaches primarily rely on white-box access through optimization or inversion, while concept awakening under black-box constraints remains underexplored. In this work, we revisit the denoising process from a trajectory perspective and show that concept erasure mainly disrupts early-stage text-semantic alignment but does not fully prevent semantic information from propagating along the denoising dynamics. As generation proceeds, the model increasingly depends on the evolving noisy state rather than textual conditions, which creates an opportunity to bypass erased mappings. Motivated by this observation, we propose ConceptAgent, a training-free, black-box, multi-agent framework that awakens erased concepts by initializing the denoising trajectory from surrogate-guided noisy states. Extensive experiments demonstrate that ConceptAgent enables accurate and controllable awakening of erased concepts under black-box settings without access to model parameters, gradients, or internal representations. These results highlight fundamental limitations of current concept erasure methods and provide new insights into the dynamic nature of semantic control in DMs.

扩散模型概念擦除黑盒攻击安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。