恶搞触发器让删除概念失效,模型仍会生成有害内容
Erased but Not Forgotten: How Backdoors Compromise Concept Erasure
- 用恶意触发器绑定目标概念,躲过模型概念删除
- 6种主流删除方法均被攻破,最高94%有害内容泄露
- 适合研究模型安全与隐私保护的学者参考
文本到图像扩散模型的扩张引发了对有害输出的担忧,如伪造公众人物形象或色情内容。为缓解风险,已有方法通过微调实现概念擦除,但尚不清楚这些方法是否真正移除了所有关联,还是仅隐藏了表面连接。本文揭示了一种关键漏洞——擦除规避后门(EEB):攻击者将后门触发器绑定至待删除概念,该恶意关联在后续擦除中依然存活。我们证明黑盒与白盒攻击均可实现此威胁。在六种前沿擦除方法中,包括明确搜索目标概念替代表示的鲁棒方法,EEB始终暴露有害内容:对名人身份擦除最高达82%成功率,对象擦除达94%,显性内容暴露最多放大16倍。尽管暴露了现有擦除方法的盲点,EEB亦可作为未来概念擦除技术的压力测试工具。
原文摘要 · Abstract (English)
The expansion of text-to-image diffusion models has raised concerns about harmful outputs, from fabricated depictions of public figures to sexually explicit imagery. To mitigate such risks, prior work has proposed concept erasure methods that aim to sever unwanted concepts from the model via fine-tuning, yet it remains unclear whether these approaches truly remove all links to the harmful concept or merely conceal superficial connections. In this work, we reveal a critical vulnerability, the Erasure Evasion Backdoor (EEB): an adversary binds a backdoor trigger to a concept slated for removal, and this malicious link survives subsequent erasure. We show that both black-box and white-box adversaries can instantiate this threat. Across six state-of-the-art erasure methods, including robust ones that explicitly search for alternative representations of the target concept, EEB consistently exposes harmful content: up to 82% success against celebrity-identity unlearning, up to 94% for object erasure, and up to 16 times amplification of explicit-content exposure. While EEB uncovers a blind spot in current erasure methods, it also provides a diagnostic tool for stress-testing future concept erasure techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。