arXiv:2412.07249cs.CVcs.AI2024-12被引 3

用语义后门阻止文本生成不当内容,高效且不降质。

Buster: Implanting Semantic Backdoor into Text Encoder to Mitigate NSFW Content Generation

  • 在文本编码器中植入语义触发的后门,拦截有害提示。
  • 91.2%以上消除不当内容,五分钟内完成微调。
  • 适合需快速部署安全机制的生成模型应用。

深度学习模型在数字时代兴起,引发了对生成不当内容(NSFW)的严重担忧。现有防御方法主要依赖模型微调和事后内容审核,但普遍存在可扩展性差、良性图像生成质量下降或推理成本高等问题。为此,我们提出创新框架Buster,通过向文本编码器植入后门以防止不当内容生成。Buster利用深层语义信息而非显式提示作为触发条件,将不当提示引导至目标良性提示。此外,Buster采用基于能量的训练数据生成方法,通过Langevin动力学实现对抗性知识增强,确保有害概念定义的鲁棒性。该方法在消除不当内容方面展现出卓越的韧性与可扩展性。特别地,Buster仅需五分钟即可完成文本到图像模型文本编码器的微调,效率极高。大量实验表明,Buster优于九个最先进基线,在至少91.2%的范围内有效移除不当内容,同时保持无害图像生成质量。

原文摘要 · Abstract (English)

The rise of deep learning models in the digital era has raised substantial concerns regarding the generation of Not-Safe-for-Work (NSFW) content. Existing defense methods primarily involve model fine-tuning and post-hoc content moderation. Nevertheless, these approaches largely lack scalability in eliminating harmful content, degrade the quality of benign image generation, or incur high inference costs. To address these challenges, we propose an innovative framework named \textit{Buster}, which injects backdoors into the text encoder to prevent NSFW content generation. Buster leverages deep semantic information rather than explicit prompts as triggers, redirecting NSFW prompts towards targeted benign prompts. Additionally, Buster employs energy-based training data generation through Langevin dynamics for adversarial knowledge augmentation, thereby ensuring robustness in harmful concept definition. This approach demonstrates exceptional resilience and scalability in mitigating NSFW content. Particularly, Buster fine-tunes the text encoder of Text-to-Image models within merely five minutes, showcasing its efficiency. Our extensive experiments denote that Buster outperforms nine state-of-the-art baselines, achieving a superior NSFW content removal rate of at least 91.2\% while preserving the quality of harmless images.

内容安全文本生成后门防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。