只调一个神经元,就能精准清除文本生成图像中的有害内容
A Single Neuron Works: Precise Concept Erasure in Text-to-Image Diffusion Models
- 用稀疏自编码器让每个神经元对应一个语义概念
- 通过激活模式识别出有害概念专属神经元并抑制其响应
- 仅改一个神经元就实现精准擦除,且不影响其他图像生成
文本到图像模型具备强大的图像生成能力,但也存在生成有害内容的安全风险。现有概念擦除方法难以在精确移除目标概念的同时保持图像质量。本文提出基于单神经元的概念擦除方法(SNCE),通过训练稀疏自编码器(SAE)将文本嵌入映射到稀疏解耦的潜在空间,使单个神经元紧密对齐原子语义概念。为准确识别有害概念相关神经元,设计基于激活模式调制频率评分的新方法。通过抑制特定神经元激活,SNCE实现对有害内容的精准擦除,对图像质量影响极小。在多个基准测试中,该方法在目标概念擦除上达到当前最优效果,同时保留模型对非目标概念的生成能力。此外,该方法对对抗攻击表现出强鲁棒性,显著优于现有方法。
原文摘要 · Abstract (English)
Text-to-image models exhibit remarkable capabilities in image generation. However, they also pose safety risks of generating harmful content. A key challenge of existing concept erasure methods is the precise removal of target concepts while minimizing degradation of image quality. In this paper, we propose Single Neuron-based Concept Erasure (SNCE), a novel approach that can precisely prevent harmful content generation by manipulating only a single neuron. Specifically, we train a Sparse Autoencoder (SAE) to map text embeddings into a sparse, disentangled latent space, where individual neurons align tightly with atomic semantic concepts. To accurately locate neurons responsible for harmful concepts, we design a novel neuron identification method based on the modulated frequency scoring of activation patterns. By suppressing activations of the harmful concept-specific neuron, SNCE achieves surgical precision in concept erasure with minimal disruption to image quality. Experiments on various benchmarks demonstrate that SNCE achieves state-of-the-art results in target concept erasure, while preserving the model's generation capabilities for non-target concepts. Additionally, our method exhibits strong robustness against adversarial attacks, significantly outperforming existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。