arXiv:2509.21008cs.CV2025-09被引 8

只调一个神经元,就能精准清除文本生成图像中的有害内容

A Single Neuron Works: Precise Concept Erasure in Text-to-Image Diffusion Models

  • 用稀疏自编码器让每个神经元对应一个语义概念
  • 通过激活模式识别出有害概念专属神经元并抑制其响应
  • 仅改一个神经元就实现精准擦除,且不影响其他图像生成

文本到图像模型具备强大的图像生成能力,但也存在生成有害内容的安全风险。现有概念擦除方法难以在精确移除目标概念的同时保持图像质量。本文提出基于单神经元的概念擦除方法(SNCE),通过训练稀疏自编码器(SAE)将文本嵌入映射到稀疏解耦的潜在空间,使单个神经元紧密对齐原子语义概念。为准确识别有害概念相关神经元,设计基于激活模式调制频率评分的新方法。通过抑制特定神经元激活,SNCE实现对有害内容的精准擦除,对图像质量影响极小。在多个基准测试中,该方法在目标概念擦除上达到当前最优效果,同时保留模型对非目标概念的生成能力。此外,该方法对对抗攻击表现出强鲁棒性,显著优于现有方法。

原文摘要 · Abstract (English)

Text-to-image models exhibit remarkable capabilities in image generation. However, they also pose safety risks of generating harmful content. A key challenge of existing concept erasure methods is the precise removal of target concepts while minimizing degradation of image quality. In this paper, we propose Single Neuron-based Concept Erasure (SNCE), a novel approach that can precisely prevent harmful content generation by manipulating only a single neuron. Specifically, we train a Sparse Autoencoder (SAE) to map text embeddings into a sparse, disentangled latent space, where individual neurons align tightly with atomic semantic concepts. To accurately locate neurons responsible for harmful concepts, we design a novel neuron identification method based on the modulated frequency scoring of activation patterns. By suppressing activations of the harmful concept-specific neuron, SNCE achieves surgical precision in concept erasure with minimal disruption to image quality. Experiments on various benchmarks demonstrate that SNCE achieves state-of-the-art results in target concept erasure, while preserving the model's generation capabilities for non-target concepts. Additionally, our method exhibits strong robustness against adversarial attacks, significantly outperforming existing methods.

概念擦除扩散模型安全生成神经元操控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。