让AI不再生成特定敏感内容,提升生成模型的安全性。
Erasing Concepts, Steering Generations: A Comprehensive Survey of Concept Suppression

- 按干预层级、优化结构、语义范围三维度分类概念删除方法
- 揭示删除精度、泛化能力与计算开销间的权衡关系
- 适合关注AI伦理安全与可控生成的研究者参考
文本到图像(T2I)模型在从自然语言提示生成高质量、多样化视觉内容方面展现出强大能力。然而,对敏感、受版权保护或有害图像的不受控生成带来了严重的伦理、法律与安全挑战。为应对这些问题,概念擦除范式应运而生,能够在保留模型整体功能的前提下,选择性移除特定语义概念。本文系统综述了T2I扩散模型中的概念擦除技术,从干预层级、优化结构和语义范围三个维度进行分类,构建多维分类体系,实现对不同方法的清晰对比,揭示擦除精度、泛化性与计算复杂度之间的根本权衡。文章还讨论了现有评估基准、标准化指标与实用数据集,指出现有评估在鲁棒性与实际效果方面的不足。最后,提出关键挑战与未来方向,包括概念表示解耦、自适应与增量擦除策略、对抗鲁棒性及新型生成架构。本综述旨在引导研究人员开发更安全、更符合伦理的生成模型,提供基础认知与可操作建议,推动生成式AI负责任发展。
原文摘要 · Abstract (English)
Text-to-Image (T2I) models have demonstrated impressive capabilities in generating high-quality and diverse visual content from natural language prompts. However, uncontrolled reproduction of sensitive, copyrighted, or harmful imagery poses serious ethical, legal, and safety challenges. To address these concerns, the concept erasure paradigm has emerged as a promising direction, enabling the selective removal of specific semantic concepts from generative models while preserving their overall utility. This survey provides a comprehensive overview and in-depth synthesis of concept erasure techniques in T2I diffusion models. We systematically categorize existing approaches along three key dimensions: intervention level, which identifies specific model components targeted for concept removal; optimization structure, referring to the algorithmic strategies employed to achieve suppression; and semantic scope, concerning the complexity and nature of the concepts addressed. This multi-dimensional taxonomy enables clear, structured comparisons across diverse methodologies, highlighting fundamental trade-offs between erasure specificity, generalization, and computational complexity. We further discuss current evaluation benchmarks, standardized metrics, and practical datasets, emphasizing gaps that limit comprehensive assessment, particularly regarding robustness and practical effectiveness. Finally, we outline major challenges and promising future directions, including disentanglement of concept representations, adaptive and incremental erasure strategies, adversarial robustness, and new generative architectures. This survey aims to guide researchers toward safer, more ethically aligned generative models, providing foundational knowledge and actionable recommendations to advance responsible development in generative AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。