让AI绘画模型自动屏蔽特定风格或内容,避免侵权与有害生成。
A Comprehensive Survey on Concept Erasure in Text-to-Image Diffusion Models
- 按修改方式分三类:调参微调、快速公式修正、推理时干预。
- 提出评估体系,涵盖数据集与指标,统一测试擦除效果。
- 适合关注AI伦理、内容安全的研究者和开发者。
文本到图像(T2I)模型在生成高质量、多样化视觉内容方面取得了显著进展。然而,其再现受版权保护的风格、敏感图像及有害内容的能力引发了重大伦理与法律问题。概念擦除通过主动修改T2I模型,防止生成不期望的内容,提供了一种替代外部过滤的解决方案。本文系统梳理了概念擦除方法,依据优化策略与修改架构组件的不同,将其分为三类:参数更新的微调方法、高效编辑的闭式解法,以及无需修改权重的推理时干预。此外,还探讨了绕过擦除技术的对抗攻击及新兴防御手段。为支持后续研究,本文整合了关键数据集、评估指标与基准测试,用于衡量擦除有效性与模型鲁棒性。本综述全面呈现了概念擦除的发展现状、挑战与未来方向。
原文摘要 · Abstract (English)
Text-to-Image (T2I) models have made remarkable progress in generating high-quality, diverse visual content from natural language prompts. However, their ability to reproduce copyrighted styles, sensitive imagery, and harmful content raises significant ethical and legal concerns. Concept erasure offers a proactive alternative to external filtering by modifying T2I models to prevent the generation of undesired content. In this survey, we provide a structured overview of concept erasure, categorizing existing methods based on their optimization strategies and the architectural components they modify. We categorize concept erasure methods into fine-tuning for parameter updates, closed-form solutions for efficient edits, and inference-time interventions for content restriction without weight modification. Additionally, we explore adversarial attacks that bypass erasure techniques and discuss emerging defenses. To support further research, we consolidate key datasets, evaluation metrics, and benchmarks for assessing erasure effectiveness and model robustness. This survey serves as a comprehensive resource, offering insights into the evolving landscape of concept erasure, its challenges, and future directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。