arXiv:2505.15450cs.CV2025-05被引 3

系统评估文本到图像模型中NSFW内容消除方法的有效性

Comprehensive Evaluation and Analysis for NSFW Concept Erasure in Text-to-Image Diffusion Models

  • 构建全流程工具链,统一测试不同消除方法
  • 揭示机制与实测效果间的关联,提供实用指导
  • 适合关注生成安全、内容过滤的研究者与开发者

文本到图像扩散模型在多个领域广泛应用,展现出强大的创造力。然而,其强大的泛化能力可能导致生成不当内容(NSFW),威胁安全部署。尽管已有多种概念消除方法被提出以缓解此问题,但缺乏在多种场景下的系统性评估。为此,我们引入一个全链条工具包,开展首个针对NSFW概念消除方法的系统性研究。通过分析底层机制与实证结果之间的相互作用,提供深入洞察与实际应用建议,旨在推动对扩散模型内容安全的理解,并为该关键领域的未来研究与发展奠定坚实基础。

原文摘要 · Abstract (English)

Text-to-image diffusion models have gained widespread application across various domains, demonstrating remarkable creative potential. However, the strong generalization capabilities of diffusion models can inadvertently lead to the generation of not-safe-for-work (NSFW) content, posing significant risks to their safe deployment. While several concept erasure methods have been proposed to mitigate the issue associated with NSFW content, a comprehensive evaluation of their effectiveness across various scenarios remains absent. To bridge this gap, we introduce a full-pipeline toolkit specifically designed for concept erasure and conduct the first systematic study of NSFW concept erasure methods. By examining the interplay between the underlying mechanisms and empirical observations, we provide in-depth insights and practical guidance for the effective application of concept erasure methods in various real-world scenarios, with the aim of advancing the understanding of content safety in diffusion models and establishing a solid foundation for future research and development in this critical area.

内容安全扩散模型图像生成去污

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。