在生成过程中提前识别色情内容,提升扩散模型安全
Seeing It Before It Happens: In-Generation NSFW Detection for Diffusion-Based Text-to-Image Models
- 利用扩散过程中的噪声预测作为内部信号检测有害内容
- 对7类有害内容平均准确率达91.32%,优于7种基线方法
- 适合需要实时内容过滤的生成式AI系统开发者
基于扩散的文本到图像(T2I)模型虽能生成高质量图像,但也存在被滥用于生成不适宜工作场所(NSFW)内容的风险。以往检测方法主要聚焦于生成前的提示词过滤或生成后的图像审核,而扩散模型生成过程中的中间阶段尚未被充分探索用于NSFW检测。本文提出一种名为“生成中检测”(In-Generation Detection, IGD)的方法,利用扩散过程中预测的噪声作为内部信号,识别潜在有害内容。实验表明,该方法在7个NSFW类别上对普通和对抗性有害提示的平均检测准确率达91.32%,显著优于7种基线方法。
原文摘要 · Abstract (English)
Diffusion-based text-to-image (T2I) models enable high-quality image generation but also pose significant risks of misuse, particularly in producing not-safe-for-work (NSFW) content. While prior detection methods have focused on filtering prompts before generation or moderating images afterward, the in-generation phase of diffusion models remains largely unexplored for NSFW detection. In this paper, we introduce In-Generation Detection (IGD), a simple yet effective approach that leverages the predicted noise during the diffusion process as an internal signal to identify NSFW content. This approach is motivated by preliminary findings suggesting that the predicted noise may capture semantic cues that differentiate NSFW from benign prompts, even when the prompts are adversarially crafted. Experiments conducted on seven NSFW categories show that IGD achieves an average detection accuracy of 91.32% over naive and adversarial NSFW prompts, outperforming seven baseline methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。