提出噪声驱动框架NDM,精准识别并抑制文本生成中的隐性性意图。
NDM: A Noise-driven Detection and Mitigation Framework against Implicit Sexual Intentions in Text-to-Image Generation
- 利用生成早期的噪声分离特性,实现高效隐性恶意意图检测。
- 通过增强负向引导机制,抑制关键区域注意力,提升性内容压制效果。
- 在自然与对抗数据集上优于现有方法,且不损害模型生成能力。
尽管文生图扩散模型生成能力强大,但对隐性性暗示提示仍易生成不当内容。这类微妙线索常伪装为看似无害的词汇,因模型内部偏见意外触发性内容,引发严重伦理问题。现有检测方法多针对显性有害内容,难以捕捉此类隐性提示;微调虽部分有效,却可能降低生成质量,形成权衡。为此,我们提出首个噪声驱动的检测与缓解框架NDM,可在保持模型原有生成能力的同时,精准识别并抑制隐性恶意意图。核心创新包括:1)利用早期预测噪声的可分离性,构建高精度、高效噪声检测方法;2)提出噪声增强的自适应负向引导机制,通过优化初始噪声以抑制显著区域注意力,从而提升负向引导对性内容的抑制效果。实验在自然与对抗数据集上验证了NDM性能优于当前SOTA方法(如SLD、UCE、RECE等),代码与资源已开源。
原文摘要 · Abstract (English)
Despite the impressive generative capabilities of text-to-image (T2I) diffusion models, they remain vulnerable to generating inappropriate content, especially when confronted with implicit sexual prompts. Unlike explicit harmful prompts, these subtle cues, often disguised as seemingly benign terms, can unexpectedly trigger sexual content due to underlying model biases, raising significant ethical concerns. However, existing detection methods are primarily designed to identify explicit sexual content and therefore struggle to detect these implicit cues. Fine-tuning approaches, while effective to some extent, risk degrading the model's generative quality, creating an undesirable trade-off. To address this, we propose NDM, the first noise-driven detection and mitigation framework, which could detect and mitigate implicit malicious intention in T2I generation while preserving the model's original generative capabilities. Specifically, we introduce two key innovations: first, we leverage the separability of early-stage predicted noise to develop a noise-based detection method that could identify malicious content with high accuracy and efficiency; second, we propose a noise-enhanced adaptive negative guidance mechanism that could optimize the initial noise by suppressing the prominent region's attention, thereby enhancing the effectiveness of adaptive negative guidance for sexual mitigation. Experimentally, we validate NDM on both natural and adversarial datasets, demonstrating its superior performance over existing SOTA methods, including SLD, UCE, and RECE, etc. Code and resources are available at https://github.com/lorraine021/NDM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。