arXiv:2503.15197cs.CV2025-03CVPR被引 12

通过优化提示词令牌实现扩散模型生成内容的自我检测与安全调控。

Detect-and-Guide: Self-regulation of Diffusion Models for Safe Text-to-Image Generation via Guideline Token Optimization

  • 利用优化令牌的交叉注意力图检测生成过程中的有害概念。
  • 自适应调节安全引导强度与编辑区域,精准抑制不安全内容。
  • 无需微调模型,兼顾安全性与文本跟随能力,适合实际应用。

文本到图像扩散模型在合成任务中达到顶尖水平,但其生成有害内容的潜在风险日益引发关注。现有事后干预技术如概念消解和安全引导虽有效,但通过微调模型权重或调整隐藏状态的方式难以解释,且严重影响采样轨迹,阻碍其在真实场景中的应用。本文提出安全生成框架Detect-and-Guide(DAG),利用扩散模型内部知识,在采样过程中实现自我诊断与细粒度自我调节。DAG首先通过优化令牌的精细化交叉注意力图从噪声潜在表示中检测有害概念,随后以自适应强度和编辑区域施加安全引导,消除不安全生成。该优化仅需少量标注数据,可生成具有泛化性与概念特异性的精确检测图。此外,DAG无需微调扩散模型,因此不损失生成多样性。在去除性内容的实验中,DAG在多概念真实提示下实现了最先进的安全生成性能,平衡了有害性缓解与文本跟随能力。

原文摘要 · Abstract (English)

Text-to-image diffusion models have achieved state-of-the-art results in synthesis tasks; however, there is a growing concern about their potential misuse in creating harmful content. To mitigate these risks, post-hoc model intervention techniques, such as concept unlearning and safety guidance, have been developed. However, fine-tuning model weights or adapting the hidden states of the diffusion model operates in an uninterpretable way, making it unclear which part of the intermediate variables is responsible for unsafe generation. These interventions severely affect the sampling trajectory when erasing harmful concepts from complex, multi-concept prompts, thus hindering their practical use in real-world settings. In this work, we propose the safe generation framework Detect-and-Guide (DAG), leveraging the internal knowledge of diffusion models to perform self-diagnosis and fine-grained self-regulation during the sampling process. DAG first detects harmful concepts from noisy latents using refined cross-attention maps of optimized tokens, then applies safety guidance with adaptive strength and editing regions to negate unsafe generation. The optimization only requires a small annotated dataset and can provide precise detection maps with generalizability and concept specificity. Moreover, DAG does not require fine-tuning of diffusion models, and therefore introduces no loss to their generation diversity. Experiments on erasing sexual content show that DAG achieves state-of-the-art safe generation performance, balancing harmfulness mitigation and text-following performance on multi-concept real-world prompts.

扩散模型安全生成自调节

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。