arXiv:2412.16039cs.CV2024-12被引 1

通过动态引导控制有害特征,实现生成图像的高质与安全兼顾。

SafeCFG: Controlling Harmful Features with Dynamic Safe Guidance for Safe Generation

  • 基于提示有害性动态调节无分类器指引过程
  • 在有害生成时显著偏离,清洁生成质量不变
  • 无需标注即可实现无监督安全对齐,适合实际部署

扩散模型在文本到图像任务中表现卓越,得益于无分类器指引(CFG)技术的引入,生成图像质量大幅提升。然而,攻击者可利用CFG恶意引导生成有害内容。现有安全对齐方法虽能降低风险,但常损害清洁图像质量。为此,本文提出SafeCFG,通过动态安全引导自适应调控CFG生成过程:根据提示有害性实时调整,仅在有害生成时引入显著偏离,保持清洁生成高质量。该方法可同时处理多种有害生成路径,有效消除有害元素且不损失质量。此外,SafeCFG具备图像有害性检测能力,支持无监督安全对齐,无需预先定义干净或有害标签。实验表明,使用SafeCFG生成的图像兼具高质与安全;采用该方法训练的安全扩散模型也展现出良好的安全性能。

原文摘要 · Abstract (English)

Diffusion models (DMs) have demonstrated exceptional performance in text-to-image tasks, leading to their widespread use. With the introduction of classifier-free guidance (CFG), the quality of images generated by DMs is significantly improved. However, one can use DMs to generate more harmful images by maliciously guiding the image generation process through CFG. Existing safe alignment methods aim to mitigate the risk of generating harmful images but often reduce the quality of clean image generation. To address this issue, we propose SafeCFG to adaptively control harmful features with dynamic safe guidance by modulating the CFG generation process. It dynamically guides the CFG generation process based on the harmfulness of the prompts, inducing significant deviations only in harmful CFG generations, achieving high quality and safety generation. SafeCFG can simultaneously modulate different harmful CFG generation processes, so it could eliminate harmful elements while preserving high-quality generation. Additionally, SafeCFG provides the ability to detect image harmfulness, allowing unsupervised safe alignment on DMs without pre-defined clean or harmful labels. Experimental results show that images generated by SafeCFG achieve both high quality and safety, and safe DMs trained in our unsupervised manner also exhibit good safety performance.

扩散模型安全生成无监督对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。