动态校准负向引导,让生成图像更安全且不失真
Concept Removal Guidance: Evidence-Calibrated Negative Guidance for Safe Diffusion Sampling

- 根据噪声预测估算不想要概念的出现概率,自适应调整负向引导强度
- 在红队测试中降低攻击成功率,同时保持正常提示的生成质量
- 无需训练即可抑制风格、暴力等额外内容,适合实际部署场景
文本到图像扩散模型仍易受恶意提示诱导生成违规内容,亟需可靠的推理时控制机制。现有负向引导方法使用固定权重,常导致安全与保真度之间的权衡:过度使用会引入伪影或提示漂移,不足则无法抵御攻击。动态变体依赖后验几率信号重调权重,对开放词汇组合提示敏感;轻量级相似性方法则忽略去噪过程中图像证据的演化。本文提出无需训练的「概念移除引导」(CRG),通过模型噪声预测估计每一步中不希望出现的概念存在程度,并通过闭式约束更新自适应校准负向引导,使目标概念存在阈值达标的同时最小化对条件轨迹的扰动。在多个红队测试基准上,CRG有效降低攻击成功率并保持良性生成质量,且无需微调或外部分类器即可扩展至抑制艺术家风格、暴力等额外目标。
原文摘要 · Abstract (English)
Text-to-image diffusion models remain vulnerable to adversarial prompts that elicit disallowed content, motivating reliable inference-time controls. A popular approach is negative guidance, which subtracts a negative prompt direction with a fixed weight. However, it often forces a safety-fidelity trade-off, causing artifacts or prompt drift when over-applied and failing under attacks when under-applied. Dynamic variants reweight guidance using posterior-odds signals, which can be brittle for open-vocabulary compositional prompts, while lightweight similarity-based methods ignore the evolving image evidence along the denoising trajectory. We introduce Concept Removal Guidance (CRG), a training-free method that estimates unwanted-concept presence at each diffusion step from the model's noise predictions, and adaptively calibrates negative guidance via a closed-form constrained update enforcing a target presence threshold while minimally perturbing the conditional trajectory. Across red-teaming benchmarks, CRG reduces attack success rates while preserving benign fidelity, and extends to additional suppression targets such as artist style and violence without fine-tuning or external classifiers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。