提出统一框架,让生成模型在早期阶段强化安全引导,后期自动减弱以保证质量。
Safety-Guided Flow (SGF): A Unified Framework for Negative Guidance in Safe Generation
- 用能量函数统一处理几何约束与数据驱动的安全引导
- 发现安全引导需在早期强、后期弱,存在关键时间窗口
- 适用于机器人路径生成等真实安全生成场景
扩散模型和流模型的安全机制目前沿两条路径发展:一是机器人规划中使用控制屏障函数,在每步去噪时显式施加几何约束以避开障碍物;二是近期基于数据的负向引导方法,能抑制有害内容并提升生成多样性,但依赖启发式规则,未明确安全引导的必要时机。本文首次提出基于最大均值差异(MMD)势能的统一概率框架,将Shielded Diffusion和Safe Denoiser均视为针对不安全样本的能量型负向引导实例。进一步通过控制屏障函数分析,证明存在一个关键时间窗口,此时负向引导必须较强;超出该窗口后,引导应衰减至零,以保障安全且高质量生成。我们在多个真实安全生成场景中验证了该框架,结果表明负向引导应在去噪过程早期应用,才能实现有效安全生成。
原文摘要 · Abstract (English)
Safety mechanisms for diffusion and flow models have recently been developed along two distinct paths. In robot planning, control barrier functions are employed to guide generative trajectories away from obstacles at every denoising step by explicitly imposing geometric constraints. In parallel, recent data-driven, negative guidance approaches have been shown to suppress harmful content and promote diversity in generated samples. However, they rely on heuristics without clearly stating when safety guidance is actually necessary. In this paper, we first introduce a unified probabilistic framework using a Maximum Mean Discrepancy (MMD) potential for image generation tasks that recasts both Shielded Diffusion and Safe Denoiser as instances of our energy-based negative guidance against unsafe data samples. Furthermore, we leverage control-barrier functions analysis to justify the existence of a critical time window in which negative guidance must be strong; outside of this window, the guidance should decay to zero to ensure safe and high-quality generation. We evaluate our unified framework on several realistic safe generation scenarios, confirming that negative guidance should be applied in the early stages of the denoising process for successful safe generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。