arXiv:2604.02265cs.CV2026-04

无需训练,用预训练模型实时引导生成安全图像。

Modular Energy Steering for Safe Text-to-Image Generation with Foundation Models

  • 在推理阶段通过梯度反馈调节生成过程,不修改原模型。
  • 在NSFW对抗测试中表现优于现有方法,且保持高质量生成。
  • 适用于扩散模型与流匹配模型,适合需要安全控制的场景。

控制文本到图像生成模型的行为对安全部署至关重要。现有安全方法多依赖微调或定制数据集,可能降低生成质量或限制可扩展性。我们提出一种推理时的可控框架,利用冻结的预训练视觉-语言基础模型的梯度反馈,指导生成过程而不修改底层生成器。关键观察是,这些基础模型编码了丰富的语义表示,可作为现成的监督信号。通过在每一步采样中注入干净的潜在估计反馈,将安全控制建模为基于能量的采样问题。该设计实现模块化、免训练的安全控制,兼容扩散模型与流匹配模型,并能泛化至多样视觉概念。实验表明,在对抗性NSFW测试中达到业界最优鲁棒性,同时实现多目标可控生成,且在无目标提示下仍保持高生成质量。本框架为利用基础模型作为语义能量估计算子提供了原则性方案,实现可靠、可扩展的文本到图像生成安全控制。

原文摘要 · Abstract (English)

Controlling the behavior of text-to-image generative models is critical for safe and practical deployment. Existing safety approaches typically rely on model fine-tuning or curated datasets, which can degrade generation quality or limit scalability. We propose an inference-time steering framework that leverages gradient feedback from frozen pretrained foundation models to guide the generation process without modifying the underlying generator. Our key observation is that vision-language foundation models encode rich semantic representations that can be repurposed as off-the-shelf supervisory signals during generation. By injecting such feedback through clean latent estimates at each sampling step, our method formulates safety steering as an energy-based sampling problem. This design enables modular, training-free safety control that is compatible with both diffusion and flow-matching models and can generalize across diverse visual concepts. Experiments demonstrate state-of-the-art robustness against NSFW red-teaming benchmarks and effective multi-target steering, while preserving high generation quality on benign non-targeted prompts. Our framework provides a principled approach for utilizing foundation models as semantic energy estimators, enabling reliable and scalable safety control for text-to-image generation.

文本生成图像安全扩散模型能量引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。