arXiv:2510.24012cs.LGcs.AI2025-10NeurIPS被引 5

无需训练即可让文生图模型生成更安全图像

Training-Free Safe Text Embedding Guidance for Text-to-Image Diffusion Models

  • 采样时动态调整文本嵌入,基于预期图像安全性判断
  • 在去裸露、去暴力等场景中显著降低有害内容产出
  • 适合关注生成安全性的研究者与应用开发者

文生图模型近年来在生成真实且语义一致的图像方面取得显著进展,主要得益于先进的扩散模型和大规模网络爬取数据集。然而,这些数据集常包含不当或偏见内容,导致恶意文本提示可能引发有害输出。我们提出训练无关的安全文本嵌入引导(STG),通过在采样过程中根据预期最终去噪图像的安全性函数调整文本嵌入,实现无需额外训练即可生成更安全的输出。理论上,STG使模型分布与安全约束对齐,从而在最小影响生成质量的前提下提升安全性。在去裸露、去暴力及艺术家风格移除等多种安全场景下的实验表明,STG在去除有害内容方面持续优于基于训练和无训练的基线方法,同时保留输入提示的核心语义意图。

原文摘要 · Abstract (English)

Text-to-image models have recently made significant advances in generating realistic and semantically coherent images, driven by advanced diffusion models and large-scale web-crawled datasets. However, these datasets often contain inappropriate or biased content, raising concerns about the generation of harmful outputs when provided with malicious text prompts. We propose Safe Text embedding Guidance (STG), a training-free approach to improve the safety of diffusion models by guiding the text embeddings during sampling. STG adjusts the text embeddings based on a safety function evaluated on the expected final denoised image, allowing the model to generate safer outputs without additional training. Theoretically, we show that STG aligns the underlying model distribution with safety constraints, thereby achieving safer outputs while minimally affecting generation quality. Experiments on various safety scenarios, including nudity, violence, and artist-style removal, show that STG consistently outperforms both training-based and training-free baselines in removing unsafe content while preserving the core semantic intent of input prompts. Our code is available at https://github.com/aailab-kaist/STG.

文生图扩散模型安全生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。