arXiv:2605.08116cs.LGcs.AI2026-05

为文本扩散模型设计安全引导去噪器,生成更安全文本。

The Safety-Aware Denoiser for Text Diffusion Models

论文配图:The Safety-Aware Denoiser for Text Diffusion Models
图 1 · 摘自论文原文
  • 在去噪过程中动态引导文本走向安全区域。
  • 显著降低有害内容生成率,保持生成质量。
  • 无需重训练模型,适合实际部署使用。

近期的文本扩散模型为非自回归生成提供了有前景的替代方案,但其安全性控制仍待深入探索。现有安全方法主要针对自回归模型,通常依赖事后过滤或推理时干预,难以有效应对文本扩散模型的安全风险。本文提出安全感知去噪器(SAD),一种在文本扩散模型中引入安全引导的框架。SAD通过修改迭代去噪过程,使最终生成的文本样本被引导至文本空间中的可证明安全区域。该推理阶段方法可将安全约束融入去噪器,避免对底层扩散模型进行昂贵的重新训练,并实现灵活、轻量的安全引导。我们在危险分类、记忆泄露和越狱攻击等维度评估了SAD的安全性。实验结果表明,SAD显著减少了不安全生成,同时保持了生成质量与流畅性,优于现有方法。这些结果证明,在去噪过程中加入安全引导是一种有效且可扩展的文本扩散模型安全机制。

原文摘要 · Abstract (English)

Recent work on text diffusion models offers a promising alternative to autoregressive generation, but controlling their safety remains underexplored. Existing safety approaches are geared toward autoregressive models and typically rely on post-hoc filtering or inference-time interventions. These are inadequate for effectively addressing safety risks in text diffusion models. We propose the Safety-Aware Denoiser (SAD), a safety-guidance framework in text diffusion models. The SAD modifies the iterative denoising process such that the text sample at the final denoising step is steered toward provably safe regions of the text space. This inference-time method can integrate safety constraints into the denoiser, avoiding computationally expensive retraining of the underlying diffusion model and enabling flexible, lightweight safety guidance. We evaluate the safety of the generated text using the SAD, with respect to hazard taxonomy, memorization, and jailbreak. Experimental results show that SAD substantially reduces unsafe generations while preserving generation quality and fluency, outperforming existing methods. These results demonstrate that our safety guidance during denoising provides an effective and scalable mechanism for enforcing safety in text diffusion models.

文本生成扩散模型安全引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。