arXiv:2605.29626cs.CLcs.AI2026-05被引 1

无需训练即可控制扩散语言模型的生成风格与安全属性。

DLM-SWAI: Steering Diffusion Language Models Before They Unmask

论文配图:DLM-SWAI: Steering Diffusion Language Models Before They Unmask
图 1 · 摘自论文原文
  • 通过预计算的分词级风格分数,在去噪每一步调整词分布。
  • 在风格与安全性任务上有效,且保持生成质量,计算开销极低。
  • 可调节控制强度与流畅性的权衡,适合需要灵活控制的场景。

在不重新训练的情况下控制语言模型生成内容的文本特性对实际部署至关重要,而推理阶段方法尤其吸引人,因其无需重训练即可实现可控生成。近期研究指出,扩散语言模型(DLMs)作为一种新兴生成范式,具有独特的解码特性。然而,现有大多数控制方法依赖辅助模型或专为自回归逐词生成设计,难以适用于通过迭代去噪部分掩码序列生成文本的扩散语言模型。为此,我们提出 DLM-SWAI,一种无需训练的控制方法,通过预计算的分词级风格分数,在每个去噪步骤中偏置词分布。在风格与安全控制任务上的实验表明,DLM-SWAI 能有效引导扩散语言模型生成,同时保持生成质量,并仅需极少计算开销。消融实验揭示了控制强度与流畅性间的可调控权衡,分析还发现类别级可控制性与分词级属性提示强度相关。

原文摘要 · Abstract (English)

Steering language model generation toward desired textual properties is essential for practical deployment, and inference-time methods are particularly appealing because they enable controllable generation without retraining. Recent work has also highlighted diffusion language models as an emerging generation paradigm with distinct decoding properties. However, most existing steering approaches either rely on auxiliary models or are designed for autoregressive next-token decoding, making them difficult to apply to diffusion language models DLMs, which generate text through iterative denoising of partially masked sequences. Therefore, we propose DLM-SWAI, a simple training-free steering method that biases the token distribution at each denoising step using pre-computed token-level style scores. Experiments on style and safety control tasks show that DLM-SWAI effectively steers diffusion language models while preserving generation quality and requiring minimal computational overhead. Ablations further reveal a controllable trade-off between steering strength and fluency, and our analysis links class-wise steerability to the strength of token-level attribute cues.

扩散语言模型风格控制生成安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。