通过过程奖励提升扩散语言模型的推理能力,让生成更稳定可解释。
Advancing Reasoning in Diffusion Language Models with Denoising Process Rewards
- 设计基于去噪轨迹的中间过程奖励,引导推理走向正确答案。
- 在多个复杂推理基准上显著提升稳定性与任务性能。
- 方法高效可扩展,适合追求可解释推理的AI研究者使用。
基于扩散的大型语言模型为文本生成提供了非自回归替代方案,但实现复杂推理仍具挑战。强化学习近年成为提升其性能的有效后训练策略,但现有方法主要依赖结果奖励,无法直接监督去噪过程,常导致推理结构松散、难以解释且对最终预测支持不一致。为此,我们提出“去噪过程奖励”,一种定义在扩散语言模型去噪轨迹上的过程级强化信号。该奖励通过估计中间去噪阶段对最终任务结果的贡献,鼓励模型选择持续引导生成向正确预测的推理路径。我们进一步提出一种高效的随机估计算法,复用标准训练轨迹,在大规模下实现过程级监督。在多个高难度推理基准上的实验表明,该方法显著提升了推理稳定性、可解释性与整体任务表现。
原文摘要 · Abstract (English)
Diffusion-based large language models offer a non-autoregressive alternative for text generation, but enabling them to perform complex reasoning remains challenging. Reinforcement learning has recently emerged as an effective post-training strategy for improving their performance; however, existing methods rely primarily on outcome-based rewards, which provide no direct supervision over the denoising process and often result in poorly structured reasoning that is difficult to interpret and inconsistently supports the final prediction. To address this limitation, we introduce \emph{denoising process reward}, a process-level reinforcement signal defined over the denoising trajectory of diffusion language models. This reward is obtained by estimating the contribution of intermediate denoising intervals to the final task outcome, encouraging the model to favor reasoning trajectories that consistently guide generation toward correct predictions. We further propose an efficient stochastic estimator that reuses standard training rollouts, enabling practical process-level supervision at scale. Experiments on challenging reasoning benchmarks demonstrate that our approach yields consistent improvements in reasoning stability, interpretability, and overall task performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。