arXiv:2510.21473cs.CL2025-10NeurIPS被引 4

通过多奖励优化提升扩散语言模型的推理能力

MRO: Enhancing Reasoning in Diffusion Language Models via Multi-Reward Optimization

  • 设计多奖励机制,引导模型在去噪过程中关注词元关联性
  • 在少步去噪下仍保持高推理性能,且采样速度显著提升
  • 适合关注高效推理与生成质量平衡的研究者

近期扩散语言模型(DLMs)为传统自回归大模型提供了有前景的替代方案,但在推理性能上仍落后于大语言模型(LLMs),尤其在去噪步数较少时。分析表明,问题主要源于各去噪步骤中掩码词元的独立生成,未能捕捉词元间的相关性。本文定义了序列内相关性和序列间相关性,并证明增强这些相关性可提升推理表现。为此,提出多奖励优化(MRO)方法,通过测试时缩放、拒绝采样和强化学习,在去噪过程中直接优化词元相关性,引入分组步长与重要性采样策略以降低奖励方差并提高采样效率。大量实验表明,MRO不仅显著提升推理性能,还在保持高表现的同时实现显著采样加速。

原文摘要 · Abstract (English)

Recent advances in diffusion language models (DLMs) have presented a promising alternative to traditional autoregressive large language models (LLMs). However, DLMs still lag behind LLMs in reasoning performance, especially as the number of denoising steps decreases. Our analysis reveals that this shortcoming arises primarily from the independent generation of masked tokens across denoising steps, which fails to capture the token correlation. In this paper, we define two types of token correlation: intra-sequence correlation and inter-sequence correlation, and demonstrate that enhancing these correlations improves reasoning performance. To this end, we propose a Multi-Reward Optimization (MRO) approach, which encourages DLMs to consider the token correlation during the denoising process. More specifically, our MRO approach leverages test-time scaling, reject sampling, and reinforcement learning to directly optimize the token correlation with multiple elaborate rewards. Additionally, we introduce group step and importance sampling strategies to mitigate reward variance and enhance sampling efficiency. Through extensive experiments, we demonstrate that MRO not only improves reasoning performance but also achieves significant sampling speedups while maintaining high performance on reasoning benchmarks.

扩散模型推理增强多奖励优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。