提出RSPO方法,让扩散语言模型推理更稳定高效。
Relative Score Policy Optimization for Diffusion Language Models
- 用奖励优势校准噪声得分,改进策略优化
- 在规划任务上提升显著,数学推理表现优异
- 适合追求高稳定性推理的模型优化者
扩散大语言模型(dLLMs)为并行高效文本生成提供了前景,但其推理能力的提升依赖有效的后训练。强化学习结合可验证奖励(RLVR)是自然选择,但其应用受限于难以获得可处理的序列级对数比,这是标准策略优化的核心。现有方法被迫使用高方差的ELBO近似,导致高验证奖励放大错误得分估计,进而破坏强化学习训练稳定性。为此,我们提出相对得分策略优化(RSPO),一种利用可验证奖励校准dLLMs中噪声似然估计的简单RLVR方法。核心思想是:奖励优势不仅可作为更新方向,也可作为当前与参考策略间相对对数比的目标值。因此,RSPO通过比较实际奖励优势与奖励隐含目标相对对数比,基于当前估计与目标间的差距而非原始优势来更新策略。在数学推理和规划基准测试上的实验表明,RSPO在规划任务上取得显著提升,且在数学推理任务上表现具有竞争力。
原文摘要 · Abstract (English)
Diffusion large language models (dLLMs) offer a promising route to parallel and efficient text generation, but improving their reasoning ability requires effective post-training. Reinforcement learning with verifiable rewards (RLVR) is a natural choice for this purpose, yet its application to dLLMs is hindered by the absence of tractable sequence-level log-ratios, which are central to standard policy optimization. The lack of tractable sequence-level log-ratios forces existing methods to rely on high-variance ELBO-based approximations, where high verifier rewards can amplify inaccurate score estimates and destabilize RL training. To overcome this issue, we propose \textbf{R}elative \textbf{S}core \textbf{P}olicy \textbf{O}ptimization (RSPO), a simple RLVR method that uses verifiable rewards to calibrate noisy likelihood estimates in dLLMs. The core of our algorithm relies on a key observation: a reward advantage can be interpreted not only as an update direction, but also as a target for the relative log-ratio between the current and reference policies. Accordingly, RSPO calibrates this noisy relative log-ratio estimate by comparing its reward advantage with the reward-implied target relative log-ratio, updating the policy according to the gap between the current estimate and the target rather than the raw advantage alone. Experiments on mathematical reasoning and planning benchmarks show that RSPO yields especially strong gains on planning tasks and competitive mathematical-reasoning performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。