arXiv:2609.07303cs.LG2026-09

用分段匹配奖励让长任务语言模型强化学习更快更准

Long-Horizon Language Model Reinforcement Learning via Progressive Point Matching

  • 用分段进度匹配设计无偏密集奖励,提升长序列任务学习效率
  • 在高难度数学题上,稀疏奖励无法进步,分段奖励显著提升成功率
  • 只需一条参考轨迹即可实现,适合复杂长程推理任务

当前基于强化学习训练语言模型的方法严重依赖稀疏结果奖励。但当任务需要更长、更复杂的执行轨迹时,这类策略学习速度极慢。已有工作尝试通过奖励部分进展缓解问题,但简单做法常引入偏差并收敛至次优策略。本文提出一种简单且无偏的密集奖励机制——渐进式点匹配(progressive point matching),该方法在合成环境中通过分段级奖励部分进展,理论与实证均证明其能指数级提升长时序任务的学习效率。进一步展示如何仅用每项任务一条参考轨迹即可实际应用该方法。在极难的数学推理问题中,稀疏奖励无法产生任何进展,而分段级奖励在更大测试阶段词元预算下显著提升成功率或pass@k指标。

原文摘要 · Abstract (English)

Current paradigms for training language models via reinforcement learning rely heavily on sparse outcome rewards. However, as we pursue tasks that require longer and more complicated trajectories, such strategies result in slow learning. Prior work has attempted to address this problem by rewarding partial progress; however, naive formulations are often biased and converge to suboptimal policies. We show that a simple and unbiased dense reward formulation, which we term progressive point matching, scales exponentially more efficiently to long-horizon tasks by rewarding partial progress on a segment level, both theoretically and empirically via synthetic environments. We then show how progressive point matching can be practically instantiated using a single reference trajectory per task. On extremely hard math reasoning problems, sparse outcome rewards cannot make any progress, whereas segment-level rewards enable improvements at larger test-time token budgets when measured by success rate or pass@k.

强化学习长程推理奖励设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。