用事后信息优化多轮决策奖励,提升长任务中的信用分配可靠性。
HISR: Hindsight Information Modulated Segmental Process Rewards For Multi-turn Agentic Reinforcement Learning

- 通过分段过程奖励机制,按子目标分配奖励,避免过度细化到每一步。
- 引入事后模型计算动作重要性比率,突出关键决策片段。
- 在三个公开基准上验证有效,适合复杂长序列决策任务研究者。
尽管大语言模型在多个领域表现优异,但在复杂的长时序代理决策任务中仍存在性能瓶颈。现有方法多聚焦于设计有效的奖励模型(RMs)以通过多轮强化学习提升性能,但面临稀疏结果奖励传播延迟以及过于细粒度、不聚焦的回合级过程奖励导致信用分配不可靠的问题。本文提出HISR(Hindsight Information Modulated Segmental Process Rewards),利用事后信息调节分段过程奖励,使奖励更贴近子目标,并强化重要段落以提升信用分配可靠性。具体而言,提出一种分段级过程奖励模型,为任务中的每个子目标分配奖励,避免对单个回合过度精细化;设计事后模型,基于轨迹最终结果反映特定动作执行后的偏好;通过比较事后模型与策略模型的序列似然比,衡量动作重要性,并据此聚合分段重要性得分,进而调制分段过程奖励。在三个公开基准上的实验结果证明了该方法的有效性。
原文摘要 · Abstract (English)
While large language models excel in diverse domains, their performance on complex longhorizon agentic decision-making tasks remains limited. Most existing methods concentrate on designing effective reward models (RMs) to advance performance via multi-turn reinforcement learning. However, they suffer from delayed propagation in sparse outcome rewards and unreliable credit assignment with potentially overly fine-grained and unfocused turnlevel process rewards. In this paper, we propose (HISR) exploiting Hindsight Information to modulate Segmental process Rewards, which closely aligns rewards with sub-goals and underscores significant segments to enhance the reliability of credit assignment. Specifically, a segment-level process RM is presented to assign rewards for each sub-goal in the task, avoiding excessively granular allocation to turns. To emphasize significant segments in the trajectory, a hindsight model is devised to reflect the preference of performing a certain action after knowing the trajectory outcome. With this characteristic, we design the ratios of sequence likelihoods between hindsight and policy model to measure action importance. The ratios are subsequently employed to aggregate segment importance scores, which in turn modulate segmental process rewards, enhancing credit assignment reliability. Extensive experimental results on three publicly benchmarks demonstrate the validity of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。