让大模型决策的奖励更精准,提升长任务表现。
Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization

- 用成功经验提炼可执行动作的专属监督信号
- 在三个基准上比GRPO提升10.6%,泛化更好
- 适合做复杂长流程任务的大模型训练
基于结果的强化学习为语言模型智能体提供验证反馈,但将轨迹级优势均匀分配给所有决策,导致长时程交互中信用分配粗糙。在线策略自蒸馏通过利用训练期独有的特权信息(PI)重新评估采样行为,实现更细粒度的监督。然而,细粒度监督不等于细粒度信用:PI引起的似然变化仅反映额外信息如何改变策略偏好,并不直接决定可执行动作应如何继承已验证的任务成果,从而产生监督-信用鸿沟。特权信号可能与当前交互状态无关,以词元粒度运行,与可执行决策不匹配,且缺乏强化学习所需的成果语义。本文提出TASPO,将特权监督转化为基于成果的动作信用。TASPO从验证成功的经验中构建适用于决策的特权信息,将PI引起的似然变化聚合至可执行动作层级,并将相对动作支持转换为正向、有界、均值保持的原始轨迹优势权重。因此,验证结果决定更新方向与平均尺度,而特权信息仅重新分配信用。在三个智能体基准上,TASPO相比GRPO提升10.6%,并展现出更好的未见任务泛化能力。进一步分析表明,TASPO降低了监督错配,且动作级分配稳定了策略优化过程。这些发现为社区提供了新的视角。
原文摘要 · Abstract (English)
Outcome-based reinforcement learning provides verified feedback for language-model agents, but assigns trajectory-level advantage uniformly to all decisions, yielding coarse credit over long-horizon interactions. On-policy self-distillation offers finer supervision by re-evaluating sampled behavior with privileged information (PI) available only during training. However, fine-grained supervision is not necessarily fine-grained credit: PI-induced likelihood changes describe how additional information alters policy preference, but do not directly determine how an executable action should inherit the verified task outcome. This creates a supervision-credit gap. Privileged signals may be irrelevant to the current interaction state, operate at a token granularity misaligned with executable decisions, and lack the outcome semantics required for reinforcement. We introduce TASPO, which converts privileged supervision into outcome-grounded action credit. TASPO constructs decision-applicable PI from verified successful experience, aggregates PI-induced likelihood shifts at the executable-action level, and converts relative action support into positive, bounded, mean-preserving weights on the original trajectory advantage. Thus, the verified outcome determines the update direction and average scale, while PI only redistributes credit across actions. Across three agentic benchmarks, TASPO improves over GRPO by 10.6\% and generalizes better to unseen tasks. Further analysis indicates that TASPO reduces supervision mismatch and that action-level assignment stabilizes the policy optimization process. These findings offer the community another interesting perspective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。