让AI更会规划:通过精细奖励提升多轮任务效率
PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMs

- 用粗到细的奖励机制,区分成功路径的效率差异
- 在多个复杂任务上平均提升27.2%表现,优于现有方法
- 适合需要高效规划能力的对话系统与智能体研究
群体相对策略优化已成为训练多轮交互式智能体大模型的关键范式。然而,现有方法难以区分不同成功轨迹间的交互效率差异,导致迂回成功的路径获得相同奖励,引发优势坍缩和性能瓶颈。为此,我们提出一种名为计划感知策略优化(PlanPO)的简单而有效的强化学习方法,旨在学习超越特定任务的通用规划能力。PlanPO引入粗粒度到细粒度的优势信号,捕捉同一任务下成功轨迹在整体长度和每轮响应长度上的相对差异。在群体相对优化框架中,使智能体能从高质量采样轨迹中主动学习涵盖交互规划与文本生成的通用、审慎行为,避免退化为单纯的长度最小化。实验表明,PlanPO在挑战性多轮基准测试ALFWorld、WebShop和SciWorld上平均比GRPO提升27.2%,超越近期强大基线,且训练开销几乎不变。
原文摘要 · Abstract (English)
Group-relative policy optimization has emerged as a key paradigm for training agentic large language models (LLMs) on multi-turn interactive tasks. However, most existing variants fail to distinguish advantages among successful trajectories even when these trajectories differ substantially in their interaction efficiency. For instance, circuitous successes are often assigned the identical outcome reward, causing advantage collapse and severe performance bottlenecks. To this end, we propose Group Planning-aware Policy Optimization (PlanPO), a simple yet effective RL method for learning generalizable planning abilities beyond task-specific high-quality behavior patterns. Specifically, PlanPO introduces coarse-to-fine advantage signals, which capture the relative differences in trajectory-level lengths and turn-level response lengths conditioned on successful trajectories sampled for the same task. Within the group-relative optimization structure, this enables agents to actively learn generalizable and deliberate behaviors spanning interaction planning and textual generation from high-quality rollouts, without degenerating into vanilla length minimization. Experimentally, PlanPO improves over GRPO by 27.2\% on average across the challenging multi-turn benchmarks ALFWorld, WebShop, and SciWorld, outperforming recent powerful baselines while incurring negligible additional training cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。