通过局部证据识别关键里程碑,提升长周期智能体的信用分配效果。
MileGPO: Milestone Inference with Local Evidence for Graph-Based Policy Optimization of Long-Horizon LLM Agents

- 基于成功与失败轨迹发现候选里程碑,结合可信度加权筛选可靠节点。
- 在ALFWorld和WebShop上达到当前最优性能,且分布内到分布外差距小。
- 无需额外模型或环境交互,适合长周期决策类强化学习任务研究者。
长周期智能体强化学习中的信用分配极具挑战,因监督信号通常仅来自最终奖励。现有方法通过步骤分组或基于图的优势估计将轨迹级信号细化为步骤级信用,但常忽略有意义的中间里程碑。本文提出MileGPO(基于局部证据的里程碑推断图基策略优化),通过三个设计从有政策回放中推导过程级信用:里程碑发现(在成功轨迹上识别候选里程碑,在失败轨迹上识别重复陷阱);可靠性校准着色(根据结果置信度加权候选节点,增强可靠里程碑和陷阱,抑制不确定项);进度对比校准(检验候选节点是否反映局部进展,并验证其转移路径是否优于同状态下的其他选择)。MileGPO无需辅助模型或额外环境交互。在ALFWorld和WebShop上的实验表明其性能达到当前最优,且在ALFWorld上分布内到分布外的性能差距较小。消融实验与信用诊断显示,可靠性加权、局部进展检测及同状态分支证据共同补充了里程碑发现,有效解决了中间信用模糊问题。
原文摘要 · Abstract (English)
Credit assignment is challenging in long-horizon agentic reinforcement learning, where supervision often comes only from final rewards. Existing methods refine trajectory-level signals into step-level credits through step grouping or graph-based advantage estimation, but can overlook meaningful intermediate milestones. We propose MileGPO (Milestone Inference with Local Evidence for Graph-Based Policy Optimization), which derives process-level credit from grouped on-policy rollouts through three designs. Milestone Discovery identifies candidate milestones on successful rollouts and recurring traps on failed ones. Reliability-Calibrated Shaping (RCS) weights these candidates by outcome-based confidence, strengthening reliable milestones and traps while down-weighting uncertain ones. Progress-Contrastive Calibration (PCC) further tests whether a candidate reflects local progress and whether its incoming transition outperforms observed alternatives from the same state. MileGPO requires neither auxiliary models nor additional environment interaction. Experiments on ALFWorld and WebShop show state-of-the-art performance and a small in-distribution to out-of-distribution gap on ALFWorld. Ablations and credit diagnostics indicate that reliability weighting, local progress, and same-state branch evidence complement milestone discovery and resolve ambiguous intermediate credit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。