针对多轮智能体的反馈选择性提炼,提升长程奖励分配效果。
What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents

- 根据环境反馈选择关键动作进行强化学习更新
- 在ALFWorld和WebShop上分别达90.0%和80.1%成功率
- 适合需要精准反馈的复杂交互任务研究者
强化学习可从稀疏的任务奖励中训练大模型智能体,但长时程信用分配仍具挑战:单一成功或失败信号需分配至多个动作。现有方法依赖轨迹级奖励或代理信号,未能充分利用每步环境反馈。多轮智能体场景下,反馈可包括错误信息、页面变化、观测结果或参考轨迹。我们系统研究了五类反馈源与两种插入粒度,提出SERL——一种选择性环境重加权学习框架。SERL以任务奖励决定更新方向,环境反馈调整更新位置与强度,聚焦关键动作。在ALFWorld和WebShop上,SERL分别实现90.0%和80.1%的成功率,优于强基线的RL与蒸馏方法。分析表明,在有意义时刻使用具体且与动作相关的反馈,始终优于盲目使用更长或更丰富的上下文。
原文摘要 · Abstract (English)
Reinforcement learning can train LLM agents from sparse task rewards, but long-horizon credit assignment remains challenging: a single success-or-failure signal must be distributed across many actions. Existing methods rely on trajectory-level rewards or proxy signals, without fully leveraging per-step environmental feedback. Multi-turn agent settings are underexplored, where feedback can include error messages, page changes, observations, or reference trajectories. We systematically study five feedback sources and two insertion granularities and introduce SERL, a selective environment-reweighted learning framework. SERL uses the task reward to determine update direction, while environment feedback adjusts placement and magnitude, focusing on critical actions. On ALFWorld and WebShop, SERL achieves 90.0% and 80.1% success, outperforming strong RL and distillation baselines. Analysis shows that grounded, action-relevant feedback at meaningful points consistently outperforms indiscriminate use of longer or richer context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。