arXiv:2608.08255cs.LGcs.CL2026-08

利用环境反馈实现多时尺度奖励分配,提升长程智能体强化学习效果

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning

论文配图:Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning
图 1 · 摘自论文原文
  • 基于环境反馈提取短期与中期过程信号,补充长期奖励
  • 在ALFWorld和WebShop上任务成功率和质量显著优于基线
  • 适合需要长序列决策的智能体学习场景

智能体强化学习在真实环境中常面临奖励延迟且稀疏的问题。现有信用分配方法忽略交互过程中自然生成的丰富过程信息(如交互历史)。本文提出环境反馈信用分配(EFCA),一种多时尺度信用分配方法,通过直接从环境反馈中提取短期反馈信号(当前动作的即时影响)和中短期状态历史信号(近期无效模式识别),并以回报重加权机制融合,为中间决策提供更细粒度监督。在ALFWorld和WebShop上的实验表明,EFCA在任务成功率和任务质量上均持续优于强基线,验证了基于环境的多时尺度信用分配在长时程智能体强化学习中的有效性。

原文摘要 · Abstract (English)

Agentic reinforcement learning (RL) often suffers from delayed and sparse rewards in real-world environments. A promising solution to this challenge is credit assignment, which aims to decompose trajectory-level rewards and provide more fine-grained supervision for intermediate decisions. However, existing credit assignment approaches ignore the rich process information naturally generated during environment interaction, e.g., interaction history. We argue that such information provides valuable supervision for identifying the contribution of individual actions. To this end, we propose Environmental Feedback-based Credit Assignment (EFCA), a multi-timescale credit assignment approach for long-horizon agentic RL. EFCA complements the long-term outcome signal with two environment-grounded process signals: a short-term feedback signal that captures the immediate effect of the current action and a medium-term state-history signal that identifies ineffective patterns from recent interactions. Both signals are directly extracted from environment feedback and integrated through a return reweighting mechanism. Experiments on ALFWorld and WebShop demonstrate that EFCA consistently improves both task success and task quality over strong baselines, highlighting the effectiveness of environment-grounded multi-timescale credit assignment for long-horizon agentic RL.

强化学习信用分配智能体多时尺度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。