arXiv:2609.08404cs.LGcs.AI2026-09

通过增强环境反馈,让大模型在长任务中自进化更稳定高效

Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks

论文配图:Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks
图 1 · 摘自论文原文
  • 用反馈丰富的环境替代传统指令,引导模型自主探索
  • 在SciWorld和BFCL上提升性能,最高达18.7%的奖励增益
  • 适合研究长周期强化学习与自主智能体的学者

大型语言模型在静态推理中表现优异,但在长期任务中通过强化学习训练为自主智能体时,常因奖励稀疏而受阻。传统依赖监督微调的智能体预热方法受限于数据稀缺和探索能力不足。为此,我们提出环境侧适应的新范式——构建反馈增强环境(FEE)。通过试点研究,确立了在任务内探索与跨轮次演化后期,将环境从动作指导转向观察增强的反馈设计策略。大规模实验在SciWorld与BFCL基准上使用不同规模的Qwen3模型及GRPO、GSPO、DAPO等算法验证,结果表明FEE在所有设置下均显著优于标准环境。分析显示,使用FEE能:(1) 降低熵波动,稳定训练动态;(2) 在困难任务中促进主动状态空间探索;(3) 确保环境引导被内化为策略权重,而非仅作为推理阶段先验;(4) 识别组内反馈一致性是稳定优化的关键边界。

原文摘要 · Abstract (English)

Large Language Models demonstrate remarkable proficiency in static reasoning, yet training them as autonomous agents through Reinforcement Learning (RL) for long-horizon tasks is often hindered by severe reward sparsity. While conventional \textit{agent-side warming} up via supervised fine-tuning (SFT) can alleviate this, it is frequently limited by data scarcity and constrained exploration. To address this, we propose a paradigm shift to \textit{environment-side adaptation} by constructing \textbf{F}eedback-\textbf{E}nriched \textbf{E}nvironments (\textbf{FEEs}). Through a pilot study, we establish a feedback design strategy that reformulates environments by transitioning from action guidance to observation enrichment during the later stages of both intra-episode exploration and inter-episode evolution. Large-scale experiments on SciWorld and BFCL benchmarks using various Qwen3 model scales and RL algorithms such as GRPO, GSPO, and DAPO demonstrate that FEEs consistently yield performance improvements over standard settings. Furthermore, our analysis reveals that training with FEEs \textbf{(1)} stabilizes training dynamics by reducing entropy volatility, \textbf{(2)} facilitates proactive state-space exploration in difficult tasks, \textbf{(3) }ensures the internalization of environmental guidance into policy weights rather than acting as a mere inference-time prior, and \textbf{(4) }identifies intra-group feedback consistency as a critical boundary for stable optimization.

强化学习智能体反馈机制长程任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。