arXiv:2602.23008cs.LGcs.AI2026-02中稿 · ICLR被引 16

让大模型在新环境中更会探索,且不依赖记忆也能稳定表现。

Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization

  • 用记忆辅助探索,结合在线与离线策略更新提升学习效率。
  • 在ScienceWorld和WebShop上分别比GRPO提升128.6%和11.3%。
  • 新任务只需少量尝试即可适应,适合构建通用智能体。

探索仍是大语言模型智能体通过强化学习训练时的关键瓶颈。尽管先前方法利用预训练知识,但在需要发现新状态的环境中表现不佳。我们提出探索性记忆增强的混合在线与离线策略优化框架(EMPO²),利用记忆促进探索,并结合在线与离线更新机制,使大模型在有记忆时表现优异,无记忆时仍具鲁棒性。在ScienceWorld和WebShop上,EMPO²分别较GRPO提升128.6%和11.3%。此外,在分布外测试中,EMPO²展现出更强适应能力,仅需少量试错即可完成新任务,无需参数更新。结果表明,该框架为构建更具探索性和泛化能力的大模型智能体提供了新路径。

原文摘要 · Abstract (English)

Exploration remains the key bottleneck for large language model agents trained with reinforcement learning. While prior methods exploit pretrained knowledge, they fail in environments requiring the discovery of novel states. We propose Exploratory Memory-Augmented On- and Off-Policy Optimization (EMPO$^2$), a hybrid RL framework that leverages memory for exploration and combines on- and off-policy updates to make LLMs perform well with memory while also ensuring robustness without it. On ScienceWorld and WebShop, EMPO$^2$ achieves 128.6% and 11.3% improvements over GRPO, respectively. Moreover, in out-of-distribution tests, EMPO$^2$ demonstrates superior adaptability to new tasks, requiring only a few trials with memory and no parameter updates. These results highlight EMPO$^2$ as a promising framework for building more exploratory and generalizable LLM-based agents.

强化学习大模型智能体探索机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。