让智能体通过与环境互动学习,提升决策能力。
Environment Scaling for Interactive Agentic Experience Collection: A Survey
- 以生成-执行-反馈循环为核心,让环境主动挑战智能体。
- 环境需更复杂真实,才能提供有效学习数据。
- 适合研究智能体训练与交互式体验构建的学者。
基于大模型的智能体可在多个领域自主完成复杂任务。然而,要培养适应性行为和长期决策能力,仅靠静态的人类知识数据集训练已不足。这类数据集成本高,缺乏动态性和真实性。越来越多研究认为,智能体应直接与环境交互,通过强化学习积累经验。本文将此过程形式化为生成-执行-反馈(GEF)循环:环境生成任务挑战智能体,执行中返回观测结果,并在轨迹完成后提供评估反馈。在此范式下,环境成为经验数据的关键生产者,亟需向更高复杂度、真实感和交互性扩展。本综述从环境中心视角系统梳理代表性方法,按GEF循环的三个阶段——任务生成、任务执行与反馈进行组织,分析实现框架、挑战与应用,整合零散进展,并指明未来研究方向。
原文摘要 · Abstract (English)
LLM-based agents can autonomously accomplish complex tasks across various domains. However, to further cultivate capabilities such as adaptive behavior and long-term decision-making, training on static datasets built from human-level knowledge is insufficient. These datasets are costly to construct and lack both dynamism and realism. A growing consensus is that agents should instead interact directly with environments and learn from experience through reinforcement learning. We formalize this iterative process as the Generation-Execution-Feedback (GEF) loop, where environments generate tasks to challenge agents, return observations in response to agents' actions during task execution, and provide evaluative feedback on rollouts for subsequent learning. Under this paradigm, environments function as indispensable producers of experiential data, highlighting the need to scale them toward greater complexity, realism, and interactivity. In this survey, we systematically review representative methods for environment scaling from a pioneering environment-centric perspective and organize them along the stages of the GEF loop, namely task generation, task execution, and feedback. We further analyze implementation frameworks, challenges, and applications, consolidating fragmented advances and outlining future research directions for agent intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。