让训练环境自动适应智能体水平,提升学习效率与泛化能力。
Towards Adaptive Environment Generation for Training Embodied Agents
- 基于智能体表现反馈动态调整环境难度
- 生成更符合当前短板的挑战性场景
- 适合需要高效泛化训练的具身智能体研究
具身智能体在新环境中难以泛化,即使其结构与训练环境相似。现有环境生成多采用开环范式,未考虑智能体表现。尽管程序化生成可创造多样场景,但缺乏反馈则效率低下,生成环境可能过于简单,学习信号有限。为此,我们提出闭环环境生成的可行性方案,根据智能体当前能力自适应调整难度。系统采用可控环境表示,提取细粒度性能反馈(非仅成功/失败),并通过闭环机制将反馈转化为环境修改。该反馈驱动方法生成的训练环境能针对性地在智能体需改进的方向上增加挑战,实现更高效的训练,并提升对新场景的泛化能力。
原文摘要 · Abstract (English)
Embodied agents struggle to generalize to new environments, even when those environments share similar underlying structures to their training settings. Most current approaches to generating these training environments follow an open-loop paradigm, without considering the agent's current performance. While procedural generation methods can produce diverse scenes, diversity without feedback from the agent is inefficient. The generated environments may be trivially easy, providing limited learning signal. To address this, we present a proof-of-concept for closed-loop environment generation that adapts difficulty to the agent's current capabilities. Our system employs a controllable environment representation, extracts fine-grained performance feedback beyond binary success or failure, and implements a closed-loop adaptation mechanism that translates this feedback into environment modifications. This feedback-driven approach generates training environments that more challenging in the ways the agent needs to improve, enabling more efficient learning and better generalization to novel settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。