arXiv:2409.03256cs.CLcs.AI2024-09EMNLP被引 10

让智能体通过试错纠错,更适应真实环境。

E2CL: Exploration-based Error Correction Learning for Embodied Agents

论文配图:E2CL: Exploration-based Error Correction Learning for Embodied Agents
图 1 · 摘自论文原文
  • 利用探索中的错误和环境反馈进行自我修正
  • 在VirtualHome中优于基线方法,纠错能力更强
  • 适合研究具身智能与自主学习的学者

语言模型在知识利用和推理方面能力日益增强。然而,当作为具身智能体应用于环境中时,其内在知识常与环境知识不匹配,导致不可行动作。传统环境对齐方法如专家轨迹监督学习和强化学习,分别受限于知识覆盖不足和收敛效率低。受人类学习启发,我们提出探索式错误修正学习(E2CL),一种新框架,利用探索引发的错误和环境反馈,提升具身智能体的环境对齐能力。E2CL结合教师引导和无教师探索,收集环境反馈并修正错误行为。智能体学会提供反馈并自我修正,从而增强对目标环境的适应性。在VirtualHome环境的大量实验表明,E2CL训练的智能体优于基线方法,展现出更强的自我修正能力。

原文摘要 · Abstract (English)

Language models are exhibiting increasing capability in knowledge utilization and reasoning. However, when applied as agents in embodied environments, they often suffer from misalignment between their intrinsic knowledge and environmental knowledge, leading to infeasible actions. Traditional environment alignment methods, such as supervised learning on expert trajectories and reinforcement learning, encounter limitations in covering environmental knowledge and achieving efficient convergence, respectively. Inspired by human learning, we propose Exploration-based Error Correction Learning (E2CL), a novel framework that leverages exploration-induced errors and environmental feedback to enhance environment alignment for embodied agents. E2CL incorporates teacher-guided and teacher-free explorations to gather environmental feedback and correct erroneous actions. The agent learns to provide feedback and self-correct, thereby enhancing its adaptability to target environments. Extensive experiments in the VirtualHome environment demonstrate that E2CL-trained agents outperform those trained by baseline methods and exhibit superior self-correction capabilities.

具身智能自纠正探索学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。