arXiv:2509.23155cs.RO2025-09被引 2

让机器人用语言反思错误,自动改进操作能力。

LaGEA: Language Guided Embodied Agents for Robotic Manipulation

  • 用视觉语言模型生成带时间定位的错误分析语言反馈
  • 在元世界和抓取任务上成功提升17%的完成率
  • 适合需要自纠错能力的智能机器人研究者

机器人操作得益于能描述目标的基础模型,但现有智能体仍缺乏从自身错误中学习的系统方法。本文探讨自然语言能否作为反馈信号,帮助具身智能体诊断失败原因并调整策略。提出LaGEA框架,将视觉语言模型(VLM)生成的、具有结构约束的片段式反思转化为强化学习中的时序锚定指导。该框架对每次尝试进行简洁语言总结,定位轨迹中的关键决策时刻,将反馈与视觉状态对齐于共享表征空间,并将目标进展与反馈一致性转化为有界、分步的奖励信号,其影响由自适应的故障感知系数调节。该设计在探索初期提供密集引导,随能力提升逐渐减弱。在Meta-World MT10和Robotic Fetch具身操作基准测试中,LaGEA在随机目标下平均成功率提升9.0%,固定目标下提升5.3%,抓取任务提升17%,且收敛更快。结果验证了假设:当语言结构化并时空锚定时,可有效赋能机器人自我反思与决策优化。

原文摘要 · Abstract (English)

Robotic manipulation benefits from foundation models that describe goals, but today's agents still lack a principled way to learn from their own mistakes. We ask whether natural language can serve as feedback, an error-reasoning signal that helps embodied agents diagnose what went wrong and correct course. We introduce LaGEA (Language Guided Embodied Agents), a framework that turns episodic, schema-constrained reflections from a vision language model (VLM) into temporally grounded guidance for reinforcement learning. LaGEA summarizes each attempt in concise language, localizes the decisive moments in the trajectory, aligns feedback with visual state in a shared representation, and converts goal progress and feedback agreement into bounded, step-wise shaping rewards whose influence is modulated by an adaptive, failure-aware coefficient. This design yields dense signals early when exploration needs direction and gracefully recedes as competence grows. On the Meta-World MT10 and Robotic Fetch embodied manipulation benchmark, LaGEA improves average success over the state-of-the-art (SOTA) methods by 9.0% on random goals, 5.3% on fixed goals, and 17% on fetch tasks, while converging faster. These results support our hypothesis: language, when structured and grounded in time, is an effective mechanism for teaching robots to self-reflect on mistakes and make better choices.

具身智能语言反馈机器人操作自纠错

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。