arXiv:2509.18597cs.RO2025-09被引 4

让机器人通过人机协作持续学习长程操作技能,纠正错误并固化为可复用知识。

Growing with Your Embodied Agent: A Human-in-the-Loop Lifelong Code Generation Framework for Long-Horizon Manipulation Skills

  • 引入人类反馈,将纠错转化为可复用技能,结合外部记忆与动态检索增强生成。
  • 在多个环境和真实场景中实现0.93成功率,纠错轮次效率提升42%。
  • 适合需要长期规划与人机协同的复杂机器人任务研究者使用。

基于大语言模型(LLM)的代码生成在机器人操作中展现出潜力,可直接将人类指令转为可执行代码,但现有方法存在噪声大、受限于固定操作基元、上下文窗口有限,且难以处理长程任务。尽管闭环反馈已被探索,但修正知识常以不恰当格式存储,影响泛化能力并导致灾难性遗忘,凸显了可复用技能学习的必要性。此外,仅依赖LLM指导的方法在极长程任务中常因其在机器人领域推理能力不足而失败,而此类问题对人类而言通常易识别。为此,我们提出一种人机协同的终身代码生成框架,将修正信息编码为可复用技能,依托外部记忆与带提示机制的检索增强生成实现动态重用。在Ravens、Franka Kitchen、MetaWorld及真实场景中的实验表明,该框架成功率达0.93(比基线最高提升27%),纠错轮次效率提升42%,能稳健解决如“建房子”等需规划超过20个基元的极端长程任务。

原文摘要 · Abstract (English)

Large language models (LLMs)-based code generation for robotic manipulation has recently shown promise by directly translating human instructions into executable code, but existing methods remain noisy, constrained by fixed primitives and limited context windows, and struggle with long-horizon tasks. While closed-loop feedback has been explored, corrected knowledge is often stored in improper formats, restricting generalization and causing catastrophic forgetting, which highlights the need for learning reusable skills. Moreover, approaches that rely solely on LLM guidance frequently fail in extremely long-horizon scenarios due to LLMs' limited reasoning capability in the robotic domain, where such issues are often straightforward for humans to identify. To address these challenges, we propose a human-in-the-loop framework that encodes corrections into reusable skills, supported by external memory and Retrieval-Augmented Generation with a hint mechanism for dynamic reuse. Experiments on Ravens, Franka Kitchen, and MetaWorld, as well as real-world settings, show that our framework achieves a 0.93 success rate (up to 27% higher than baselines) and a 42% efficiency improvement in correction rounds. It can robustly solve extremely long-horizon tasks such as "build a house", which requires planning over 20 primitives.

人机协作长程任务技能学习机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。