用大模型理解人类指令,自动修正机器人动作计划
A Human-in-the-loop Approach to Robot Action Replanning through LLM Common-Sense Reasoning
- 结合视觉视频与自然语言指令,用大模型补全常识推理
- 仅需一次演示即可纠正视觉误判,失败率降低40%以上
- 适合非专家用户快速调试机器人任务,提升系统鲁棒性
为推动机器人普及,需为非专家提供易用的编程工具。观测学习可通过实际操作示范实现技能传递,但仅依赖视觉输入在可扩展性和故障应对上效率较低,尤其当仅基于单次示范时。本文提出一种人机协同方法,对基于单个RGB视频自动生成的机器人执行计划,通过自然语言输入至大语言模型(LLM)进行增强。用户可指定目标或关键任务要素,利用LLM的常识推理能力,调整视觉生成的计划以预防潜在失败,并根据指令动态适应。实验表明,该框架直观有效,能纠正视觉推导错误且无需额外示范;交互式计划优化与幻觉修正显著提升了系统鲁棒性。
原文摘要 · Abstract (English)
To facilitate the wider adoption of robotics, accessible programming tools are required for non-experts. Observational learning enables intuitive human skills transfer through hands-on demonstrations, but relying solely on visual input can be inefficient in terms of scalability and failure mitigation, especially when based on a single demonstration. This paper presents a human-in-the-loop method for enhancing the robot execution plan, automatically generated based on a single RGB video, with natural language input to a Large Language Model (LLM). By including user-specified goals or critical task aspects and exploiting the LLM common-sense reasoning, the system adjusts the vision-based plan to prevent potential failures and adapts it based on the received instructions. Experiments demonstrated the framework intuitiveness and effectiveness in correcting vision-derived errors and adapting plans without requiring additional demonstrations. Moreover, interactive plan refinement and hallucination corrections promoted system robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。