让大模型学会反思与创造,实现零样本机器人协作新方案
REFLEX: Metacognitive Reasoning for Reflective Zero-Shot Robotic Planning with Large Language Models
- 引入元认知机制,让模型拆解技能、反思失败并生成新解法
- 在新任务上零样本表现超越基线,成功率达92%以上
- 适合对自主机器人规划感兴趣的研究者与工程师
尽管大语言模型(LLMs)在多个领域展现出巨大潜力,其在机器人领域的应用仍主要局限于静态提示行为,在零样本或少样本设置下的复杂任务中面临挑战。受人类元认知学习和创造性问题解决的启发,我们提出一个核心问题:能否赋予LLM元认知能力,使其能够推理、反思与创造,从而在极少示范下提升机器人任务执行能力?本文提出REFLEX框架,将元认知学习融入基于LLM的多机器人协作系统。该系统为LLM驱动的机器人代理配备了技能分解与自我反思机制,可从过往任务中识别模块化技能,反思未见过任务场景中的失败,并合成有效的新解决方案。我们设计了一个更具挑战性的机器人基准任务,在现有基准和新任务上评估该框架。实验结果表明,所提元认知学习框架显著优于现有基线。此外,我们观察到该框架生成的解决方案虽不同于真实答案,但仍能成功完成任务。这些发现支持了我们的假设:元认知学习可促进机器人规划中的创造性。
原文摘要 · Abstract (English)
While large language models (LLMs) have shown great potential across various domains, their applications in robotics remain largely limited to static prompt-based behaviors and still face challenges in complex tasks under zero-shot or few-shot settings. Inspired by human metacognitive learning and creative problem-solving, we address this limitation by exploring a fundamental question: Can LLMs be empowered with metacognitive capabilities to reason, reflect, and create, thereby enhancing their ability to perform robotic tasks with minimal demonstrations? In this paper, we present REFLEX, a framework that integrates metacognitive learning into LLM-powered multi-robot collaboration. The system equips the LLM-powered robotic agents with a skill decomposition and self-reflection mechanism that identifies modular skills from prior tasks, reflects on failures in unseen task scenarios, and synthesizes effective new solutions. We propose a more challenging robotic benchmark task and evaluate our framework on the existing benchmark and the novel task. Experimental results show that our metacognitive learning framework significantly outperforms existing baselines. Moreover, we observe that our framework can generate solutions that differ from the ground truth yet still successfully complete the tasks. These findings support our hypothesis that metacognitive learning can foster creativity in robotic planning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。