让机器人团队自我反思并动态调整计划,提升复杂任务执行成功率。
REMAC: Self-Reflective and Self-Evolving Multi-Agent Collaboration for Long-Horizon Robot Manipulation
- 引入自反思与自演化模块,实现无需预设提示的环境适应性规划。
- 在多机器人协作下,任务成功率提升40%,执行效率提高52.7%。
- 适合需要长期、复杂、动态环境下的机器人任务设计与部署。
视觉语言模型(VLMs)在机器人长周期任务规划中展现出强大能力,尤其在需全局理解环境以分解任务的场景中。现有方法依赖先验环境知识或特定任务提示,难以应对动态变化或意外情况,如机器人发现微波炉门关闭而无法放置胡萝卜。这暴露了适应性与效率两大挑战。为此,本文提出自反思与自演化多智能体协作框架REMAC,通过持续反思与自我进化,实现高效、无场景依赖的多机器人长周期任务规划与执行。REMAC包含两个核心模块:自反思模块在循环中进行前提与后置条件检查,评估进展并优化计划;自演化模块基于场景特定推理动态调整策略。其优势包括:1)无需复杂提示即可初始探索与推理环境;2)持续识别规划错误并根据任务洞察调整方案;3)经迭代后可调用其他机器人并行协作,提升执行效率。为验证效果,我们基于RoboCasa构建多智能体长周期机器人操作与导航环境,涵盖4类任务、27种任务形式及50+不同物体。在此基础上,对DeepSeek-R1、o3-mini、QwQ和Grok3等先进推理模型进行基准测试,结果显示REMAC相比单机器人基线平均成功率提升40%,执行效率提高52.7%。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have demonstrated remarkable capabilities in robotic planning, particularly for long-horizon tasks that require a holistic understanding of the environment for task decomposition. Existing methods typically rely on prior environmental knowledge or carefully designed task-specific prompts, making them struggle with dynamic scene changes or unexpected task conditions, e.g., a robot attempting to put a carrot in the microwave but finds the door was closed. Such challenges underscore two critical issues: adaptability and efficiency. To address them, in this work, we propose an adaptive multi-agent planning framework, termed REMAC, that enables efficient, scene-agnostic multi-robot long-horizon task planning and execution through continuous reflection and self-evolution. REMAC incorporates two key modules: a self-reflection module performing pre-condition and post-condition checks in the loop to evaluate progress and refine plans, and a self-evolvement module dynamically adapting plans based on scene-specific reasoning. It offers several appealing benefits: 1) Robots can initially explore and reason about the environment without complex prompt design. 2) Robots can keep reflecting on potential planning errors and adapting the plan based on task-specific insights. 3) After iterations, a robot can call another one to coordinate tasks in parallel, maximizing the task execution efficiency. To validate REMAC's effectiveness, we build a multi-agent environment for long-horizon robot manipulation and navigation based on RoboCasa, featuring 4 task categories with 27 task styles and 50+ different objects. Based on it, we further benchmark state-of-the-art reasoning models, including DeepSeek-R1, o3-mini, QwQ, and Grok3, demonstrating REMAC's superiority by boosting average success rates by 40% and execution efficiency by 52.7% over the single robot baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。