用分层架构让机器人更懂指令并精准执行复杂任务。
RoBridge: A Hierarchical Architecture Bridging Cognition and Execution for General Robotic Manipulation
- 用大模型理解指令,符号化表示连接认知与执行
- 新任务成功率75%,仅需每任务5个真实样本即可泛化
- 适合研究通用机器人操作与人机协作的开发者
在开放场景中执行多样化任务是机器人领域的重要方向。尽管自然语言处理和大规模多模态模型提升了机器人理解复杂指令的能力,但其操作仍面临程序性技能与陈述性技能的双重困境。现有方法常在认知与执行能力间妥协。为此,本文提出RoBridge,一种分层智能架构,包含基于大模型的高层认知规划器(HCP)、作为符号桥梁的不变可操作表示(IOR)以及通用具身代理(GEA)。该架构既保留了大模型的陈述性知识,又释放强化学习的程序性能力,有效弥合认知与执行的鸿沟。RoBridge在新任务上实现75%的成功率,在模拟到现实的泛化中达到83%的平均成功率,且每个任务仅需5个真实世界数据样本。本工作为机器人系统中认知推理与物理执行的融合提供了新范式。
原文摘要 · Abstract (English)
Operating robots in open-ended scenarios with diverse tasks is a crucial research and application direction in robotics. While recent progress in natural language processing and large multimodal models has enhanced robots' ability to understand complex instructions, robot manipulation still faces the procedural skill dilemma and the declarative skill dilemma in open environments. Existing methods often compromise cognitive and executive capabilities. To address these challenges, in this paper, we propose RoBridge, a hierarchical intelligent architecture for general robotic manipulation. It consists of a high-level cognitive planner (HCP) based on a large-scale pre-trained vision-language model (VLM), an invariant operable representation (IOR) serving as a symbolic bridge, and a generalist embodied agent (GEA). RoBridge maintains the declarative skill of VLM and unleashes the procedural skill of reinforcement learning, effectively bridging the gap between cognition and execution. RoBridge demonstrates significant performance improvements over existing baselines, achieving a 75% success rate on new tasks and an 83% average success rate in sim-to-real generalization using only five real-world data samples per task. This work represents a significant step towards integrating cognitive reasoning with physical execution in robotic systems, offering a new paradigm for general robotic manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。