arXiv:2506.21627cs.ROcs.AI2025-06

用类脑架构整合视觉语言模型,实现无需微调的高效机器人操作

FrankenBot: Brain-Morphic Modular Orchestration for Robotic Manipulation with Vision-Language Models

  • 模仿人脑结构,分模块设计任务规划、策略生成等核心功能
  • 在仿真与真实机器人上均实现异常处理、长期记忆和高效运行
  • 无需微调或重训,适合复杂动态环境下的通用机器人系统

构建能在复杂、动态、非结构化真实环境中执行多种任务的通用机器人操作系统长期面临挑战。实现类人级的高效与鲁棒性操作,要求机器人具备任务规划、策略生成、异常监控与处理、长期记忆等综合功能,并在所有功能间保持高效率协同。视觉-语言模型(VLM)通过大规模多模态数据预训练,已具备丰富的世界知识、出色的场景理解与多模态推理能力。然而现有方法通常仅实现部分功能,缺乏统一认知架构的整合。受分治策略与人脑结构启发,我们提出FrankenBot:一种由VLM驱动、类脑架构的机器人操作框架,兼顾功能完备性与系统效率。该框架将任务规划、策略生成、记忆管理、底层接口分别映射至皮层、小脑、颞叶-海马复合体和脑干,并设计高效的模块协调机制。在仿真与真实机器人环境中开展全面实验,结果表明本方法在异常检测与处理、长期记忆、运行效率与稳定性方面均具显著优势,且无需任何微调或重训练。

原文摘要 · Abstract (English)

Developing a general robot manipulation system capable of performing a wide range of tasks in complex, dynamic, and unstructured real-world environments has long been a challenging task. It is widely recognized that achieving human-like efficiency and robustness manipulation requires the robotic brain to integrate a comprehensive set of functions, such as task planning, policy generation, anomaly monitoring and handling, and long-term memory, achieving high-efficiency operation across all functions. Vision-Language Models (VLMs), pretrained on massive multimodal data, have acquired rich world knowledge, exhibiting exceptional scene understanding and multimodal reasoning capabilities. However, existing methods typically focus on realizing only a single function or a subset of functions within the robotic brain, without integrating them into a unified cognitive architecture. Inspired by a divide-and-conquer strategy and the architecture of the human brain, we propose FrankenBot, a VLM-driven, brain-morphic robotic manipulation framework that achieves both comprehensive functionality and high operational efficiency. Our framework includes a suite of components, decoupling a part of key functions from frequent VLM calls, striking an optimal balance between functional completeness and system efficiency. Specifically, we map task planning, policy generation, memory management, and low-level interfacing to the cortex, cerebellum, temporal lobe-hippocampus complex, and brainstem, respectively, and design efficient coordination mechanisms for the modules. We conducted comprehensive experiments in both simulation and real-world robotic environments, demonstrating that our method offers significant advantages in anomaly detection and handling, long-term memory, operational efficiency, and stability -- all without requiring any fine-tuning or retraining.

机器人操作类脑架构视觉语言模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。