arXiv:2506.00123cs.CVcs.RO2025-06被引 38

让大模型像人一样看、想、控制机器人,实现视觉与行动统一。

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

  • 将机器人控制转为文本任务,在2D视觉空间中统一感知与决策
  • 在13个基准上超越Qwen2.5-VL,腿式机器人平均提升50%
  • 适配真实机器人,支持复杂任务组合与灵活应对

多模态大语言模型(MLLMs)的快速发展推动其向具身实体如四足机器人延伸。这要求模型兼具多模态理解、视觉空间推理与物理交互能力,但现有方法因本质差异难以统一。本文提出视觉具身大脑(VeBrain),构建一个统一的现实世界感知、推理与控制框架。VeBrain将机器人控制重构为二维视觉空间中的通用文本任务,统一任务目标与映射空间;并设计新型机器人适配器,将大模型输出的文本指令转化为真实机器人的运动策略。从数据角度,构建了高质量指令数据集VeBrain-600k,耗时数百小时采集、清洗与标注,采用多模态思维链(CoT)融合多种能力于单一流程。在13个多模态基准与5个空间智能基准上的实验表明,VeBrain显著优于现有模型如Qwen2.5-VL。部署至四足机器人与机械臂后,展现更强适应性、灵活性与组合能力:相比Qwen2.5-VL,MMVet得分提升+5.6%,腿式机器人任务平均提升+50%。

原文摘要 · Abstract (English)

The remarkable progress of Multimodal Large Language Models (MLLMs) has attracted increasing attention to extend them to physical entities like legged robot. This typically requires MLLMs to not only grasp multimodal understanding abilities, but also integrate visual-spatial reasoning and physical interaction capabilities. Nevertheless,existing methods struggle to unify these capabilities due to their fundamental differences.In this paper, we present the Visual Embodied Brain (VeBrain), a unified framework for perception, reasoning, and control in real world. VeBrain reformulates robotic control into common text-based MLLM tasks in the 2D visual space, thus unifying the objectives and mapping spaces of different tasks. Then, a novel robotic adapter is proposed to convert textual control signals from MLLMs to motion policies of real robots. From the data perspective, we further introduce VeBrain-600k, a high-quality instruction dataset encompassing various capabilities of VeBrain. In VeBrain-600k, we take hundreds of hours to collect, curate and annotate the data, and adopt multimodal chain-of-thought(CoT) to mix the different capabilities into a single conversation. Extensive experiments on 13 multimodal benchmarks and 5 spatial intelligence benchmarks demonstrate the superior performance of VeBrain to existing MLLMs like Qwen2.5-VL. When deployed to legged robots and robotic arms, VeBrain shows strong adaptability, flexibility, and compositional capabilities compared to existing methods. For example, compared to Qwen2.5-VL, VeBrain not only achieves substantial gains on MMVet by +5.6%, but also excels in legged robot tasks with +50% average gains.

具身智能多模态大模型机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。