一个能看会想还会说话的机器人大脑,让机械臂听懂复杂指令并自主规划任务。
Robix: A Unified Model for Robot Interaction, Reasoning and Planning
- 用视觉语言模型统一处理机器人的思考、规划和对话
- 在洗碗、购物等任务中表现超越GPT-4o和Gemini 2.5 Pro
- 支持主动对话、打断响应和常识推理,适合人机协作场景
我们提出Robix,一个统一模型,将机器人推理、任务规划与自然语言交互集成于单一视觉语言架构中。作为分层机器人系统中的高层认知层,Robix动态生成低层控制器的原子命令和面向人类的语音回应,实现对复杂指令的理解、长时程任务的规划以及自然的人机交互,所有过程在端到端框架内完成。Robix引入了主动对话、实时中断处理和执行过程中的上下文感知常识推理等新能力。其核心采用思维链推理,并通过三阶段训练策略:(1)持续预训练以增强基础具身推理能力,包括3D空间理解、视觉定位和任务中心推理;(2)监督微调以建模人机交互与任务规划为统一的推理-行动序列;(3)强化学习提升推理-行动一致性与长时程任务连贯性。大量实验表明,Robix在交互式任务执行中优于开源及商用基线模型(如GPT-4o和Gemini 2.5 Pro),在多样化指令类型(开放式、多阶段、受限、无效、中断)及用户参与任务(如桌椅清理、杂货采购、饮食过滤)中均表现出强泛化能力。
原文摘要 · Abstract (English)
We introduce Robix, a unified model that integrates robot reasoning, task planning, and natural language interaction within a single vision-language architecture. Acting as the high-level cognitive layer in a hierarchical robot system, Robix dynamically generates atomic commands for the low-level controller and verbal responses for human interaction, enabling robots to follow complex instructions, plan long-horizon tasks, and interact naturally with human within an end-to-end framework. Robix further introduces novel capabilities such as proactive dialogue, real-time interruption handling, and context-aware commonsense reasoning during task execution. At its core, Robix leverages chain-of-thought reasoning and adopts a three-stage training strategy: (1) continued pretraining to enhance foundational embodied reasoning abilities including 3D spatial understanding, visual grounding, and task-centric reasoning; (2) supervised finetuning to model human-robot interaction and task planning as a unified reasoning-action sequence; and (3) reinforcement learning to improve reasoning-action consistency and long-horizon task coherence. Extensive experiments demonstrate that Robix outperforms both open-source and commercial baselines (e.g., GPT-4o and Gemini 2.5 Pro) in interactive task execution, demonstrating strong generalization across diverse instruction types (e.g., open-ended, multi-stage, constrained, invalid, and interrupted) and various user-involved tasks such as table bussing, grocery shopping, and dietary filtering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。