arXiv:2606.07723cs.RO2026-06被引 4

让机器人像指挥家一样协调多工具完成复杂长时任务。

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation

论文配图:VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation
图 1 · 摘自论文原文
  • 用视觉语言模型当指挥官,动态调度机械臂等工具
  • 在真实机器人上实现90%以上任务成功率,优于单一模型
  • 适合需要多步推理和容错的智能服务场景

开放词汇长时程操作要求机器人能理解灵活指令、处理多物体复杂场景,并自适应地规划、执行、监控与恢复失败。我们提出一个闭环智能体系统,其中视觉语言模型(VLM)将异构机器人能力作为可中断工具进行编排。不同于虚拟AI代理,物理世界中决策、动作和工具调用的时间至关重要。我们称此为物理编排(Physical Orchestration),并提出VoLoAgent,一个将视觉语言动作模型(VLA)/WAM机械臂视为可中断工具,在推理过程中动态控制的系统,同时协同视觉模型与动作原语。为评估其长时程能力,我们引入RoboVoLo基准,该基准涵盖常识、记忆/状态追踪、复杂指代和世界知识,提供任务级成功率与故障模式诊断。实验表明,VoLoAgent显著优于单个VLA/VLM或基于工具的系统,并在真实机器人上得到验证。

原文摘要 · Abstract (English)

Open-vocabulary long-horizon manipulation requires robots to reason over flexible instructions and complex multi-object scenes while adaptively planning, executing, monitoring, and recovering from failures. We address these demands with a closed agent loop in which a VLM orchestrates heterogeneous robot capabilities as interruptible tools. Unlike in virtual AI agents, the timing of decisions, actions and tool calls is important in a physical world that does not pause for reasoning. We refer to this setting as Physical Orchestration, and propose VoLoAgent, a VLM that plans, monitors, and recovers by treating a VLA/WAM as an interruptible tool it steers mid-rollout alongside vision models and action primitives. To evaluate these long-horizon capabilities, we introduce RoboVoLo, a high-fidelity benchmark for open-vocabulary long-horizon manipulation across common sense, memory/state tracking, complex references, and world knowledge, with both task-level success and failure-mode diagnostics. Experiments show VoLoAgent substantially outperforms single VLA/VLM or tool-based systems, with validation on real-robot experiments. Project page: https://chicychen.github.io/VoLo/

机器人视觉语言模型长时程操作工具调度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。