用智能调度器让通用机器人自主完成复杂任务,性能提升超4倍。
Addressing the Orchestration Gap in Generalist Robots via Physical Agency

- 将机器人能力拆分为语言指令代理和高层调度器协同工作。
- 在真实机器人上实现90%成功率,比原策略提升超4倍。
- 无需额外训练,即可让冻结技能发挥更高水平,适合做通用机器人研发。
通用机器人需要整合感知、世界知识、规划、成功检测、恢复与底层控制。现有方法通过大规模预训练将所有能力集成到策略中,但本文提出将这些能力解耦:由一个语言驱动的通用策略/控制代理和一个高层调度器组成。我们构建了闭环物理代理调度器(Pigey),能进行高层规划、分解目标为可执行子目标、下达低层动作指令、根据观测追踪结果并处理失败。Pigey可操控现有的视觉-语言-动作(VLA)策略及参数化技能,在不收集新数据或后训练的情况下解决真实世界的复杂推理任务。在仿真基准和真实机器人操控任务中广泛评估,表现显著优于现有通用策略。在LIBERO-PRO上,性能从12.8%提升至53.3%,超越现有水平超4倍;在真实机器人上,使冻结策略在推理受限任务中的成功率从接近零提升至90%以上。我们称冻结技能单独与在智能体循环中表现之间的差距为‘调度鸿沟’。
原文摘要 · Abstract (English)
General-purpose robots need to reason about their actions, combining perception, world knowledge, planning, success detection, recovery, and low-level control. Today's state-of-the-art models attempt to combine all these capabilities into the learned policy via large-scale pre-training. Instead, we show that these capabilities can be decomposed into a general language-conditioned policy/control agent and a high-level agent manager/orchestrator. Rather than training policies to reason via pre-training, we build a closed-loop physical agent orchestrator that can do high-level planning, decompose the goal into achievable subgoals, command low-level motor commands, track and verify the outcome from low-level observations, and recover from failures. Our Physical Agency orchestrator (Pigey) can control existing vision-language-action (VLA) policies as well as parametrized skills to solve complex reasoning tasks in the real world, without any additional data collection or post-training. We evaluate Pigey extensively across simulation benchmarks and challenging real-world robotic manipulation tasks, and demonstrate significant performance improvements over existing generalist policies. On LIBERO-PRO, Pigey advances the state-of-the-art by over 4x (12.8% -> 53.3%) with no task-specific fine-tuning. On a real robot, Pigey lifts the frozen policy from near-zero to over 90% on reasoning-limited tasks. We call the difference between what frozen motor skills achieve alone and inside the agentic loop the orchestration gap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。