arXiv:2511.00917cs.ROcs.AI2025-11被引 8

用视觉语言模型指挥模块化工具,实现零样本通用机器人操控

Maestro: Orchestrating Robotics Modules with Vision-Language Models for Zero-Shot Generalist Robots

  • 以VLM为大脑,动态组合感知、规划、控制模块构建策略
  • 在复杂操作任务上零样本性能超越现有端到端模型
  • 模块可扩展,适配新机器人形态,少量代码修改即可落地

当前通用机器人研究多依赖大规模‘观察-动作’数据集训练端到端模型,沿袭视觉语言模型的成功路径。本文另辟蹊径:直接以视觉语言模型(VLM)为核心,通过精心设计的感知、规划与控制模块增强其能力。在Maestro系统中,一个VLM编码代理会根据当前任务和场景动态组合这些模块,生成程序化策略。该架构具备简洁闭环接口,无大量人工结构约束,并拥有丰富多样的工具库。结果表明,Maestro在复杂操作技能上的零样本表现显著优于现有视觉语言动作模型。此外,系统易于扩展新模块,可灵活适配如四足机器人搭载机械臂等新形态,甚至仅通过局部代码修改即可从极少真实经验中快速适应。

原文摘要 · Abstract (English)

Today's best-explored routes towards generalist robots center on collecting ever larger "observations-in actions-out" robotics datasets to train large end-to-end models, copying a recipe that has worked for vision-language models (VLMs). We pursue a road less traveled: building generalist policies directly around VLMs by augmenting their general capabilities with specific robot capabilities encapsulated in a carefully curated set of perception, planning, and control modules. In Maestro, a VLM coding agent dynamically composes these modules into a programmatic policy for the current task and scenario. Maestro's architecture benefits from a streamlined closed-loop interface without many manually imposed structural constraints, and a comprehensive and diverse tool repertoire. As a result, it largely surpasses today's VLA models for zero-shot performance on challenging manipulation skills. Further, Maestro is easily extensible to incorporate new modules, easily editable to suit new embodiments such as a quadruped-mounted arm, and even easily adapts from minimal real-world experiences through local code edits.

通用机器人视觉语言模型模块化零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。