用视觉语言模型生成可解释的机器人行为树,实现零样本迁移。
Learning Structured Robot Policies from Vision-Language Models via Synthetic Neuro-Symbolic Supervision

- 通过合成多模态数据,让大模型自动生成带符号约束的行为树。
- 120亿参数模型仅靠仿真训练就学会空间符号映射,真实机器人零样本成功执行。
- 适合需要可解释性与模块化控制的工业级机器人应用。
视觉语言模型(VLMs)在将多模态观测映射到机器人行为方面展现出强大能力。然而,现有方法多依赖黑箱式端到端视觉运动策略,难以分析,限制了其在真实场景中的应用。相比之下,经典机器人系统常采用结构化策略表示,具备可解释性、模块化和反应式执行优势。本文研究如何将基础模型专门化,以生成基于多模态感知的结构化机器人策略,融合高维学习与符号控制。提出一种神经符号方法:由VLM根据视觉输入、自然语言指令和结构化系统规范,合成可执行的行为树(Behavior Tree, BT)策略。为实现无人工标注的规模化监督,设计自动化流程,生成领域随机化的场景与指令-策略配对的合成数据集。通过解耦受符号语法约束的任务分解与硬件特定的运动控制,实验表明120亿参数模型仅通过仿真监督即可学习到执行所需的空间-符号映射。在两种异构机械臂上的真实世界实验验证了这些结构化策略具备零样本迁移能力。结果表明,可通过程序化合成高质量神经符号训练数据,绕过机器人规划中的数据瓶颈。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have recently demonstrated strong capabilities in mapping multimodal observations to robot behaviors. However, most current approaches rely on end-to-end visuomotor policies that remain opaque and difficult to analyze, limiting their use in real-world robotic applications. In contrast, classical robotic systems often rely on structured policy representations that provide interpretability, modularity, and reactive execution. This work investigates how foundation models can be specialized to generate structured robot policies grounded in multimodal perception, bridging high-dimensional learning and symbolic control. We propose a neuro-symbolic approach in which a VLM synthesizes executable Behavior Tree policies from visual observations, natural language instructions, and structured system specifications. To enable scalable supervision without manual annotation, we introduce an automated pipeline that generates a synthetic multimodal dataset of domain-randomized scenes paired with instruction-policy examples produced by a foundation model. By decoupling structured task decomposition under constrained symbolic grammars from hardware-specific motor control, we demonstrate that a 12B-parameter model can learn structured spatial-symbolic mappings required for executable BT synthesis, solely through in-silico supervision. Real-world physical experiments on two heterogeneous robotic manipulators confirm that these structurally constrained policies achieve zero-shot transfer to real-world environments. The results emphasize that the data bottleneck in robotic planning can be bypassed by procedurally synthesizing high-fidelity, neuro-symbolic training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。