让视觉语言模型直接指挥机器人动作,提升复杂任务的泛化能力。
Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control
- 用多层次抽象指令训练视觉语言动作模型,实现精细控制
- 在真实世界操作中,长时序和泛化任务性能超越基线方法
- 适配学习型推理器与现成大模型,支持灵活部署
预训练视觉语言模型(VLM)可在多种场景中进行语义与视觉推理,为机器人控制提供有价值的常识先验。然而,如何有效将这些知识融入机器人行为仍是开放挑战。现有方法通常采用分层架构,由VLM对高层指令进行推理,并交由独立的低层策略执行,如视觉语言动作模型(VLA)。但其接口常为自然语言任务指令,严重限制了VLM推理对底层行为的引导能力。为此,我们提出可调控策略:在多层级抽象的丰富合成指令上训练的VLA,包括子任务、运动动作及具体像素坐标。通过增强低层可控性,可充分释放预训练VLM中的知识,实现更好的任务泛化。我们在真实世界操作实验中验证,结合学习型高层具身推理器或现成的VLM通过上下文学习推理命令抽象,均显著优于以往的具身推理VLA及基于VLM的分层基线,在复杂泛化与长时序任务中表现优异。
原文摘要 · Abstract (English)
Pretrained vision-language models (VLMs) can make semantic and visual inferences across diverse settings, providing valuable common-sense priors for robotic control. However, effectively grounding this knowledge in robot behaviors remains an open challenge. Prior methods often employ a hierarchical approach where VLMs reason over high-level commands to be executed by separate low-level policies, e.g., vision-language-action models (VLAs). The interface between VLMs and VLAs is usually natural language task instructions, which fundamentally limits how much VLM reasoning can steer low-level behavior. We thus introduce Steerable Policies: VLAs trained on rich synthetic commands at various levels of abstraction, like subtasks, motions, and grounded pixel coordinates. By improving low-level controllability, Steerable Policies can unlock pretrained knowledge in VLMs, enabling improved task generalization. We demonstrate this benefit by controlling our Steerable Policies with both a learned high-level embodied reasoner and an off-the-shelf VLM prompted to reason over command abstractions via in-context learning. Across extensive real-world manipulation experiments, these two novel methods outperform prior embodied reasoning VLAs and VLM-based hierarchical baselines, including on challenging generalization and long-horizon tasks. Website: steerable-policies.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。