arXiv:2506.23919cs.RO2025-06被引 11

用图像生成模型做机器人抓取的世界模型,零样本泛化能力强

Goal-VLA: Image-Generative VLMs as Object-Centric World Models Empowering Zero-shot Robot Manipulation

  • 用图像生成VLM构建目标状态世界模型,自动推导物体位姿
  • 在模拟和真实场景中均实现高泛化能力的零样本操作
  • 无需额外训练,适合快速部署到新任务和新机器人

泛化能力仍是机器人操作的核心挑战。现有视觉-语言-动作(VLA)模型依赖视觉-语言模型(VLM)的开放世界语义知识,但其零样本性能远低于基础VLM,因指令-视觉-动作数据难以覆盖多样场景、任务与机器人形态。本文提出Goal-VLA,利用图像生成式VLM作为世界模型,生成期望的目标状态图像,并从中提取目标物体位姿,实现可泛化的操作。核心思想是将物体状态表示作为通用接口,分离高层策略与低层控制。该表示不依赖显式动作标注,使高泛化能力的VLM得以使用,同时提供空间线索以实现免训练的底层控制。为进一步提升鲁棒性,引入反射合成机制,在执行前迭代验证并优化生成的目标图像。仿真与真实实验均表明,该方法在多种操作任务中表现优异且具备出色泛化能力。补充材料见 https://nus-lins-lab.github.io/goalvlaweb/

原文摘要 · Abstract (English)

Generalization remains a fundamental challenge in robotic manipulation. To tackle this challenge, recent Vision-Language-Action (VLA) models build policies on top of Vision-Language Models (VLMs), seeking to transfer their open-world semantic knowledge. However, their zero-shot capability lags significantly behind the base VLMs, as the instruction-vision-action data is too limited to cover diverse scenarios, tasks, and robot embodiments. In this work, we present Goal-VLA, a zero-shot framework that leverages Image-Generative VLMs as world models to generate desired goal states, from which the target object pose is derived to enable generalizable manipulation. The key insight is that object state representation is the golden interface, naturally separating a manipulation system into high-level and low-level policies. This representation abstracts away explicit action annotations, allowing the use of highly generalizable VLMs while simultaneously providing spatial cues for training-free low-level control. To further improve robustness, we introduce a Reflection-through-Synthesis process that iteratively validates and refines the generated goal image before execution. Both simulated and real-world experiments demonstrate that our \name achieves strong performance and inspiring generalizability in manipulation tasks. Supplementary materials are available at https://nus-lins-lab.github.io/goalvlaweb/.

机器人操作零样本学习世界模型图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。