arXiv:2602.22010cs.ROcs.CV2026-02被引 31

用压缩条件建模未来,让视觉语言动作模型更精准地生成动作

World Guidance: World Modeling in Condition Space for Action Generation

  • 将未来观测压缩为条件注入动作推理流程
  • 在仿真与真实环境上均显著优于现有方法
  • 适合需要高精度动作生成的机器人任务

利用未来观测建模来提升视觉-语言-动作(VLA)模型的能力具有广阔前景。然而,现有方法难以在保持高效可预测的未来表示与保留足够精细信息以指导精确动作生成之间取得平衡。为此,我们提出WoG(World Guidance)框架,通过将未来观测映射为紧凑条件并注入动作推理流程,使VLA模型在预测未来动作的同时,也预测这些压缩条件,从而在条件空间中实现有效的世界建模。实验表明,对条件空间的建模不仅有助于精细化动作生成,还具备更强的泛化能力,并能从大量人类操作视频中有效学习。在仿真与真实环境中的广泛实验验证了该方法显著优于基于未来预测的现有方法。

原文摘要 · Abstract (English)

Leveraging future observation modeling to facilitate action generation presents a promising avenue for enhancing the capabilities of Vision-Language-Action (VLA) models. However, existing approaches struggle to strike a balance between maintaining efficient, predictable future representations and preserving sufficient fine-grained information to guide precise action generation. To address this limitation, we propose WoG (World Guidance), a framework that maps future observations into compact conditions by injecting them into the action inference pipeline. The VLA is then trained to simultaneously predict these compressed conditions alongside future actions, thereby achieving effective world modeling within the condition space for action inference. We demonstrate that modeling and predicting this condition space not only facilitates fine-grained action generation but also exhibits superior generalization capabilities. Moreover, it learns effectively from substantial human manipulation videos. Extensive experiments across both simulation and real-world environments validate that our method significantly outperforms existing methods based on future prediction. Project page is available at: https://selen-suyue.github.io/WoGNet/

动作生成世界建模VLA模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。