提出几何感知的视觉-语言-动作模型,提升机器人操作在环境变化下的鲁棒性。
GWM-VLA: Geometry-Aware Latent World Modeling for Vision-Language-Action Learning

- 融合多视角几何信息构建状态表示,保留视图间空间关系。
- 以机械臂末端视角为目标预测下一帧局部图像,聚焦抓取动作细节。
- 联合使用隐式预测与真实动作监督,优化动作表征,适合复杂操控任务。
视觉-语言-动作(VLA)模型在机器人操作中表现优异,但在视觉和环境变化下性能常下降。潜在世界建模为提升鲁棒性提供了新思路,但现有方法通常独立编码相机视角,且不显式建模其几何关系。本文提出GWM-VLA,一种面向VLA学习的几何感知潜在世界建模框架。该框架结合几何感知的多视角状态编码、全局上下文条件的目标视角预测,以及由机器人动作监督引导的共享隐式动作表征。具体地,VGGT-Ω在每一步聚合多视角观测,生成具有几何感知能力的多视角状态。潜在世界模型利用聚合后的块令牌和注册令牌,预测选定目标视角的下一步块令牌,从而保留多视角几何信息,而无需预测完整多视角状态。实验中以腕部视角为目标,更关注末端执行器运动和夹爪-物体局部交互。最后,共享的隐式动作表征同时作用于潜在世界模型和流匹配动作头,使隐式预测监督与真实机器人动作监督共同塑造同一动作表征。仿真与真实环境中的实验验证了GWM-VLA的有效性与鲁棒性。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but often degrade under visual and environmental shifts. Latent world modeling offers a promising approach to improving robustness, yet existing methods commonly encode camera views independently and predict holistic scene dynamics without explicitly modeling their geometric relationships. We propose GWM-VLA, a geometry-aware latent world modeling framework for VLA learning. GWM-VLA combines geometry-aware multi-view state encoding, global context-conditioned target-view prediction, and shared latent-action representations grounded by robot-action supervision. Specifically, VGGT-$Ω$ jointly aggregates multi-view observations at each timestep to construct geometry-aware multi-view states. The latent world model predicts the next-step patch tokens of a selected target view using patch and register tokens obtained after multi-view aggregation, thereby retaining multi-view geometric information without predicting the complete multi-view state. We use the wrist view as the target in our experiments, placing greater emphasis on end-effector motion and local gripper-object interactions. Finally, the shared latent-action representations condition both the latent world model and the flow-matching action head, allowing latent-prediction supervision and ground-truth robot-action supervision to jointly shape the same latent-action representations. Experiments across both simulation and real-world environments demonstrate the effectiveness and robustness of GWM-VLA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。