让视觉语言动作模型学会预判未来并规划运动路径
FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models

- 联合预测未来特征与稀疏点轨迹,实现目标与路径的协同建模
- 在三个基准上达到最优性能,零样本泛化能力显著提升
- 适合关注具身智能与连续动作规划的研究者
视觉-语言-动作(VLA)模型在视觉运动策略学习中表现优异,但本质上仍是反应式系统,仅将当前观测与语言映射为动作,缺乏对世界动态的显式前向预测。现有视觉前瞻方法可预测未来视觉状态,但缺乏明确的运动引导:能指明去向却无法说明如何抵达。我们提出FoMoVLA框架,通过联合学习未来特征前瞻与稀疏2D点跟踪,为VLA表征引入显式的时空监督。该框架采用紧凑的前瞻标记解码未来特征状态,解码稀疏时序2D点轨迹以建模紧凑几何运动,并通过轻量级未来条件交叉注意力模块耦合二者,实现预期状态与点动态的一致推理。在LIBERO、RoboCasa GR-1 Tabletop和LIBERO-Plus上的大量实验表明,该方法达到当前最优性能,并展现出强大的零样本泛化能力。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have achieved impressive results in visuomotor policy learning, yet remain fundamentally reactive, mapping current observations and language to actions without explicit forward prediction of world dynamics. Existing visual foresight methods predict future visual states but lack explicit motion guidance: they show where to go but not how to get there. We argue that future feature prediction and sparse point tracking are naturally complementary: the former provides the goal state, while the latter captures the continuous motion path toward it. We propose FoMoVLA, a framework that augments VLA representations with explicit spatio-temporal supervision by jointly learning future feature foresight and sparse 2D point tracking, enhancing the continuous action policy. FoMoVLA introduces compact foresight tokens to decode future feature states, decodes sparse temporal 2D point trajectories to model compact geometric motion, and couples both through a lightweight future-conditioned cross-attention module that enables consistent reasoning between anticipated states and point dynamics. Extensive experiments on LIBERO, RoboCasa GR-1 Tabletop, and LIBERO-Plus demonstrate state-of-the-art performance and strong zero-shot generalization. Project page is available at https://liauto-research.github.io/FoMoVLA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。