arXiv:2602.10098cs.ROcs.CV2026-02被引 70

用隐空间状态预测提升视觉语言动作模型的泛化能力

VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model

  • 未来帧仅作监督目标,不作为输入,避免信息泄露
  • 在隐空间预测状态变化,对镜头移动和背景干扰更鲁棒
  • 两阶段训练简单高效,适合真实世界机械臂任务

在互联网规模视频上预训练视觉语言动作(VLA)策略具有吸引力,但现有隐动作目标常学习错误:仍依赖像素变化而非与动作相关的状态转移,易受外观偏差、无关运动和信息泄露影响。我们提出 VLA-JEPA,一种基于 JEPA 风格的预训练框架,从设计上规避这些问题。核心思想是无泄露状态预测:目标编码器从未来帧生成隐表示,学生路径仅观察当前观测——未来信息仅作为监督目标,从不作为输入。在隐空间而非像素空间进行预测,使 VLA-JEPA 学习到对相机运动和无关背景变化鲁棒的动力学抽象。这带来一个简单的两阶段流程——JEPA 预训练后接动作头微调——无需以往多阶段隐动作流水线的复杂性。在 LIBERO、LIBERO-Plus、SimplerEnv 和真实世界操作任务上的实验表明,VLA-JEPA 在泛化性和鲁棒性上持续优于现有方法。

原文摘要 · Abstract (English)

Pretraining Vision-Language-Action (VLA) policies on internet-scale video is appealing, yet current latent-action objectives often learn the wrong thing: they remain anchored to pixel variation rather than action-relevant state transitions, making them vulnerable to appearance bias, nuisance motion, and information leakage. We introduce VLA-JEPA, a JEPA-style pretraining framework that sidesteps these pitfalls by design. The key idea is leakage-free state prediction: a target encoder produces latent representations from future frames, while the student pathway sees only the current observation -- future information is used solely as supervision targets, never as input. By predicting in latent space rather than pixel space, VLA-JEPA learns dynamics abstractions that are robust to camera motion and irrelevant background changes. This yields a simple two-stage recipe -- JEPA pretraining followed by action-head fine-tuning -- without the multi-stage complexity of prior latent-action pipelines. Experiments on LIBERO, LIBERO-Plus, SimplerEnv and real-world manipulation tasks show that VLA-JEPA achieves consistent gains in generalization and robustness over existing methods.

视觉语言动作隐空间建模机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。