arXiv:2602.11832cs.CVcs.RO2026-02被引 16

用视频预测嵌入提升机器人视觉-语言-动作模型的泛化与采样效率。

JEPA-VLA: Video Predictive Embedding is Needed for VLA Models

  • 引入视频预测嵌入增强视觉表征,捕捉任务相关时序动态。
  • 在多个基准上实现显著性能提升,包括真实机器人任务。
  • 适合关注机器人学习泛化能力的研究者与工程师。

基于预训练视觉-语言模型(VLM)的视觉-语言-动作(VLA)模型在机器人操作中取得显著进展,但现有VLA仍存在样本效率低、泛化能力有限的问题。本文指出,这些问题根源在于被忽视的预训练视觉表示:当前主流的图文对比或图像自监督学习得到的表征,在环境理解与策略先验方面知识不足。分析发现,这些表征难以捕捉任务相关的环境信息,也缺乏对成功执行任务时环境演化的预测性知识。相比之下,基于视频预训练的预测嵌入(如V-JEPA 2)能灵活忽略不可预测因素,有效编码任务相关的时间动态。基于此,我们提出JEPA-VLA,一种简单有效的融合方式,将预测嵌入自适应地集成到现有VLA中。实验表明,JEPA-VLA在LIBERO、LIBERO-plus、RoboTwin2.0及真实机器人任务上均实现显著性能提升。

原文摘要 · Abstract (English)

Recent vision-language-action (VLA) models built upon pretrained vision-language models (VLMs) have achieved significant improvements in robotic manipulation. However, current VLAs still suffer from low sample efficiency and limited generalization. This paper argues that these limitations are closely tied to an overlooked component, pretrained visual representation, which offers insufficient knowledge on both aspects of environment understanding and policy prior. Through an in-depth analysis, we find that commonly used visual representations in VLAs, whether pretrained via language-image contrastive learning or image-based self-supervised learning, remain inadequate at capturing crucial, task-relevant environment information and at inducing effective policy priors, i.e., anticipatory knowledge of how the environment evolves under successful task execution. In contrast, we discover that predictive embeddings pretrained on videos, in particular V-JEPA 2, are adept at flexibly discarding unpredictable environment factors and encoding task-relevant temporal dynamics, thereby effectively compensating for key shortcomings of existing visual representations in VLAs. Building on these observations, we introduce JEPA-VLA, a simple yet effective approach that adaptively integrates predictive embeddings into existing VLAs. Our experiments demonstrate that JEPA-VLA yields substantial performance gains across a range of benchmarks, including LIBERO, LIBERO-plus, RoboTwin2.0, and real-robot tasks.

机器人学习视频预测视觉表征VLA模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。