arXiv:2608.09381cs.RO2026-08被引 1

用联合嵌入建模视觉-语言-动作,实现高效机器人控制。

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling

论文配图:JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling
图 1 · 摘自论文原文
  • 在预训练V-JEPA空间中构建隐式世界模型,共享预测器耦合状态转移与动作生成。
  • 在LIBERO-Plus上达到79.2%成功率,无需大规模机器人预训练即可领先现有方法。
  • 适用于多场景迁移,实测在视觉和空间变化下仍具强泛化能力,适合实际部署。

稳健的机器人控制依赖于对状态转移的显式建模,但视频生成型世界动作模型(WAM)带来显著部署成本。现有隐式WAM虽避免显式未来生成,却常压缩预测表示或分离预测与动作生成路径。本文提出JEPA-WAM,一种基于预训练V-JEPA空间的隐式WAM,通过共享预测器将隐式状态转移预测与连续动作生成耦合。该模型预测一个空间结构化的当前-未来联合目标,捕捉视觉时序结构的同时保持像素级对应关系。共享预测器使转移监督直接作用于主干网络,从中提取专用动作预测表示。相同设计可集成至预训练的VLA策略中,保留原有感知与动作路径。在LIBERO-Plus上,JEPA-WAM达79.2%成功率,为无大规模机器人预训练下的最佳结果;其预训练版本π_{0.5}更达86.3%,取得最优整体性能。在RoboTwin 2.0及真实双臂操作任务中,亦展现出强泛化性,尤其在视觉与空间偏移下表现优异。

原文摘要 · Abstract (English)

Robust robot control benefits from explicitly modeling state transitions, but video-generation world action models (WAMs) introduce substantial deployment cost. Existing latent WAMs avoid explicit future generation, but often compress predictive representations or separate predictive modeling from the representations used for action generation. We introduce JEPA-WAM, a latent WAM built in a pretrained V-JEPA space, which couples latent transition prediction with continuous action generation through a shared predictor. JEPA-WAM predicts a spatially structured joint current-future target that captures task-shared visual temporal structure between current and future observations, while preserving dense patch-level correspondence. Through the shared predictor, transition supervision directly shapes the backbone, from which dedicated representations are extracted for action prediction. The same design can also be instantiated in pretrained VLA policies while preserving their original perception and action pathways. On LIBERO-Plus, JEPA-WAM achieves 79.2%, the best result without large-scale robot-policy pretraining, while its pretrained $π_{0.5}$ instantiation reaches 86.3%, achieving the best overall performance. Experiments on RoboTwin 2.0 and real-world bimanual manipulation further demonstrate strong generalization under visual and spatial shifts.

机器人控制世界模型联合嵌入泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。