用物理轨迹引导视觉模型,让预测更准、控制更稳。
JEPA-x: Cross-Predictive Physics Grounding for Forecastable Latent Dynamics

- 用物理状态和视觉图像作为双视角,共同训练预测器
- 滚动预测误差从0.361降至0.104,控制成功率升至78.2%
- 适合需要高精度动态建模的机器人控制任务
潜在世界模型通过预测候选动作如何推进已学习的潜在动态来进行规划。然而,在自预测模型中,编码器与预测器联合优化,可能共适应出易于预测但与场景物理演化弱相关的潜在转移。我们提出交叉预测的JEPA(JEPA-x),将潜在动态锚定在特权物理轨迹上。JEPA-x将视觉观测与物理状态视为同一动作条件轨迹的对应视图,通过共享预测器同时推进两者,并将每个预测与两种模态的未来表征对齐。这促使动作条件预测器学习跨模态的统一转移规则。特权物理状态仅用于训练,部署时为纯视觉模型。实验证明,新拟合预测器的滚动漂移从0.361降至0.104,多任务套件上平均控制成功率从53.6%提升至78.2%,覆盖六个评估子类别。此外,直接回归物理状态虽提升解码能力,但未改善可预测性或控制性能,表明收益源于塑造潜在动态,而非简单编码物理变量。
原文摘要 · Abstract (English)
Latent world models plan by predicting how candidate actions advance learned latent dynamics. In self-predictive models, however, the encoder and predictor are optimized jointly and can co-adapt to latent transitions that are easy to predict but weakly constrained by the physical evolution of the scene. We introduce the cross-predictive JEPA (JEPA-x), which grounds latent dynamics in privileged physical trajectories. JEPA-x treats visual observations and physical states as corresponding views of the same action-conditioned trajectory, advances both through a shared predictor, and matches each prediction to the future representations of both modalities. This encourages the action-conditioned predictor to learn a common transition rule across the two views. Privileged physical state is used only during training, leaving a visual-only model at deployment. Empirical results show that JEPA-x reduces the rollout drift of a newly fitted predictor from $0.361$ to $0.104$ and increases mean control success from $53.6\%$ to $78.2\%$ on a multi-task suite spanning six evaluation subfamilies. We additionally show that direct physical-state regression improves decodability without improving forecastability or control, indicating that the benefit comes from shaping latent dynamics rather than merely encoding physical variables.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。