arXiv:2601.14354cs.LG2026-01被引 12

提出可建模不确定性的自监督世界模型,提升复杂环境下的鲁棒规划能力。

VJEPA: Variational Joint Embedding Predictive Architectures as Probabilistic World Models

  • 用变分推断学习未来隐状态的预测分布,突破传统确定性预测限制。
  • 在噪声环境中有效过滤干扰项,避免表征坍塌,优于生成基线方法。
  • 适合需要不确定性估计的机器人控制、强化学习等高维动态系统任务。

联合嵌入预测架构(JEPA)通过预测隐表示而非重建高熵观测,提供了一种可扩展的自监督学习范式。然而,现有方法依赖确定性回归目标,掩盖了概率语义,限制了其在随机控制中的应用。本文提出变分JEPA(VJEPA),一种概率化推广,通过变分目标学习未来隐状态的预测分布。我们证明VJEPA统一了表征学习与预测状态表示(PSRs)及贝叶斯滤波,表明序列建模无需自回归观测似然。理论上,VJEPA表征可作为最优控制的充分信息状态,无需像素重建,并提供防坍塌的正式保证。我们进一步提出贝叶斯JEPA(BJEPA),将预测信念分解为学习的动力学专家和模块化先验专家,实现零样本任务迁移及约束(如目标、物理)满足。实验表明,在噪声环境中,VJEPA和BJEPA能成功过滤高方差干扰项,避免生成基线中的表征坍塌。通过无需观测似然即可进行严谨的不确定性估计(如采样构建可信区间),VJEPA为高维、噪声环境中的可扩展、鲁棒、不确定性感知规划提供了基础框架。

原文摘要 · Abstract (English)

Joint Embedding Predictive Architectures (JEPA) offer a scalable paradigm for self-supervised learning by predicting latent representations rather than reconstructing high-entropy observations. However, existing formulations rely on \textit{deterministic} regression objectives, which mask probabilistic semantics and limit its applicability in stochastic control. In this work, we introduce \emph{Variational JEPA (VJEPA)}, a \textit{probabilistic} generalization that learns a predictive distribution over future latent states via a variational objective. We show that VJEPA unifies representation learning with Predictive State Representations (PSRs) and Bayesian filtering, establishing that sequential modeling does not require autoregressive observation likelihoods. Theoretically, we prove that VJEPA representations can serve as sufficient information states for optimal control without pixel reconstruction, while providing formal guarantees for collapse avoidance. We further propose \emph{Bayesian JEPA (BJEPA)}, an extension that factorizes the predictive belief into a learned dynamics expert and a modular prior expert, enabling zero-shot task transfer and constraint (e.g. goal, physics) satisfaction via a Product of Experts. Empirically, through a noisy environment experiment, we demonstrate that VJEPA and BJEPA successfully filter out high-variance nuisance distractors that cause representation collapse in generative baselines. By enabling principled uncertainty estimation (e.g. constructing credible intervals via sampling) while remaining likelihood-free regarding observations, VJEPA provides a foundational framework for scalable, robust, uncertainty-aware planning in high-dimensional, noisy environments.

自监督学习世界模型不确定性估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。