arXiv:2606.27014cs.LG2026-06被引 2

首次为基于JEPA的世界模型建立泛化理论,揭示隐空间预测的优劣权衡。

A Generalization Theory for JEPA-Based World Models

论文配图:A Generalization Theory for JEPA-Based World Models
图 1 · 摘自论文原文
  • 将JEPA预训练建模为条件谱图学习问题,等价于动作相关共现矩阵的低秩分解。
  • 证明了预训练误差与下游规划遗憾之间的关联,给出有限样本泛化界。
  • 揭示隐空间维度影响近似误差与采样误差的内在权衡,适合研究世界模型理论者。

联合嵌入预测架构(JEPA)近期作为世界建模的新范式兴起,通过在隐空间中学习预测动态,而非在输入层面生成未来观测。尽管其在实践中表现良好,但对基于JEPA的世界模型的理论理解仍不充分。本文首次为基于JEPA的世界模型构建泛化理论。我们将JEPA预训练形式化为一个条件谱图学习问题,证明其目标等价于动作条件共现矩阵的低秩分解。基于这一刻画,我们建立了预训练误差与下游规划遗憾之间的联系,从而导出基于JEPA的世界模型的有限样本泛化界。分析揭示了隐空间维度在近似误差与样本误差间的固有权衡,为隐空间预测模型相较于输入级预测方法的优势与局限提供了理论洞察。

原文摘要 · Abstract (English)

Joint Embedding Predictive Architectures (JEPAs) have recently emerged as a promising paradigm for world modeling by learning predictive dynamics in a latent space rather than generating future observations at the input level. Despite their empirical success, the theoretical understanding of JEPA-based world models remains limited. In this paper, we develop the first generalization theory for JEPA-based world models. We formulate JEPA pretraining as a conditional spectral graph learning problem and show that the JEPA objective is equivalent to a low-rank factorization of an action-conditioned co-occurrence matrix. Building on this characterization, we establish a connection between JEPA pretraining error and downstream planning regret, leading to a finite-sample generalization bound for JEPA-based world models. Our analysis reveals an inherent trade-off between approximation and sample errors with respect to the latent dimension, providing theoretical insights into the advantages and limitations of latent predictive models compared with input-level predictive approaches.

世界模型泛化理论隐空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。