arXiv:2605.26379stat.MLcs.LG2026-05被引 11

LeJEPA可准确还原世界潜在变量,仅当潜变量服从高斯分布时成立。

When Does LeJEPA Learn a World Model?

论文配图:When Does LeJEPA Learn a World Model?
图 1 · 摘自论文原文
  • 通过对齐与高斯正则化,线性恢复非线性观测中的真实潜变量。
  • 在潜变量为高斯分布时,线性可识别性成立;其他分布均不满足该性质。
  • 理论验证适用于2D到1024维潜空间,适合构建可信赖的世界模型。

一种能混淆世界真实自由度的表示无法支持可靠规划或组合泛化。我们证明,在潜变量遵循平稳加性噪声转移的广泛世界中,LeJEPA(对齐+高斯正则化)能线性恢复世界的潜在变量,这一性质称为线性可识别性。核心结论是:在所有此类世界中,仅当潜变量服从高斯分布时,该保证才成立。正向证明基于谱分解,其中每个非线性度均被对齐严格惩罚,使线性映射成为最优解;反向证明排除了所有非高斯替代方案。此外,我们还证明了近似可识别性结果,其性能随扰动平滑下降,并表明线性正交可识别性支持最优潜空间规划。实验涵盖2D示例到1024维潜空间,包括分布消融和基于像素的机器人控制。本理论将经验上成功的配方转化为数学保证,为构建可证明还原世界结构的世界模型奠定基础。

原文摘要 · Abstract (English)

A representation that scrambles the true degrees of freedom of the world cannot support reliable planning or compositional generalization. We prove that LeJEPA (alignment plus Gaussian regularization) linearly recovers the world's latent variables from nonlinear observations, a property known as linear identifiability, in a broad class of worlds where latents evolve under stationary, additive-noise transitions. Our main result is that among all such worlds, the Gaussian is the unique latent distribution for which this guarantee holds. The forward direction rests on a spectral decomposition in which each degree of nonlinearity is strictly penalized by alignment, making the linear map the optimum; the converse rules out every non-Gaussian alternative. We further prove an approximate identifiability result where the guarantee degrades gracefully, and show that linear, orthogonal identifiability enables optimal latent-space planning. We validate the theory with experiments ranging from 2D examples to 1024-dimensional latents, including distributional ablations and pixel-based robotic control. Our theory turns an empirically successful recipe into a mathematical guarantee, providing the foundation for building World Models that provably recover the structure of the world.

世界模型可识别性表示学习高斯分布

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。