揭示了控制型世界模型可识别性的关键条件,解决非线性观测下的动态恢复难题。
On the Identifiability of Controlled World Models

- 提出由表征可识别与转移可识别组成的联合识别条件。
- 证明预测误差上界与谱分离度成反比,反事实误差放大与动作激励度相关。
- 理论在四个非线性场景中验证,适用于视觉控制与潜在空间规划研究者。
世界模型是处理高维观测与候选动作下环境动态的有力工具。近期LeCun提出的JEPA框架为表示空间中学习此类模型提供了有效方案。其动作条件扩展在视觉控制与潜在空间规划中起核心作用,但存在根本问题:能否从非线性观测中恢复受控动态?本文针对高斯潜态的控制型世界模型,提出一个联合可识别性条件,包含两个耦合部分:(1) 表征可识别性,依赖谱分离性质;(2) 转移可识别性,与条件动作的非退化变化相关。证明当该条件满足时,最小化LeJEPA风格的预测目标可恢复潜态与受控动态(正交变换意义下)。进一步证明转移预测误差上界与谱分离裕度成反比,且反事实预测误差的可达到放大程度与最弱条件动作激发裕度成反比。理论预测在四种非线性观测设置中得到实证支持。
原文摘要 · Abstract (English)
World model serves as a promising tool to infer environment dynamics under high-dimensional observations and candidate actions. Recently, LeCun's JEPA provides a compelling framework for learning such models in representation space. Its action-conditioned extension plays a central role in visual control and latent-space planning, but leaves a fundamental question: can it recover the controlled dynamics from nonlinear observations? This paper presents a joint identifiability condition for controlled world models with Gaussian latent states, which consists of two coupled components: (1) representation identifiability and (2) transition identifiability. The former depends on the spectral separation property while the latter is related to non-degenerate variation of conditional action. We prove that when this condition holds, minimizing the LeJEPA-style predictive objective can recover both latent states and controlled dynamics in the sense of orthogonal transformation. We further prove that the upper bound of transition prediction error is inversely proportional to the spectral separation margin. We also characterize an attainable amplification of counterfactual prediction error that scales inversely with the weakest conditional action-excitation margin. The theoretical predictions are empirically supported across four nonlinear observation settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。