揭示JEPA世界模型的理论基础,证明SIGReg可实现主动推理的精确自由能优化。
The SIGReg Objective as Variational Free Energy: A Theoretical Active-Inference Account of JEPA World Models
- 提出熵估计层级框架,用先验错配度判断正则化项是否适配主动推理
- 证明使用SIGReg时目标函数等价于信息瓶颈,且潜空间代价精准反映主动推理价值
- 指出现有模型缺失未来状态覆盖信号,为后续实验提供理论指引
联合嵌入预测架构(JEPAs)是潜空间世界模型的主流设计,但通常仅以实证性能作为依据,缺乏规范性原理支持。本文表明,反坍塌正则化器的选择决定了JEPA训练目标(预测损失加加权嵌入正则项)是否构成有效的主动推理(AIF)变分自由能。我们将四种非对比正则化器(VICReg、LogDet、PairDist、SIGReg)组织成基于先验错配间隙的熵估计层级,发现该间隙符号决定主动推理惊喜上界能否成立:VICReg和LogDet为不安全的上界,PairDist为安全下界,而SIGReg完全消除间隙。进一步证明对应定理:在标准恒定噪声编码器模型下,若成功施加SIGReg(各向同性高斯嵌入),间隙消失,目标函数成为精确信息瓶颈,惊喜上界得以保留,潜空间目标代价成为主动推理功利价值的精确代理;而VICReg则残留不可消除的二阶各向异性项。该对应关系扩展至多步期望自由能、集成认知价值及学习策略场景,并识别出当前所有JEPA世界模型均未计算的唯一一个主动推理项——状态认知价值,即未来状态覆盖信号。这些预测在本质上而非程度上存在差异,且作为理论结论留待独立工作验证;完整证明见附录A,所有结果的代数核心已在Lean 4中机器验证(附录D)。
原文摘要 · Abstract (English)
Joint-Embedding Predictive Architectures (JEPAs) are the dominant design for latent world models, yet they are usually justified by empirical performance rather than a normative principle. We show that the choice of anti-collapse regulariser determines whether a JEPA's training objective, a prediction loss plus a weighted embedding regulariser, is a valid Active Inference (AIF) variational free energy. We organise four non-contrastive regularisers (VICReg, LogDet, PairDist, and SIGReg) into an entropy-estimator hierarchy indexed by a prior-miscalibration gap, and show that the gap's sign, whether the estimator bounds the latent entropy from above or below, decides whether the AIF surprise bound survives: VICReg and LogDet are unsafe upper bounds, PairDist a safe lower bound, and SIGReg eliminates the gap. We then prove a correspondence theorem: under the standard constant-noise encoder model and successful SIGReg enforcement (isotropic-Gaussian embeddings), the gap vanishes, the objective becomes an exact information bottleneck, the surprise bound is preserved, and the latent goal cost becomes an exact proxy for AIF pragmatic value, whereas VICReg leaves an irreducible second-order anisotropy term. We extend the correspondence to multi-step expected free energy, ensemble epistemic value, and a learned-policy regime, and we identify the one AIF term no current JEPA world model computes: the state-epistemic value, a future-state coverage signal. The predictions differ in kind, not degree, and are stated here as theoretical consequences left for empirical test in separate work; full proofs are in Appendix A, and the algebraic core of every result is machine-verified in Lean 4 (Appendix D).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。