用稀疏表示建模世界模型,能降低预测复杂度并揭示可解释的动态结构。
LpWM: A Case for Sparse Representations in World Models

- 通过稀疏编码替代稠密表示,实现更高效的动态建模。
- 在PushT任务上,稀疏模型规划成功率比稠密模型高57%。
- 适合需要低复杂度控制与可解释性的强化学习研究者。
联合嵌入预测架构(JEPAs)通过匹配特征到各向同性高斯等最大熵分布来学习潜在动态,产生稠密表示。然而,稠密表示是否最优尚不明确。本文探讨稀疏表示能否使动作条件下的潜在动态更易建模,并分析其产生的动力学结构。我们证明,在足够高维的一热隐空间中,非线性李普希茨动态可被动作条件下的线性动态任意逼近,且滚动误差随维度增加趋近于零。这启发了分布式稀疏表示作为一热稀疏性的实用放松。我们提出LpWorldModel(LpWM),通过修正广义高斯分布匹配正则化(RDMReg)将编码器特征匹配到非负稀疏码。实验表明,稀疏性降低了预测器所需复杂度:在PushT任务上,稀疏LpWM在中等预测器容量下规划成功率比稠密LeWM高57%。该优势也超越高斯匹配,LpWM在多种预测器族中优于稠密VICReg表示。此外,学习到的稀疏表示具有模式分解特性,支持编码离散动力学模式,特征幅值捕捉模式内连续状态。结果表明,稀疏表示可降低控制所需的预测器复杂度,并揭示可解释结构。
原文摘要 · Abstract (English)
Joint-embedding predictive architectures (JEPAs) learn latent dynamics for planning and avoid representation collapse by matching features to maximum-entropy distributions such as isotropic Gaussians, yielding dense representations. However, it is unclear whether dense representations are the most favorable geometry for modeling dynamics. In this work, we ask whether a different geometry, sparse representations, can make action-conditioned latent dynamics easier to model, and what dynamical structure emerges from such representations. We first show that nonlinear Lipschitz dynamics can be approximated arbitrarily well by action-conditioned linear dynamics in a sufficiently high-dimensional one-hot latent space, with rollout error vanishing as the dimension grows. This motivates distributed sparse representations as a practical relaxation of one-hot sparsity. We introduce LpWorldModel (LpWM), a JEPA model regularized with Rectified Distribution Matching Regularization (RDMReg) to match encoder features to a Rectified Generalized Gaussian distribution, yielding non-negative sparse codes. Empirically, sparsity lowers the predictor complexity required for successful planning: on PushT, sparse LpWM outperforms dense LeWM by up to 57% in planning success at intermediate predictor capacities. This advantage also extends beyond Gaussian distribution matching, with LpWM outperforming dense VICReg representations across multiple predictor families. We further find that the learned sparse representations are mode-factored, with support encoding discrete dynamical regimes and feature magnitudes capturing continuous within-regime state. Together, these results suggest that sparse representations can reduce the predictor complexity required for control while revealing interpretable structure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。