提出新型世界模型,让机器在信息不全时更准确预测多种未来。
UWM-JEPA: Predictive World Models That Imagine in Belief Space

- 用密度矩阵潜变量和酉预测器建模信念空间,保持不确定性不变。
- 五步预测任务中准确率达0.77,扰动动作时性能平稳下降。
- 适合需要多路径推理的强化学习与部分可观测场景。
针对部分观测环境中的世界模型需想象多种兼容隐藏未来并响应反事实动作的问题,传统联合嵌入预测架构(JEPAs)在向量潜空间中难以保留信念连续性。本文提出单位酉世界模型JEPA(UWM-JEPA),采用联合系统-环境空间上的密度矩阵潜变量及学习到的酉预测器,使联合状态谱在滚动预测中精确保持,避免不确定性耗散。在目标观测掩码下的五步前向模拟任务中,UWM-JEPA达到0.77准确率,动作扰动时性能单调下降;而参数匹配的LSTM-JEPA在相同条件下退化至0.53多数类准确率。盲滚动测试中,UWM-JEPA在短时程下探针R²损失不足10点,向量潜变量基线分别损失41和68点;两者在保留上下文探针上表现一致,差异源于预测器而非编码器。行动敏感性需通过反事实目标训练实现,该结论不限于酉参数化。对JEPA世界模型而言,潜空间几何与预测动态至关重要,而非仅依赖静态上下文编码能力。
原文摘要 · Abstract (English)
World models for partially observed environments must imagine multiple compatible hidden futures and steer between them under counterfactual actions. Joint Embedding Predictive Architectures (JEPAs) do this in latent space, but a vector-valued latent has no internal structure for carrying the belief over hidden continuations through blind rollout. We introduce the Unitary World Model JEPA (UWM-JEPA), a JEPA world model with a density-matrix latent on a joint system-environment space and a learned unitary predictor. The construction preserves the joint-state spectrum exactly during rollout, so the predictor itself cannot dissipate the represented uncertainty. On a hidden-velocity indicator task requiring five-step forward simulation under a given action sequence with the target observation masked, UWM-JEPA reaches 0.77 accuracy and degrades monotonically as actions are perturbed; a parameter-matched LSTM-JEPA trained under the same counterfactual-target objective and action head collapses to majority-class accuracy (0.53) under every action condition. Under blind rollout, UWM-JEPA loses fewer than ten points of probe R^2 at short horizons while vector-latent baselines lose forty-one and sixty-eight; both nevertheless tie on a held-out context probe, locating the separation in the predictor rather than the encoder. Action sensitivity itself requires training against counterfactual rather than teacher-forced targets, a finding that applies beyond the unitary parameterisation. For JEPA world models to imagine under partial observability, latent geometry and predictor dynamics matter, not frozen context-encoding capacity alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。