将隐空间分解为进展与内容两部分,提升模型对任务阶段的感知能力。
Subspace-Decomposed JEPAs: Disentangling Progression and Content in Latent World Models

- 用正弦夹角损失和信号正则化分别约束进展与内容子空间
- 1维角度进展坐标在40个测试场景中使事件定位准确率提升0.18 AUROC
- 仅占4.2%隐空间的8维进展子空间可解释72%-95%的任务进展方差
联合嵌入预测架构(JEPAs)通过预测未来嵌入学习紧凑的隐世界模型,但未指定任一隐变量编码任务进展。本文将JEPA隐空间划分为两个正交且角色分离的子空间:低维进展子空间由余弦夹角三元组损失塑造,高维内容子空间则由LeWM的SIGReg目标正则化。我们证明这两种防坍缩机制作用于不重叠的坐标,可加性地共存而非竞争。所提方法SD-JEPA在多数控制基准上优于LeWM基线,在匹配计算条件下表现更优;在Push-T任务上超越最强非LeWM JEPA基线。子空间消融实验确认该解耦是核心贡献。此外,1维角度进展坐标作为隐空间中的场景感知罗盘,随任务推进而变化,回溯时反向移动;在受控扰动下能突变并重新定位至语义合理的新任务阶段区,将意外时刻与其意义解耦——这是传统预测误差标量无法实现的。三个定量测试验证:|Δθ_t|在40个保留立方体情景中,事件定位的联合AUROC最高提升0.18,每项±1步容差下胜出率达97.5%;跨四环境的单线性探测显示,8维进展子空间(占隐空间4.2%)可解释72%-95%的任务进展方差。
原文摘要 · Abstract (English)
Joint-Embedding Predictive Architectures (JEPAs) learn compact latent world models by predicting future embeddings, but no single coordinate of the latent is designated to encode task progression. We carve the JEPA latent into two orthogonal subspaces with disjoint roles: a low-dimensional progression subspace shaped by a cosine-margin triplet loss, and a high-dimensional content subspace regularised by the existing SIGReg objective of LeWM. We prove that the two anti-collapse forces act on disjoint coordinates, so they compose additively rather than competing on the same dimensions. Our method, SD-JEPA improves over the LeWM baseline on the majority of its control benchmarks at matched compute, and outperforms the strongest non-LeWM JEPA baseline on Push-T; a subspace-ablation falsifier confirms the split is the load-bearing ingredient. Beyond planning, the resulting 1-D angular progression coordinate functions as a scene-aware compass on the latent. It advances with task progress, regresses when the agent backtracks, and under controlled perturbations both spikes and relocalises to a semantically appropriate new task-phase sector, separating the moment of surprise from its meaning in a way that prediction-error scalars cannot. Three quantitative tests back this up: $|Δθ_t|$ outperforms the standard latent-prediction-error surprise at localising semantic events on 40 held-out cube episodes by up to +0.18 pooled AUROC (97.5% per-episode win rate at $\pm 1$-step tolerance); a within-episode linear probe across all four environments (40 episodes per env) shows the 8-dimensional progression subspace (4.2% of the latent) explains 72-95% of task-progress variance..
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。