arXiv:2606.31232cs.AI2026-06被引 5

通过潜空间差异解码,让视觉世界模型更敏感于动作,避免表征坍塌。

Delta-JEPA: Learning Action-Sensitive World Models via Latent Difference Decoding

论文配图:Delta-JEPA: Learning Action-Sensitive World Models via Latent Difference Decoding
图 1 · 摘自论文原文
  • 用潜空间位移解码动作,替代传统拼接终点嵌入的方法。
  • 在4个连续控制任务中,规划性能优于基线模型,动作敏感性更强。
  • 无需像素重建,适合追求高效动作感知的强化学习研究者。

学习用于规划的视觉世界模型需要紧凑且对动作敏感的潜空间动态,但无重建的联合嵌入目标可能导致表征对动作不敏感。本文提出Delta-JEPA,一种端到端的无重建世界模型,通过引入潜空间差异动作解码器(LDAD)增强潜空间前向预测。不同于从拼接终点嵌入中反推动作的逆解码器,LDAD从连续观测间的潜空间位移中重构执行的动作。这种位移层面的监督直接规范了转换几何:相邻嵌入若坍塌将丢失动作信息,不同动作被鼓励引发可区分的潜空间变化,以支持基于滚动的规划。Delta-JEPA仅依赖潜空间预测和动作重构,避免像素重建与分布匹配正则化。在四个视觉连续控制任务中,其规划性能优于基于JEPA及代表性学习的世界模型基线。消融实验表明,基于位移的动作解码始终优于终点拼接,且动作敏感性分析显示潜空间响应更清晰地受动作条件调控。结果表明,对潜空间差异进行监督是一种简单有效的抗坍塌、高动作敏感性的世界模型学习机制。

原文摘要 · Abstract (English)

Learning visual world models for planning requires compact latent dynamics that remain sensitive to actions, yet reconstruction-free joint-embedding objectives can collapse to action-insensitive representations. We propose Delta-JEPA, an end-to-end reconstruction-free world model that augments latent forward prediction with a Latent Difference Action Decoder (LDAD). Unlike inverse decoders that infer actions from concatenated endpoint embeddings, LDAD reconstructs the executed action from the latent displacement between consecutive observations. This displacement-level supervision directly regularizes transition geometry: adjacent embeddings cannot collapse without losing action information, and different actions are encouraged to induce distinguishable latent changes for rollout-based planning. Delta-JEPA uses only latent prediction and action reconstruction, avoiding pixel reconstruction and distribution-matching regularizers. Across four visual continuous-control tasks, Delta-JEPA improves planning over JEPA-based and representation-learning world model baselines. Ablations show that displacement-based action decoding is consistently more effective than endpoint concatenation, and action-sensitivity analyses show clearer action-conditioned latent responses. These results indicate that supervising latent differences is a simple and effective mechanism for collapse-resistant and action-sensitive world model learning.

世界模型动作敏感潜空间建模强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。