arXiv:2608.06799cs.ROcs.CV2026-08被引 2

让世界模型的隐变量更贴近物理状态,提升机器人规划与控制性能

Is Forward Prediction Enough? Physical State Grounding for JEPA World Models

论文配图:Is Forward Prediction Enough? Physical State Grounding for JEPA World Models
图 1 · 摘自论文原文
  • 在预测基础上增加双层物理约束:个体隐变量对齐本体感知,隐变量差值对齐关节变化
  • 在模拟和真实机器人上均超越现有基线,在规划与策略学习中表现更优
  • 适合关注机器人控制、世界模型可解释性的研究者使用

构建结构化且与控制相关的潜在表示仍是世界模型的核心挑战。基于JEPA的世界模型通过观测序列学习动作条件下的预测潜在动态,但其前向预测目标未显式保证从单个潜在变量或潜在对中可靠识别出以机器人为中心的物理状态及其变化,这可能限制下游规划与策略性能。我们提出PSG-JEPA,一种具有物理基础的JEPA世界模型,在训练阶段引入两个互补的接地目标:将个体潜在变量与机器人本体感知状态对齐,将潜在对与多时域关节变化对齐。两个目标仅在训练阶段应用,推理架构与计算成本保持不变。为全面评估PSG-JEPA,我们在三个层面进行实验:(1) 通过探针测试潜在可识别性;(2) 在冻结潜在变量上进行目标条件规划;(3) 在仿真和真实机器人上进行策略学习。实验表明,PSG-JEPA在所有三个层面均持续优于现有顶尖潜伏世界模型基线。

原文摘要 · Abstract (English)

Learning structured and control-relevant latent representations remains a key challenge for world models. Recent JEPA-based world models learn action-conditioned predictive latent dynamics from observation sequences. However, their forward-prediction objectives do not explicitly enforce reliable identifiability of robot-centric physical state from individual latents or state changes from latent pairs, which can limit downstream planning and policy performance. We propose PSG-JEPA, a physically grounded JEPA world model that shapes its latent space with two complementary grounding objectives beyond forward prediction: grounding individual latents in robot proprioceptive state, and grounding latent pairs in multi-horizon joint-angle changes. Both objectives are applied only during training, leaving the inference architecture and computational cost unchanged. To comprehensively evaluate PSG-JEPA, we conduct experiments at three levels: (1) latent identifiability via probing, (2) goal-conditioned planning on frozen latents, and (3) policy learning in simulation and on a real robot. Experiments demonstrate that our PSG-JEPA consistently outperforms state-of-the-art latent world-model baselines at all three levels.

世界模型机器人控制物理接地潜在空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。