让机器人模型更懂物理,规划成功率最高达100%
Toward Physically Grounded JEPA World Models for Goal-Conditioned Robotic Planning

- 用逆动力学和状态对齐增强潜空间预测
- 在四个任务中成功率超98%,两任务达100%
- 适合做目标导向的机器人规划研究者
动作条件下的JEPA世界模型可在不重建未来像素的情况下实现视觉目标规划,但仅靠潜空间预测无法确保表征对机器人控制有用。本文提出端到端的JEPA世界模型,引入逆动力学(IDM)和状态对齐(SA)。IDM防止潜空间坍缩,使潜移变与动作相关;状态对齐将连续表示锚定于实际物理构型与运动。在四个基准任务中,模型在TwoRoom(100%)、PushT(98%)和OGBench-Cube(87%)上表现最佳,Reacher上与LeWorldModel相当。消融实验表明,状态对齐在所有任务中均提升规划成功率。尽管LeWorldModel在OGBench-Cube上平均线性化更高,其转移子空间维数显著更低;本模型有效转移维度更高,优于仅用IDM的模型,验证了状态对齐对机器人规划的有效补充作用。
原文摘要 · Abstract (English)
Action-conditioned JEPA world models enable planning toward visually specified goals without reconstructing future pixels, yet latent prediction alone does not explicitly encourage the learned representations to retain information relevant to robotic control. We introduce an end-to-end JEPA world model that augments latent prediction with inverse dynamics (IDM) and state alignment (SA). While inverse dynamics discourages latent collapse and makes latent transitions informative of the actions that produced them, state alignment grounds consecutive representations in their associated physical configuration and motion. Across four benchmark tasks, our model attains the highest success rates on TwoRoom (100%), PushT (98%), and OGBench-Cube (87%), while performing comparably to LeWorldModel on Reacher. Our ablation further shows that adding state alignment consistently improves planning success over IDM alone across all four tasks. Although LeWorldModel, our primary baseline, attains higher average straightening on OGBench-Cube, transition-subspace analysis shows that its transition energy is concentrated in a substantially lower-dimensional subspace. Our state-aligned model exhibits a higher effective transition dimension than LeWorldModel and improves planning over IDM alone, supporting state alignment as an effective complement to inverse dynamics for robotic planning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。