arXiv:2606.09311cs.AI2026-06被引 3

无需目标图像,用分层潜空间规划实现长时程任务求解

FF-JEPA: Long-Horizon Planning in World Models with Latent Planners

论文配图:FF-JEPA: Long-Horizon Planning in World Models with Latent Planners
图 1 · 摘自论文原文
  • 设计双前向动态模型,分步预测子目标与动作轨迹
  • 在PushT上实现长时程规划,避免传统方法的坍缩问题
  • 适合无明确目标状态的现实世界任务规划

联合嵌入预测架构(JEPAs)在世界建模中展现出潜力,可通过交叉熵法(CEM)等优化动作轨迹实现在潜空间中的规划。然而,这些方法在长时程规划中计算成本过高且效果不佳,通常还需显式的目标图像,这在真实任务中难以满足。本文提出前向-前向JEPA(FF-JEPA),采用分层结构,结合一个动作条件的前向模型与一个无动作的潜空间规划器,该规划器根据当前状态预测下一子目标。该方法无需目标图像,通过将复杂轨迹分解为一系列可处理的短期优化问题,实现长时程规划。在PushT上的初步结果表明,FF-JEPA成功克服了平坦世界模型在长时程下的塌陷问题,为无目标规划提供了一个有前景的方向。

原文摘要 · Abstract (English)

Joint Embedding Predictive Architectures (JEPAs) have shown promising world modeling capabilities, enabling planning in latent space by optimizing action trajectories using methods like the Cross-Entropy Method (CEM). These methods are, however, too computationally expensive and ineffective for long-horizon planning. Furthermore, these methods typically require an explicit image of the goal state, which is not always possible in real-world tasks. In this work, we tackle these limitations by proposing Forward-Forward-JEPA (FF-JEPA), a hierarchical approach leveraging two forward dynamics models. Alongside a standard action-conditioned forward model, we introduce an action-free latent planner that predicts the next subgoal given the current state. This approach removes the need for goal images and enables long-horizon planning by decomposing complex trajectories into a sequence of tractable, short-term optimization problems. Preliminary results on PushT demonstrate that FF-JEPA successfully overcomes flat world models' long-horizon collapse, highlighting this approach as a promising direction for goal-free planning.

世界模型长程规划潜空间无目标规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。