arXiv:2601.00844cs.LGcs.AI2026-01被引 14

让JEPA模型学会用距离表示目标远近,提升决策规划能力

Value-guided action planning with JEPA world models

  • 通过约束嵌入空间,让状态间距离反映到达目标的代价
  • 在简单控制任务上规划性能显著优于传统JEPA模型
  • 适合需要高效环境推理与动作规划的研究者

构建能理解环境的深度学习模型,关键在于捕捉其内在动态。联合嵌入预测架构(JEPA)通过自监督预测目标学习表征与预测器,是建模动态的有前景框架。但其在支持有效动作规划方面仍受限。本文提出一种增强方案:通过训练使状态嵌入间的距离(或准距离)逼近给定环境中到达目标的负值函数(即抓取成本)。我们设计了一种实用方法,在训练中强制实现该约束,并在简单控制任务上验证,该方法显著提升了规划性能,优于标准JEPA模型。

原文摘要 · Abstract (English)

Building deep learning models that can reason about their environment requires capturing its underlying dynamics. Joint-Embedded Predictive Architectures (JEPA) provide a promising framework to model such dynamics by learning representations and predictors through a self-supervised prediction objective. However, their ability to support effective action planning remains limited. We propose an approach to enhance planning with JEPA world models by shaping their representation space so that the negative goal-conditioned value function for a reaching cost in a given environment is approximated by a distance (or quasi-distance) between state embeddings. We introduce a practical method to enforce this constraint during training and show that it leads to significantly improved planning performance compared to standard JEPA models on simple control tasks.

世界模型动作规划表示学习自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。