arXiv:2501.14622cs.LGcs.AI2025-01被引 6

将模仿学习与自监督学习结合,提升策略表征效率。

ACT-JEPA: Novel Joint-Embedding Predictive Architecture for Efficient Policy Representation Learning

  • 联合预测动作序列与隐变量观测序列,端到端训练。
  • 世界模型理解能力提升40%,任务成功率提高10%。
  • 适合需要高效策略学习与泛化能力的研究者。

在模仿学习(IL)中,高效策略表征的学习面临挑战:现有方法依赖昂贵的专家示范,且未显式建模环境,导致世界模型不完善。自监督学习(SSL)可从多样无标签数据中学习世界模型,但多数方法在原始输入空间运行,效率低下。本文提出ACT-JEPA,一种融合IL与SSL的新架构,通过端到端训练联合预测1)动作序列和2)隐变量观测序列。该模型利用联合嵌入预测架构(JEPA),在隐空间中过滤无关细节,构建稳健的世界模型。我们在多个环境和任务上评估,结果表明其在所有场景下均优于最强基线。相比基线,世界模型理解能力提升最高达40%,任务成功率最高提高10%。此外,预测隐变量观测序列能有效泛化至动作序列预测。本工作证明,结合IL与SSL可实现高效策略表征、改进世界模型并提升任务成功率。

原文摘要 · Abstract (English)

Learning efficient representations for decision-making policies is a challenge in imitation learning (IL). Current IL methods require expert demonstrations, which are expensive to collect. Additionally, they are not explicitly trained to understand the environment. Consequently, they have underdeveloped world models. Self-supervised learning (SSL) offers an alternative, as it can learn a world model from diverse, unlabeled data. However, most SSL methods are inefficient because they operate in raw input space. In this work, we propose ACT-JEPA, a novel architecture that unifies IL and SSL to enhance policy representations. It is trained end-to-end to jointly predict 1) action sequences and 2) latent observation sequences. To learn in latent space, we utilize Joint-Embedding Predictive Architecture, which allows the model to filter out irrelevant details and learn a robust world model. We evaluate ACT-JEPA in different environments and across multiple tasks. Our results show that it outperforms the strongest baseline in all environments. ACT-JEPA achieves up to 40% improvement in world model understanding and up to 10% higher task success rate. Finally, we show that predicting latent observation sequences effectively generalizes to predicting action sequences. This work demonstrates how integrating IL and SSL leads to efficient policy representation learning, an improved world model, and a higher task success rate.

模仿学习自监督学习世界模型策略表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。