用隐空间动态模型规划,能高效处理无奖励离线数据并适应新环境。
Learning from Reward-Free Offline Data: A Case for Planning with Latent Dynamics Models
- 用JEPA训练隐空间动态模型,基于模型进行规划。
- 在低质量数据下仍能实现接近顶尖无模型方法的轨迹拼接性能。
- 适合数据有限或环境多变的任务,比强化学习更省数据。
AI长期目标是让智能体能在多种环境中解决未见任务。主流方法有强化学习(试错)和最优控制(基于模型规划)。本文系统评估了二者在离线设置下的表现,使用不同质量的导航任务数据集。对比了目标条件与零样本强化学习方法,以及基于JEPA训练的隐空间动态模型进行规划的方法。结果表明:无模型强化学习依赖大量高质量数据;而基于模型的规划在未见布局上泛化能力更强,数据效率更高,且轨迹拼接性能媲美领先无模型方法。尤其在处理次优离线数据和多样环境时,隐空间动态模型规划表现出显著优势。
原文摘要 · Abstract (English)
A long-standing goal in AI is to develop agents capable of solving diverse tasks across a range of environments, including those never seen during training. Two dominant paradigms address this challenge: (i) reinforcement learning (RL), which learns policies via trial and error, and (ii) optimal control, which plans actions using a known or learned dynamics model. However, their comparative strengths in the offline setting - where agents must learn from reward-free trajectories - remain underexplored. In this work, we systematically evaluate RL and control-based methods on a suite of navigation tasks, using offline datasets of varying quality. On the RL side, we consider goal-conditioned and zero-shot methods. On the control side, we train a latent dynamics model using the Joint Embedding Predictive Architecture (JEPA) and employ it for planning. We investigate how factors such as data diversity, trajectory quality, and environment variability influence the performance of these approaches. Our results show that model-free RL benefits most from large amounts of high-quality data, whereas model-based planning generalizes better to unseen layouts and is more data-efficient, while achieving trajectory stitching performance comparable to leading model-free methods. Notably, planning with a latent dynamics model proves to be a strong approach for handling suboptimal offline data and adapting to diverse environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。