用隐空间扩散规划,让机器人从无动作示范和次优数据中学习
Latent Diffusion Planning for Imitation Learning
- 在隐空间中用扩散模型做规划,分离动作预测与策略生成
- 可在无动作示范和次优数据上训练,提升数据利用效率
- 适合复杂视觉操控任务,尤其当专家数据稀缺时
模仿学习的进展依赖于能处理复杂视觉运动任务、多模态分布和大规模数据的策略架构。然而,这些方法通常需要大量专家示范。为解决这一问题,我们提出隐空间扩散规划(LDP),一种模块化方法:规划器可利用无动作示范,逆动力学模型可利用次优数据,两者均在学习到的隐空间中运行。首先,通过变分自编码器学习紧凑隐空间,实现图像域中的未来状态有效预测。随后,使用扩散目标训练规划器和逆动力学模型。通过将规划与动作预测分离,LDP可受益于次优和无动作数据提供的更密集监督信号。在模拟的视觉机器人操控任务中,LDP优于现有顶尖模仿学习方法,因后者无法利用此类额外数据。
原文摘要 · Abstract (English)
Recent progress in imitation learning has been enabled by policy architectures that scale to complex visuomotor tasks, multimodal distributions, and large datasets. However, these methods often rely on learning from large amount of expert demonstrations. To address these shortcomings, we propose Latent Diffusion Planning (LDP), a modular approach consisting of a planner which can leverage action-free demonstrations, and an inverse dynamics model which can leverage suboptimal data, that both operate over a learned latent space. First, we learn a compact latent space through a variational autoencoder, enabling effective forecasting of future states in image-based domains. Then, we train a planner and an inverse dynamics model with diffusion objectives. By separating planning from action prediction, LDP can benefit from the denser supervision signals of suboptimal and action-free data. On simulated visual robotic manipulation tasks, LDP outperforms state-of-the-art imitation learning approaches, as they cannot leverage such additional data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。