arXiv:2410.14081cs.LG2024-10ICML被引 6

无需奖励信号,用隐空间模型实现高维复杂任务的稳定模仿学习

Reward-free World Models for Online Imitation Learning

  • 在隐空间中建模环境动态,不依赖图像重建
  • 通过逆软Q学习实现稳定优化,达到专家级表现
  • 适合高维观测与复杂动力学场景,如机器人操控

模仿学习(IL)使智能体能直接从专家示范中获取技能,是强化学习的有力替代方案。然而,以往在线模仿学习方法在高维输入和复杂动态的任务上表现不佳。本文提出一种基于无奖励世界模型的新型在线模仿学习方法。该方法在隐空间中完全学习环境动态,无需图像重建,实现高效准确建模。采用逆软Q学习目标,将优化过程重构为Q策略空间,缓解传统奖励-策略空间优化中的不稳定性。结合学习到的隐空间动态模型与规划控制,该方法在高维观测或动作空间及复杂动力学任务中持续实现稳定且达到专家水平的表现。我们在DMControl、MyoSuite和ManiSkill2等多个基准上评估,结果表明其性能优于现有方法。

原文摘要 · Abstract (English)

Imitation learning (IL) enables agents to acquire skills directly from expert demonstrations, providing a compelling alternative to reinforcement learning. However, prior online IL approaches struggle with complex tasks characterized by high-dimensional inputs and complex dynamics. In this work, we propose a novel approach to online imitation learning that leverages reward-free world models. Our method learns environmental dynamics entirely in latent spaces without reconstruction, enabling efficient and accurate modeling. We adopt the inverse soft-Q learning objective, reformulating the optimization process in the Q-policy space to mitigate the instability associated with traditional optimization in the reward-policy space. By employing a learned latent dynamics model and planning for control, our approach consistently achieves stable, expert-level performance in tasks with high-dimensional observation or action spaces and intricate dynamics. We evaluate our method on a diverse set of benchmarks, including DMControl, MyoSuite, and ManiSkill2, demonstrating superior empirical performance compared to existing approaches.

模仿学习世界模型机器人控制无奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。