arXiv:2411.06965cs.LGcs.AI2024-11被引 2

让智能体从少量示范中学习多样高质行为,突破传统模仿学习单一行为局限。

Imitation from Diverse Behaviors: Wasserstein Quality Diversity Imitation Learning with Single-Step Archive Exploration

  • 基于WAE的对抗训练提升质量多样性模仿学习稳定性
  • 单步档案探索奖励机制缓解行为过拟合,提升泛化性能
  • 在MuJoCo连续控制任务上达到或超越专家表现

从有限示范中学习多样且高性能的行为是一项重大挑战。传统模仿学习方法通常只能学习单一行为,即使提供多个示范也难以突破。为此,本文提出基于瓦瑟斯坦质量多样性优化的模仿学习方法(WQDIL),通过基于瓦瑟斯坦自编码器(WAE)的潜在对抗训练提升质量多样性设置下的模仿学习稳定性,并采用基于度量条件的奖励函数与单步档案探索奖励机制,有效缓解行为过拟合问题。实验表明,该方法在源自MuJoCo环境的复杂连续控制任务上显著优于现有顶尖模仿学习方法,实现了接近或超越专家水平的质量多样性表现。

原文摘要 · Abstract (English)

Learning diverse and high-performance behaviors from a limited set of demonstrations is a grand challenge. Traditional imitation learning methods usually fail in this task because most of them are designed to learn one specific behavior even with multiple demonstrations. Therefore, novel techniques for \textit{quality diversity imitation learning}, which bridges the quality diversity optimization and imitation learning methods, are needed to solve the above challenge. This work introduces Wasserstein Quality Diversity Imitation Learning (WQDIL), which 1) improves the stability of imitation learning in the quality diversity setting with latent adversarial training based on a Wasserstein Auto-Encoder (WAE), and 2) mitigates a behavior-overfitting issue using a measure-conditioned reward function with a single-step archive exploration bonus. Empirically, our method significantly outperforms state-of-the-art IL methods, achieving near-expert or beyond-expert QD performance on the challenging continuous control tasks derived from MuJoCo environments.

模仿学习质量多样性连续控制强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。