arXiv:2411.00610cs.LG2024-11NeurIPS被引 7

首个理论高效且实用的通用函数逼近强化学习方法

Provably and Practically Efficient Adversarial Imitation Learning with General Function Approximation

  • 基于优化的奖励与乐观贝尔曼误差最小化,实现在线学习
  • 理论证明仅需多项式数量的专家样本和交互次数
  • 适合追求理论保证又需实际落地的强化学习研究者

对抗性模仿学习(AIL)在神经网络近似下取得了显著实践成功,但现有理论研究多局限于简化场景(如表格型或线性近似),且算法设计复杂,难以实用。本文研究了具有通用函数逼近的在线AIL的理论基础,提出新方法优化型AIL(OPT-AIL),通过在线优化奖励函数并结合乐观正则化贝尔曼误差最小化来更新Q值。理论上,证明OPT-AIL在学习接近专家策略时,具有多项式级别的专家样本复杂度和交互复杂度。据我们所知,OPT-AIL是首个具备通用函数逼近的理论高效对抗性模仿学习方法。实践中,仅需近似优化两个目标,便于实现。实验证明,OPT-AIL在多个挑战性任务中优于先前最先进的深度AIL方法。

原文摘要 · Abstract (English)

As a prominent category of imitation learning methods, adversarial imitation learning (AIL) has garnered significant practical success powered by neural network approximation. However, existing theoretical studies on AIL are primarily limited to simplified scenarios such as tabular and linear function approximation and involve complex algorithmic designs that hinder practical implementation, highlighting a gap between theory and practice. In this paper, we explore the theoretical underpinnings of online AIL with general function approximation. We introduce a new method called optimization-based AIL (OPT-AIL), which centers on performing online optimization for reward functions and optimism-regularized Bellman error minimization for Q-value functions. Theoretically, we prove that OPT-AIL achieves polynomial expert sample complexity and interaction complexity for learning near-expert policies. To our best knowledge, OPT-AIL is the first provably efficient AIL method with general function approximation. Practically, OPT-AIL only requires the approximate optimization of two objectives, thereby facilitating practical implementation. Empirical studies demonstrate that OPT-AIL outperforms previous state-of-the-art deep AIL methods in several challenging tasks.

模仿学习强化学习理论保证函数逼近

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。