arXiv:2510.09487cs.LG2025-10中稿 · ICLR被引 1

提出新型模型化强化学习算法,实现近最优样本效率。

Near-Optimal Second-Order Guarantees for Model-Based Adversarial Imitation Learning

  • 基于模型的在线模仿学习框架,结合专家数据与免奖励交互
  • 理论证明样本复杂度达近最优,随策略方差减小而提升
  • 适合追求高效学习的强化学习研究者,尤其关注样本利用率

我们研究在线对抗性模仿学习(AIL),即智能体在无奖励环境下,从离线专家演示中学习并在线与环境交互。尽管已有很强的实证效果,但在线交互的优势及随机性影响仍不明确。本文提出一种基于模型的AIL算法(MB-AIL),在一般函数逼近下,建立了不依赖于时域长度的二阶样本复杂度保证。这些二阶界限可随相关策略回报方差变化,系统越趋确定性则越紧致。结合新构造的难例族上的信息论下界,我们证明:在有限专家演示下,MB-AIL在在线交互中的样本复杂度达到最小最大最优(对数因子内);在专家数据方面,其对时域$H$、精度$ε$和策略方差$σ^2$的依赖也匹配下界。实验验证了理论结果,且实际实现的MB-AIL在样本效率上优于或相当现有方法。

原文摘要 · Abstract (English)

We study online adversarial imitation learning (AIL), where an agent learns from offline expert demonstrations and interacts with the environment online without access to rewards. Despite strong empirical results, the benefits of online interaction and the impact of stochasticity remain poorly understood. We address these gaps by introducing a model-based AIL algorithm (MB-AIL) and establish its horizon-free, second-order sample-complexity guarantees under general function approximations for both expert data and reward-free interactions. These second-order bounds provide an instance-dependent result that can scale with the variance of returns under the relevant policies and therefore tighten as the system approaches determinism. Together with second-order, information-theoretic lower bounds on a newly constructed hard-instance family, we show that MB-AIL attains minimax-optimal sample complexity for online interaction (up to logarithmic factors) with limited expert demonstrations and matches the lower bound for expert demonstrations in terms of the dependence on horizon $H$, precision $ε$ and the policy variance $σ^2$. Experiments further validate our theoretical findings and demonstrate that a practical implementation of MB-AIL matches or surpasses the sample efficiency of existing methods.

模仿学习强化学习样本效率理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。