arXiv:2505.19946cs.LG2025-05NeurIPS被引 6

提出新算法 extsc{SPOIL},在离线模仿学习中实现专家级性能。

Inverse Q-Learning Done Right: Offline Imitation Learning in $Q^π$-Realizable MDPs

  • 基于$Q^π$-可实现性假设,设计鞍点优化算法
  • 仅需$ ilde{ m O}(\varepsilon^{-2})$样本即可逼近专家性能
  • 适用于深度神经网络,比行为克隆更优

研究马尔可夫决策过程中的离线模仿学习问题,目标是利用由专家策略生成的状态-动作对数据集学习出高性能策略。不同于以往假设专家属于已知易处理策略类的工作,本文从新角度出发,引入环境结构假设:在线性$Q^π$-可实现的MDP中,提出一种名为 extsc{SPOIL}的新算法,保证在$ ilde{ m O}(

原文摘要 · Abstract (English)

We study the problem of offline imitation learning in Markov decision processes (MDPs), where the goal is to learn a well-performing policy given a dataset of state-action pairs generated by an expert policy. Complementing a recent line of work on this topic that assumes the expert belongs to a tractable class of known policies, we approach this problem from a new angle and leverage a different type of structural assumption about the environment. Specifically, for the class of linear $Q^π$-realizable MDPs, we introduce a new algorithm called saddle-point offline imitation learning (\SPOIL), which is guaranteed to match the performance of any expert up to an additive error $\varepsilon$ with access to $\mathcal{O}(\varepsilon^{-2})$ samples. Moreover, we extend this result to possibly nonlinear $Q^π$-realizable MDPs at the cost of a worse sample complexity of order $\mathcal{O}(\varepsilon^{-4})$. Finally, our analysis suggests a new loss function for training critic networks from expert data in deep imitation learning. Empirical evaluations on standard benchmarks demonstrate that the neural net implementation of \SPOIL is superior to behavior cloning and competitive with state-of-the-art algorithms.

离线学习模仿学习强化学习策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。