证明了策略梯度在广义效用强化学习中的全局最优性
On the Global Optimality of Policy Gradient Methods in General Utility Reinforcement Learning
- 基于状态-动作占用测度的凹效用函数,构建新证明方法
- 表格场景下实现全局最优,大规模场景样本复杂度仅依赖近似维度
- 适用于模仿学习、安全强化学习等场景,理论突破性强
广义效用强化学习(RLGU)为超越标准期望回报的问题提供了统一框架,涵盖模仿学习、纯探索和安全强化学习等。尽管近期在标准强化学习的策略梯度(PG)方法理论上取得进展,但其在RLGU中的应用理解仍有限。本文建立了在一般凹效用函数下的全局最优性保证。在表格设置中,采用基于梯度支配的新证明技术,推动对非直接策略参数化的分析。此外,在大规模状态-动作空间中,通过最大似然估计在函数逼近类中近似占用测度,实现全局最优性,样本复杂度仅随逼近类维度增长,而非状态-动作空间规模。
原文摘要 · Abstract (English)
Reinforcement learning with general utilities (RLGU) offers a unifying framework to capture several problems beyond standard expected returns, including imitation learning, pure exploration, and safe RL. Despite recent fundamental advances in the theoretical analysis of policy gradient (PG) methods for standard RL and recent efforts in RLGU, the understanding of these PG algorithms and their scope of application in RLGU still remain limited. In this work, we establish global optimality guarantees of PG methods for RLGU in which the objective is a general concave utility function of the state-action occupancy measure. In the tabular setting, we provide global optimality results using a new proof technique building on recent theoretical developments on the convergence of PG methods for standard RL using gradient domination. Our proof technique opens avenues for analyzing policy parameterizations beyond the direct policy parameterization for RLGU. In addition, we provide global optimality results for large state-action space settings beyond prior work which has mostly focused on the tabular setting. In this large scale setting, we adapt PG methods by approximating occupancy measures within a function approximation class using maximum likelihood estimation. Our sample complexity only scales with the dimension induced by our approximation class instead of the size of the state-action space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。