arXiv:2502.11033cs.LGmath.OC2025-02ICML被引 1

突破传统限制,证明大规模策略优化可收敛至最优策略

Convergence of Policy Mirror Descent Beyond Compatible Function Approximation

  • 用变分梯度主导性替代强闭包条件,放宽理论适用范围
  • 首次给出非闭包策略类下策略镜面下降的收敛速率上界
  • 适合研究强化学习理论或复杂策略优化的学者参考

现代策略优化方法大多遵循策略镜面下降(PMD)的算法范式,已有大量理论收敛结果。但这些结果通常仅适用于表格型环境,或要求策略类满足强闭包条件,而大尺度环境下参数化策略类通常不满足该条件。本文为一般策略类构建了新的理论框架,将闭包条件替换为更弱的变分梯度主导性假设,并得到逼近最优策略类的收敛速率上界。核心结果基于一个新定义的平滑性概念——相对于当前策略占据度量诱导的局部范数,将PMD视为非欧空间中光滑非凸优化的特例。

原文摘要 · Abstract (English)

Modern policy optimization methods roughly follow the policy mirror descent (PMD) algorithmic template, for which there are by now numerous theoretical convergence results. However, most of these either target tabular environments, or can be applied effectively only when the class of policies being optimized over satisfies strong closure conditions, which is typically not the case when working with parametric policy classes in large-scale environments. In this work, we develop a theoretical framework for PMD for general policy classes where we replace the closure conditions with a strictly weaker variational gradient dominance assumption, and obtain upper bounds on the rate of convergence to the best-in-class policy. Our main result leverages a novel notion of smoothness with respect to a local norm induced by the occupancy measure of the current policy, and casts PMD as a particular instance of smooth non-convex optimization in non-Euclidean space.

强化学习策略优化理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。