提出适用于大/连续动作空间的离线强化学习新方法,突破传统算法限制。
Beyond State-Wise Mirror Descent: Offline Policy Optimization with Parametric Policies
- 将镜面下降法拓展至参数化策略,解决大动作空间下的优化难题
- 揭示上下文耦合是核心障碍,建立与自然策略梯度的新关联
- 理论统一离线RL与模仿学习,适合实际部署的策略参数化场景
我们研究了在一般函数逼近下离线强化学习的理论问题。已有工作(如Xie等,2021)建立了通过悲观性从离线数据学习好策略的理论基础,但现有计算可行的算法(如PSPI)仅适用于有限小动作空间,且依赖状态级镜面下降,要求智能体由评论家函数隐式推导,无法适配实践中普遍存在的独立策略参数化。本文克服这些局限,将理论保证扩展至大或连续动作空间的参数化策略类。在将镜面下降推广至参数化策略时,我们识别出上下文耦合为核心难点,并发现将镜面下降与自然策略梯度相连接可带来新的分析、保证与算法洞察,包括离线强化学习与模仿学习之间令人意外的统一。
原文摘要 · Abstract (English)
We investigate the theoretical aspects of offline reinforcement learning (RL) under general function approximation. While prior works (e.g., Xie et al., 2021) have established the theoretical foundations of learning a good policy from offline data via pessimism, existing algorithms that are computationally tractable (often in an oracle-efficient sense), such as PSPI, only apply to finite and small action spaces. Moreover, these algorithms rely on state-wise mirror descent and require actors to be implicitly induced from the critic functions, failing to accommodate standalone policy parameterization which is ubiquitous in practice. In this work, we address these limitations and extend the theoretical guarantees to parameterized policy classes over large or continuous action spaces. When extending mirror descent to parameterized policies, we identify contextual coupling as the core difficulty, and show how connecting mirror descent to natural policy gradient leads to novel analyses, guarantees, and algorithmic insights, including a surprising unification between offline RL and imitation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。