揭示离线强化学习的统计复杂性,给出紧致理论边界与高效算法。
On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures
- 基于实例依赖分析,构建近最优的离线策略评估与学习方法。
- 首次获得表格与线性场景下离线策略评估的紧致统计界。
- 适合关注强化学习理论基础与低适应探索的研究者。
本文综述了离线与低自适应强化学习(RL)的最新统计基础进展。首先论证离线RL是几乎所有真实机器学习问题的合适模型,即使与近期基于RL的AI突破无关。随后聚焦离线RL两大核心问题:离线策略评估(OPE)与离线策略学习(OPL)。令人惊讶的是,即便在表格和线性情形下,这些任务的紧致界直到最近才被建立。文章区分了最坏情况下的极小极大界与实例依赖界,并梳理了实现近最优实例依赖方法的关键算法思想与证明技术。最后,讨论离线RL的局限性,并介绍新兴的低适应探索问题,该问题通过在离线与在线之间提供折中方案,有效缓解现有局限。
原文摘要 · Abstract (English)
This article reviews the recent advances on the statistical foundation of reinforcement learning (RL) in the offline and low-adaptive settings. We will start by arguing why offline RL is the appropriate model for almost any real-life ML problems, even if they have nothing to do with the recent AI breakthroughs that use RL. Then we will zoom into two fundamental problems of offline RL: offline policy evaluation (OPE) and offline policy learning (OPL). It may be surprising to people that tight bounds for these problems were not known even for tabular and linear cases until recently. We delineate the differences between worst-case minimax bounds and instance-dependent bounds. We also cover key algorithmic ideas and proof techniques behind near-optimal instance-dependent methods in OPE and OPL. Finally, we discuss the limitations of offline RL and review a burgeoning problem of \emph{low-adaptive exploration} which addresses these limitations by providing a sweet middle ground between offline and online RL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。