arXiv:2506.01884cs.LGcs.AI2025-06

提出无假设强化学习框架,揭示函数逼近下的理论极限与算法设计新方向。

Agnostic Reinforcement Learning: Foundations and Algorithms

  • 在无最优策略保证下,探索策略类中最佳策略的理论方法。
  • 发现环境访问、覆盖条件与表征结构三者间存在显著统计分离现象。
  • 适合关注强化学习理论边界与函数逼近挑战的研究者阅读。

强化学习在多个复杂领域展现出巨大成功,但对大规模状态空间下带函数逼近的强化学习统计复杂性仍缺乏坚实的理论理解。本文从学习理论视角,系统研究带函数逼近的强化学习统计复杂性。不同于以往工作,本文考虑最弱形式的函数逼近——无假设策略学习,即学习者在给定策略类Π中寻找最佳策略,但不保证Π包含任务最优策略。通过三个核心维度系统探索:环境访问方式、覆盖条件(衡量策略类Π中状态占据分布的扩展性)和表征条件(对策略类Π本身的结构性假设)。在此框架内,(1) 设计具有理论保证的新算法;(2) 刻画任意算法的性能下限。结果揭示了显著的统计分离现象,凸显无假设策略学习的优势与局限。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) has demonstrated tremendous empirical success across numerous challenging domains. However, we lack a strong theoretical understanding of the statistical complexity of RL in environments with large state spaces, where function approximation is required for sample-efficient learning. This thesis addresses this gap by rigorously examining the statistical complexity of RL with function approximation from a learning theoretic perspective. Departing from a long history of prior work, we consider the weakest form of function approximation, called agnostic policy learning, in which the learner seeks to find the best policy in a given class $Π$, with no guarantee that $Π$ contains an optimal policy for the underlying task. We systematically explore agnostic policy learning along three key axes: environment access -- how a learner collects data from the environment; coverage conditions -- intrinsic properties of the underlying MDP measuring the expansiveness of state-occupancy measures for policies in the class $Π$, and representational conditions -- structural assumptions on the class $Π$ itself. Within this comprehensive framework, we (1) design new learning algorithms with theoretical guarantees and (2) characterize fundamental performance bounds of any algorithm. Our results reveal significant statistical separations that highlight the power and limitations of agnostic policy learning.

强化学习函数逼近理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。