arXiv:2507.14444stat.MLcs.AI2025-07被引 3

系统梳理强化学习的算法与理论基础,聚焦样本效率与计算效率问题。

Statistical and Algorithmic Foundations of Reinforcement Learning

  • 以马尔可夫决策过程为框架,整合模型、价值、策略优化三类主流方法。
  • 分析在线、离线、带人类反馈等场景下的样本复杂度与计算效率瓶颈。
  • 适合关注RL理论机制或算法设计的研究者阅读。

作为未知环境中序列决策的范式,强化学习近年来备受关注。然而,新兴应用中模型复杂度激增及非凸性问题加剧了在数据稀缺情境下实现高效强化学习的挑战,此类情境下数据收集成本高、耗时长或风险大(如临床试验、自动驾驶、在线广告)。因此,理解并提升强化学习算法的样本效率与计算效率具有重要意义。本教程旨在介绍强化学习中若干重要的算法与理论进展,突出新思想与经典主题的联系。以马尔可夫决策过程为核心数学模型,涵盖多种典型强化学习场景(即带模拟器的RL、在线RL、离线RL、鲁棒RL、带人类反馈的RL),并介绍主流方法(模型驱动、价值驱动、策略优化)。讨论围绕非渐近视角下的样本复杂度、计算效率,以及算法依赖与信息论下界展开。

原文摘要 · Abstract (English)

As a paradigm for sequential decision making in unknown environments, reinforcement learning (RL) has received a flurry of attention in recent years. However, the explosion of model complexity in emerging applications and the presence of nonconvexity exacerbate the challenge of achieving efficient RL in sample-starved situations, where data collection is expensive, time-consuming, or even high-stakes (e.g., in clinical trials, autonomous systems, and online advertising). How to understand and enhance the sample and computational efficacies of RL algorithms is thus of great interest. In this tutorial, we aim to introduce several important algorithmic and theoretical developments in RL, highlighting the connections between new ideas and classical topics. Employing Markov Decision Processes as the central mathematical model, we cover several distinctive RL scenarios (i.e., RL with a simulator, online RL, offline RL, robust RL, and RL with human feedback), and present several mainstream RL approaches (i.e., model-based approach, value-based approach, and policy optimization). Our discussions gravitate around the issues of sample complexity, computational efficiency, as well as algorithm-dependent and information-theoretic lower bounds from a non-asymptotic viewpoint.

强化学习理论分析样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。