arXiv:2505.01041cs.LG2025-05IJCAI被引 3

证明单时标演员-评论家在连续空间中可达到近优解

Global Optimality of Single-Timescale Actor-Critic under Continuous State-Action Space: A Study on Linear Quadratic Regulator

  • 基于线性二次调节器问题,分析单时标演员-评论家方法
  • 样本复杂度达epsilon的-2阶,可在连续空间求得近优解
  • 填补理论与实践差距,适合强化学习理论研究者

演员-评论家方法在诸多挑战性任务中表现卓越,但其性能的理论理解仍不充分。现有研究多聚焦于双循环或双时标步长等非典型变体,且仅限于有限状态或动作空间的局部收敛性。本文将研究拓展至经典单样本单时标演员-评论家算法在连续(无限)状态-动作空间的情况,以标准线性二次调节器(LQR)问题为案例。结果表明,该方法在连续状态-动作空间求解LQR时,可达到epsilon最优解,样本复杂度为epsilon^-2阶。本工作为单时标演员-评论家的性能提供了新见解,进一步弥合了理论与实践之间的鸿沟。

原文摘要 · Abstract (English)

Actor-critic methods have achieved state-of-the-art performance in various challenging tasks. However, theoretical understandings of their performance remain elusive and challenging. Existing studies mostly focus on practically uncommon variants such as double-loop or two-timescale stepsize actor-critic algorithms for simplicity. These results certify local convergence on finite state- or action-space only. We push the boundary to investigate the classic single-sample single-timescale actor-critic on continuous (infinite) state-action space, where we employ the canonical linear quadratic regulator (LQR) problem as a case study. We show that the popular single-timescale actor-critic can attain an epsilon-optimal solution with an order of epsilon to -2 sample complexity for solving LQR on the demanding continuous state-action space. Our work provides new insights into the performance of single-timescale actor-critic, which further bridges the gap between theory and practice.

强化学习演员评论家理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。