arXiv:2501.19116cs.LGstat.ML2025-01ICML被引 7

解释了异步演员-评论家算法为何在部分可观测环境更有效

A Theoretical Justification for Asymmetric Actor-Critic Algorithms

  • 用有限时间收敛分析证明异步评论家能消除状态混淆误差
  • 理论表明异步设计可加速学习,尤其在状态信息不完整时
  • 适合研究强化学习理论或设计高效训练算法的研究者

在部分可观测环境中,许多成功的强化学习算法采用异步学习范式,利用训练时的额外状态信息以加快学习速度。尽管这些方法的学习目标通常具有理论基础,但其潜在优势仍缺乏精确的理论解释。本文通过将有限时间收敛分析适配到该设置,为使用线性函数逼近器的异步演员-评论家算法提供了理论依据。所得的有限时间界显示,异步评论家能够消除由智能体状态中的混淆(aliasing)引起的误差项。

原文摘要 · Abstract (English)

In reinforcement learning for partially observable environments, many successful algorithms have been developed within the asymmetric learning paradigm. This paradigm leverages additional state information available at training time for faster learning. Although the proposed learning objectives are usually theoretically sound, these methods still lack a precise theoretical justification for their potential benefits. We propose such a justification for asymmetric actor-critic algorithms with linear function approximators by adapting a finite-time convergence analysis to this setting. The resulting finite-time bound reveals that the asymmetric critic eliminates error terms arising from aliasing in the agent state.

强化学习理论分析演员评论家

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。