arXiv:2510.04189cs.LGcs.AI2025-10

提出首个带约束的平均成本自然批评家-演员算法,实现非渐近收敛与更优采样效率。

Finite Time Analysis of Constrained Natural Critic-Actor Algorithm with Improved Sample Complexity

  • 基于函数逼近的自然批评家-演员框架,处理长期平均成本与不等式约束
  • 给出非渐近收敛保证,确定最优学习率并改进采样复杂度
  • 在Safety-Gym三环境中表现优于主流算法,适合安全强化学习场景

近期研究聚焦于演员-批评家(AC)算法的非渐近收敛分析。已有工作在折扣成本设定下提出一种双时间尺度的批评家-演员算法,其演员与批评家角色互换,但仅建立了渐近收敛性。随后,针对线性函数逼近的批评家-演员算法完成了渐近与非渐近分析。本文首次引入适用于长期平均成本设定且带不等式约束的函数逼近自然批评家-演员算法,并提供其非渐近收敛保证。我们的分析确定了最优学习率,并提出改进以提升采样复杂度。实验在三个Safety-Gym环境上验证,结果表明该算法在性能上可与现有知名算法相媲美。

原文摘要 · Abstract (English)

Recent studies have increasingly focused on non-asymptotic convergence analyses for actor-critic (AC) algorithms. One such effort introduced a two-timescale critic-actor algorithm for the discounted cost setting using a tabular representation, where the usual roles of the actor and critic are reversed. However, only asymptotic convergence was established there. Subsequently, both asymptotic and non-asymptotic analyses of the critic-actor algorithm with linear function approximation were conducted. In our work, we introduce the first natural critic-actor algorithm with function approximation for the long-run average cost setting and under inequality constraints. We provide the non-asymptotic convergence guarantees for this algorithm. Our analysis establishes optimal learning rates and we also propose a modification to enhance sample complexity. We further show the results of experiments on three different Safety-Gym environments where our algorithm is found to be competitive in comparison with other well known algorithms.

强化学习约束优化非渐近分析自然梯度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。